AI engineer interview questions in 2026 test whether you can ship reliable products on top of large language models: prompting, retrieval-augmented generation (RAG), tool-calling agents, evaluation, and cost and latency control. Most loops pair a standard coding round with an LLM system design question, usually "design a RAG system" or "design an agent," and a practical build. You are judged on engineering judgment and measurement, not on model training math.
Key Takeaways
- An AI engineer builds applications on foundation models; an ML engineer trains and serves models. Interview questions follow that split.
- The most common design prompt is a RAG system. Strong answers cover chunking, hybrid retrieval, reranking, citations, and an evaluation plan with real metrics.
- Evals are the senior signal. Saying how you would measure retrieval recall, faithfulness, and regressions matters more than naming frameworks.
- Know the fundamentals well enough to make decisions: tokens drive cost, context length affects quality and latency, and temperature changes determinism.
- Agent questions focus on control: tool schemas, loop limits, failure handling, and when not to use an agent at all.
- Practical rounds usually mean building a small RAG or tool-calling app in Python in 60 to 120 minutes.
What Is an AI Engineer, and How Is It Different From an ML Engineer?
An AI engineer is a software engineer who builds products on top of pretrained foundation models, usually through APIs or open-weight models served in-house. The work is retrieval pipelines, prompt and tool design, agents, evaluation harnesses, and production concerns like caching, rate limits, and cost.
An ML engineer trains, tunes, and serves models. Their interviews cover feature engineering, loss functions, offline and online metrics, and training infrastructure. If that is the role you are targeting, the machine learning engineer interview guide covers it in depth. The roles overlap at fine-tuning and model serving, but the interview emphasis is different.
| Dimension | AI engineer (LLM apps) | ML engineer |
|---|---|---|
| Core question | How do I make a model reliable in a product? | How do I train a better model? |
| Typical design prompt | Design a RAG assistant, design an agent | Design a recommendation or ranking system |
| Fundamentals tested | Tokens, context, sampling, embeddings | Bias-variance, losses, regularization, features |
| Evaluation | Faithfulness, retrieval recall, LLM-as-judge, human review | AUC, precision/recall, A/B tests on model versions |
| Coding focus | Python services, async I/O, API integration | Data pipelines, training loops, PyTorch |
| Practical round | Build a RAG or tool-calling app | Train or debug a model on a dataset |
| Production concerns | Token cost, latency, prompt injection, rate limits | Training cost, feature freshness, drift |
Backend and full-stack engineers already have most of what the loop tests. The usual gap is evaluation and retrieval quality, not math.
What Does an AI Engineer Interview Loop Look Like?
Most AI engineer loops have four to six stages, and the exact mix depends heavily on company size. Frontier labs like OpenAI and Anthropic run their own distinct processes for applied roles; the pattern below is typical for product companies and startups hiring for LLM features.
- Recruiter screen (30 minutes): background, what you have shipped with LLMs, and level calibration.
- Technical screen (45 to 60 minutes): a coding problem, often practical (parse and chunk documents, call an API with retries) rather than pure algorithms.
- LLM system design (45 to 60 minutes): design a RAG assistant, a support agent, a document extraction pipeline, or a code review bot.
- Practical or take-home: build something working against a model API.
- Project deep dive: walk through an LLM system you built and defend its decisions.
- Behavioral: ownership, ambiguity, and working with product on fuzzy requirements.
Company size changes the mix, so ask your recruiter which version you are getting.
LLM Fundamentals Questions
Fundamentals questions check that your design choices come from understanding, not habit. You do not need to derive attention, but you do need to connect each concept to cost, latency, or quality.
What is a token, and why does it matter?
A token is the unit a model reads and writes, typically a word fragment from a subword tokenizer such as byte-pair encoding. API pricing, context limits, and generation latency are all measured in tokens. Code, numbers, and non-English text often use more tokens per character than English prose.
What is the context window, and is bigger always better?
The context window is the maximum number of tokens the model can handle in one request, prompt and output included. Bigger is not always better. Long prompts cost more, slow time to first token, and models can use information in the middle of very long inputs less reliably. The senior answer: retrieve the few passages that matter, and test quality at the context lengths you actually use.
How do temperature and top-p work?
Temperature scales the probability distribution over next tokens. Low values make output more deterministic; high values make it more varied. Top-p (nucleus sampling) restricts sampling to the smallest set of tokens whose cumulative probability reaches p. For extraction, classification, and tool calls, use low temperature. For brainstorming or creative copy, raise it. Even temperature zero is not guaranteed to be fully deterministic on many hosted APIs, so evals should not assume identical outputs.
What are embeddings?
An embedding is a dense vector representation of text where semantically similar inputs land close together, usually compared with cosine similarity or dot product. Embeddings power semantic search in RAG. Follow-up questions often probe their limits: embeddings are weak at exact matches like product codes, error IDs, and names, which is why hybrid search exists.
Why is generation slower than reading the prompt?
Prefill processes input tokens in parallel, while decode produces output one token at a time, reusing a KV cache of earlier tokens. So output length usually drives latency more than input length, and prompt caching on repeated prefixes cuts cost and time to first token.
RAG Interview Questions
RAG questions are the center of most AI engineer loops. Retrieval-augmented generation is a pattern where the system retrieves relevant documents at query time and passes them to the model as context, so answers are grounded in your data rather than only in the model's training.
The questions interviewers ask most:
- How do you choose a chunk size? Small chunks match precisely but lose context; large chunks keep context but dilute relevance and cost more tokens. Chunk on document structure, add overlap and metadata, then tune with evals.
- Why use hybrid search? Dense vectors catch paraphrases; keyword search such as BM25 catches exact terms. Combining them, often with reciprocal rank fusion, improves recall on real queries that mix both.
- What does a reranker do? A cross-encoder reranker scores each query-passage pair jointly and is more accurate than vector similarity, but slower. Retrieve broadly (say, top 50), rerank, and pass the top few to the model.
- How do you handle a question the docs do not answer? Set a relevance threshold, instruct the model to say it does not know, and log these queries as content gaps.
- RAG or fine-tuning? Use RAG for knowledge that changes or must be cited. Use fine-tuning for format, tone, or narrow task behavior. They are complementary, not rivals.
- How do you keep the index fresh? Incremental ingestion keyed by content hash, propagated deletes, and a re-embed plan for new embedding models.
Worked Example: Design a RAG System for Support Docs
This is the prompt you are most likely to get, so here is a complete answer at the depth interviewers expect. Budget about 45 minutes in a real round.
Step 1: Clarify requirements
Ask before drawing anything. A reasonable set of assumptions to state out loud:
- Corpus: about 5,000 help center articles plus internal runbooks, updated daily.
- Users: customers in a chat widget, plus support agents using an internal tool.
- Quality bar: answers must cite sources; wrong answers about billing or account security are worse than "I don't know."
- Latency target: first token under about 2 seconds, full answer under about 8.
- Access control: internal runbooks must never reach customers.
Naming the access control requirement early is a strong signal.
Step 2: Ingestion pipeline
Pull documents from the CMS on change events and parse HTML into clean text, keeping heading structure. Chunk by section, roughly 300 to 800 tokens with small overlap, and store metadata: doc ID, title, section path, URL, product area, audience (public or internal), and last-updated time. Write embeddings to a vector index and text to a keyword index. A content hash prevents re-embedding unchanged chunks.
Step 3: Query path
- Rewrite the user's message into a standalone query using conversation history, so "what about on mobile?" becomes a complete question.
- Apply a metadata filter for audience before retrieval. Filtering after retrieval risks leaking internal content into logs or prompts.
- Run hybrid retrieval (vector plus BM25), fuse results, and take the top 30 to 50.
- Rerank with a cross-encoder and keep the top 4 to 6 chunks.
- If the best rerank score is below a threshold, return a fallback with links to search and an escalation path.
- Generate with a system prompt that says: answer only from the provided sources, cite chunk IDs, say so when the sources do not cover the question.
- Post-process: verify that every citation ID exists in the context, render links, and stream the response.
def answer(question: str, history: list[dict], user: User) -> Answer:
query = rewrite_standalone(question, history)
audience = "internal" if user.is_agent else "public"
candidates = hybrid_search(query, filters={"audience": audience}, k=40)
ranked = rerank(query, candidates)[:5]
if not ranked or ranked[0].score < RELEVANCE_THRESHOLD:
return fallback_answer(query)
draft = generate(SYSTEM_PROMPT, query, context=ranked)
cited = validate_citations(draft, ranked)
return cited if cited.ok else fallback_answer(query)
Step 4: Evaluation strategy
This is where you separate yourself from other candidates. Split evaluation into retrieval and generation, because a bad answer can come from either.
| Layer | Metric | How to measure |
|---|---|---|
| Retrieval | Recall@k | Labeled set of questions with the chunk IDs that answer them; check if any appears in top k |
| Retrieval | MRR | Rank of the first correct chunk, averaged across questions |
| Generation | Faithfulness | Every claim supported by a cited chunk; judged by an LLM grader calibrated against human labels |
| Generation | Answer correctness | Compare to a reference answer on the labeled set |
| Behavior | Abstention accuracy | On questions the docs cannot answer, does it decline instead of guessing? |
| Product | Deflection and escalation rate, thumbs down rate | Online, segmented by product area |
Build the labeled set from real support tickets: sample a few hundred, have agents mark the correct article, and include unanswerable questions on purpose. Run the suite on every change to chunking, prompts, models, or retrieval, and block releases that regress faithfulness. Spot-check any LLM judge against human labels.
Step 5: Production concerns
Cache answers for frequent normalized queries, and use prompt caching for the fixed system prompt. Route simple queries to a smaller, cheaper model and reserve the larger model for long or ambiguous ones. Log query, retrieved IDs, scores, answer, and feedback for every request, with PII redaction. Treat retrieved text as untrusted to reduce prompt injection risk, and never let the model take account actions in this design. For caching layers and invalidation details, the patterns from distributed cache design apply directly.
For the general framework behind any design round, see how to ace the system design interview.
Blanking on recall@k versus MRR halfway through an LLM design round is a real risk when the interviewer pushes on evals. TechScreen listens to the question and shows structured answer points on your screen, invisible during screen share on Zoom, Google Meet, and Teams. Start with 3 free tokens, no credit card.
AI Agent Interview Questions
An AI agent is a system where a model decides which actions to take, calls tools, observes results, and repeats until a goal is met or a limit is hit. Interviewers want to see that you can keep that loop under control.
Common questions and what a strong answer covers:
- How does tool calling work? You describe tools with a name, description, and JSON schema. The model returns a structured call; your code validates arguments, executes, and returns the result as a new message. The model never executes anything itself.
- When would you not use an agent? When the steps are known in advance. A fixed workflow (classify, retrieve, generate) is cheaper, faster, and easier to test than an open loop. Saying this unprompted shows maturity.
- How do you stop runaway loops? Cap iterations and total tokens, set per-tool timeouts, detect repeated identical calls, and return a partial result with an explanation when limits are reached.
- How do you make tools safe? Least-privilege credentials, allowlisted actions, server-side validation of every argument, and human confirmation for anything irreversible such as refunds, deletes, or emails.
- How do you defend against prompt injection? Treat tool output and retrieved content as data, not instructions, and limit what an injected instruction could actually do.
- How do you evaluate an agent? Score task success on a fixed scenario set, plus trajectory checks: right tools, reasonable step count, no forbidden actions, cost per task.
- What is MCP? The Model Context Protocol is an open protocol, introduced by Anthropic, for exposing tools and data sources to models through a standard client-server interface.
Evals and Hallucination Mitigation
Evaluation is the topic that most clearly separates mid-level from senior AI engineer candidates. An eval is a repeatable test that scores model or system output against expected behavior, and a good team runs evals the way a backend team runs unit and integration tests.
Expect questions like "How do you know a prompt change made things better?" The answer has four parts:
- A versioned dataset of real inputs with expected outputs or grading rubrics, including edge cases and adversarial inputs.
- Graders matched to the task: exact match or schema checks for structured output, code execution for code, LLM-as-judge with a written rubric for open-ended text, and human review for a sample.
- A regression gate in CI so prompt and model changes cannot ship if key scores drop.
- Online signals after release: user feedback, escalation rates, and sampled transcripts reviewed by humans.
For hallucinations, layer the defenses: ground in retrieved sources with citations, allow abstention, constrain output format, verify high-stakes claims, and measure the rate on a labeled set. "We lowered temperature" alone is a weak answer.
Cost, Latency and Production Questions
Production questions test whether you have run an LLM feature with real traffic. Interviewers often ask, "Your LLM feature costs too much and is too slow. What do you do?"
| Lever | Effect on cost | Effect on latency | Trade-off |
|---|---|---|---|
| Smaller model for easy requests (routing) | Large reduction | Faster | Needs a router and evals per route |
| Prompt caching for shared prefixes | Lower input cost | Faster time to first token | Prompt structure must keep the stable part first |
| Response caching | Large on repeated queries | Near instant on hits | Staleness, needs invalidation |
| Fewer, better retrieved chunks | Lower | Faster | Requires good reranking |
| Shorter outputs, structured formats | Lower | Faster | Less verbose answers |
| Streaming | None | Better perceived latency | More client complexity |
| Batch APIs for offline jobs | Lower on providers that discount batch | Slower, asynchronous | Only for non-interactive work |
Also cover rate limits (backoff and queueing), a fallback provider for outages, timeouts, and idempotent retries. The model is an external dependency you do not control.
Prompt Engineering Interview Questions
Prompt engineering questions in 2026 are about discipline, not clever wording. Interviewers ask how you structure a system prompt, when you use few-shot examples, and how you get reliable JSON.
Good answers: delimit user content from instructions, add few-shot examples only where the model gets the format wrong, use the provider's structured output features instead of parsing free text, and version prompts with an eval run attached to each change. For chain-of-thought, ask whether the extra reasoning improves eval scores enough to justify the added tokens and latency.
Take-Home and Practical Rounds
Practical rounds show whether you can actually ship. Typical formats include building a small RAG app over a provided document set, building a tool-calling assistant against a mock API, or improving a deliberately weak pipeline and reporting what changed.
What graders look for:
- A working end-to-end path before any polish.
- A small eval script with a handful of labeled questions and printed scores. This is often the single biggest differentiator.
- Error handling around model calls: timeouts, retries, malformed output.
- A short README explaining trade-offs and what you would do with more time.
- Clean Python. Brush up with these Python interview questions if async code or typing is rusty.
Some companies allow AI coding assistants in these rounds, so read the rules. The AI-enabled coding interview guide covers how those rounds are graded, and the take-home assignment guide covers scoping and submission.
How to Prepare in Four Weeks
A focused four-week plan works for most engineers moving from backend or full-stack roles:
- Week 1: Fundamentals (tokens, context, sampling, embeddings) and your target provider's docs on tool calling and structured output.
- Week 2: Build a RAG app over a real corpus with hybrid search, reranking, and citations. Write the eval script with recall@k and faithfulness.
- Week 3: Add a tool-calling agent feature with limits and validation. Practice two LLM system design prompts out loud, timed.
- Week 4: Mock interviews, project deep-dive rehearsal, and behavioral stories about shipping under ambiguity.
The project matters more than the reading. Be ready to explain your chunk size, your eval numbers before and after a change, and the failure case that still bothers you.
LLM design rounds move fast, and follow-ups on evals, agents, and cost can come in any order. TechScreen gives you real-time, invisible answer support during live AI engineer interviews on Zoom, Google Meet, Teams, HackerRank, and CoderPad. Try it with 3 free tokens, no credit card required.
Frequently Asked Questions
What is the difference between an AI engineer and an ML engineer interview?
An ML engineer interview focuses on training models: feature engineering, loss functions, model selection, training pipelines, and ML system design for ranking or recommendation. An AI engineer interview focuses on building products on top of foundation models: prompting, retrieval-augmented generation, tool calling, agents, evaluation, and the cost and latency of model APIs. AI engineer loops still include normal coding and system design, but the domain questions assume you call a model rather than train one.
Do AI engineer interviews include LeetCode?
Many do, though usually lighter than a pure software engineering loop. Larger companies often keep one or two standard coding rounds at medium difficulty. Startups more often replace them with a practical round, such as building a small RAG pipeline or an agent with tool calls in 60 to 90 minutes, or a take-home. Expect at least one round where you write working Python against a real or mocked model API.
How do you answer 'how would you reduce hallucinations' in an interview?
Give a layered answer. First, ground the model: retrieve relevant sources, instruct it to answer only from them, and require citations. Second, constrain the output with structured formats and allow an explicit 'I don't know' path. Third, verify: check that cited passages support each claim, use a judge model or rules for high-risk outputs, and route low-confidence answers to a human. Finally, measure the hallucination rate on a labeled eval set so changes are tested, not guessed.
What should I build to prepare for an AI engineer interview?
Build one end-to-end project you can defend in detail: a RAG app over a real document set with hybrid search, reranking, citations, and an evaluation script that reports retrieval recall and answer faithfulness on 50 to 100 labeled questions. Add one tool-calling feature with input validation. Interviewers care less about the idea and more about whether you can explain your chunking choice, your failure cases, and what your eval numbers changed.
Is prompt engineering still asked in AI engineer interviews in 2026?
Yes, but rarely as a standalone topic. Interviewers fold it into larger questions: how you structure a system prompt, when you use few-shot examples, how you get reliable structured output, and how you version and test prompts. Treating prompts like code, with version control and regression evals, is the answer that signals seniority. Clever phrasing tricks without measurement is the answer that signals inexperience.
How technical are LLM fundamentals questions for AI engineers?
You need working knowledge, not research depth. Expect to explain tokens and tokenization, context windows and why long contexts degrade, temperature and top-p sampling, embeddings and cosine similarity, and the rough idea of attention and KV caching as it affects latency and cost. You will rarely be asked to derive math, but you will be asked how these properties change your design decisions.
Ready to use AI assistance in your next interview?
TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.
Ace your next interview →