Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 23 min read

The AI Engineer Interview - Complete Deep Dive

Stage 8 - Getting Hired Lesson 36 of 36

Prerequisites: Evals, RAG End to End, AI Security, Portfolio Projects Used in: the last page of this track - practice the adjacent rounds at System Design, Machine Coding, and DSA Build it: Lesson 24 - The Agentic AI Interview implements this as runnable, tested code you can execute offline.


What is the AI Engineer Interview?

It is a software engineering interview loop with one round added and one criterion changed.

The round added is AI system design: design a feature whose core component is a probabilistic model you did not train. The criterion changed runs through every other round. Instead of β€œdoes it produce the correct output,” the question becomes β€œhow do you know it produces acceptable output often enough to ship, and what happens on the occasions it does not.” That shift is the whole interview, and it is why a candidate can be a strong backend engineer and still fail this loop: they answer as though the model were a deterministic dependency.

Processes vary a great deal between companies, so treat everything here as round types rather than a fixed sequence. Most loops assemble some subset of five: a coding round, an AI system design round, a practical build, a debugging or case round, and a depth conversation about your own work.

Real-world analogy: interviewing a structural engineer versus interviewing a bridge inspector. The first is asked to design something that holds. The second is asked how they would know whether it holds β€” what they measure, how often, what reading would make them close the bridge. An AI engineer interview is mostly the second conversation, and candidates who prepared for the first one get surprised.

The good news, stated early: the distinctive round has a structure you can learn, and the single highest-signal behaviour in it is one you fully control. Bring up measurement before anyone asks you to.


The Round Types

Round What is graded How to prepare
Coding Ordinary software engineering - problem decomposition, correct data structures, clean readable code, tests, handling edge cases. Usually not ML at all. Normal practice at DSA. Expect an average-difficulty problem and be fluent rather than clever. Some companies swap in a component build - practise that at Machine Coding.
AI system design Whether you reason about quality, latency, cost, and failure as engineering constraints, and whether you raise measurement without being prompted. The structure below, plus general design fluency from System Design and the worked AI design at /hld/ChatGPT/.
Practical build or take-home Working code under time pressure, and the judgement of what to cut. Almost always whether you wrote any evaluation at all. Build the small thing end to end, then spend the last third of your time on a tiny eval set and a README with numbers. Most submissions have neither.
Debugging or case round Diagnostic discipline. Do you look at data before changing code, and can you isolate a component in a pipeline with several plausible culprits. The triage order below. Practise on your own project by reading its failures rather than its score.
Project depth conversation Whether your claimed experience survives mechanism questions, and whether you can state measured results. The measured artifact from Portfolio Projects. Know your own numbers cold.

Two things worth internalising from that table. The coding round is usually ordinary software engineering, so do not let AI preparation crowd out fundamentals β€” candidates lose loops on the plain coding round more often than they expect. And in the take-home, writing an eval set is not extra credit; it is frequently the discriminator between two submissions that both work.


The AI System Design Round

This is the distinctive round, so it deserves the most preparation. It looks like a normal system design interview and is graded differently. Throughput and sharding matter less; quality measurement, token cost, tail latency, and failure behaviour matter far more.

Follow an explicit order. The order itself is signal, because it shows you know which decisions constrain which.

flowchart TD
    A[Clarify the task and who sees a failure] --> B[Set the quality bar and the latency and cost budgets]
    B --> C[Choose prompting or retrieval or tuning with reasons]
    C --> D[Design the retrieval path if there is one]
    D --> E[Specify the eval strategy - offline suite and online signals]
    E --> F[Only now choose models and routing against that suite]
    F --> G[Guardrails - injection defence and output validation and abstention]
    G --> H[Observability - traces and quality and cost dashboards]
    H --> I[Rollout - flag and canary and defined rollback trigger]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class A client
    class B,C data
    class D service
    class E async
    class F service
    class G,H edge
    class I client

Note where E sits. Specifying the eval strategy before discussing model choice is the move that separates strong candidates from average ones, and it is not a presentation trick. If you have not said how quality will be measured, then β€œwe will use a strong model and fall back to a cheaper one” is an unfalsifiable statement β€” there is no instrument that could tell you whether the fallback was acceptable. Candidates who name models first spend the rest of the round unable to justify any tradeoff, because every justification needs a number they never defined.

1. Clarify the task and who sees a failure

Ask what the feature does and, more importantly, what a wrong answer costs and who encounters it. An internal draft that a specialist reviews before sending tolerates a very different error rate from an answer shown directly to a customer, and that single fact drives the grounding strategy, the abstention policy, and whether a human sits in the loop.
πŸ’‘ Abstention is the system answering β€œI do not know” rather than guessing, and counting that as a success when the evidence genuinely is not there.
Also pin down volume, corpus size and update rate, languages, and whether the data is sensitive.

2. Establish the quality bar and the budgets

Give the numbers shape even when the interviewer has not supplied them: β€œI would want to agree a target success rate on a labelled suite, a p95 budget in the low seconds for an interactive surface, and a cost per request ceiling the business can defend at expected volume.” Name the three dimensions explicitly β€” quality, p95, cost per request β€” and say that a change failing any one of them fails. Grounding the budgets in traffic estimates is ordinary back of the envelope work.
πŸ’‘ A labelled suite is a fixed list of real inputs each paired with the answer you have agreed is correct, which is what turns β€œit seems good” into a number.

3. Choose between prompting, retrieval, and tuning

State the decision rule rather than the answer. Missing knowledge means retrieval. Missing behaviour β€” format adherence, tone, a narrow task done consistently β€” means tuning. Ambiguous instructions mean a better prompt. Most production systems land on prompting plus retrieval, with tuning arriving later to cut cost or lock in a format once a labelled dataset exists as a byproduct of the eval work. Deep dives: Prompt Engineering, RAG End to End, Fine-Tuning.

4. Design the retrieval path

Parsing and chunking, embeddings, the index, hybrid retrieval, reranking, metadata filters for permissions and document versions, and how the index gets updated. Say out loud that retrieval quality caps answer quality β€” no model choice repairs a retriever that did not return the evidence. See Document Parsing and Chunking, Vector Databases, Hybrid Search and Reranking.

5. Specify the eval strategy

Where the round is won. Say how the labelled set is built and from what. Say that retrieval and generation are scored separately so a regression localises, and which dimensions get deterministic assertions while others need a rubric judge validated against human labels. Then say that latency and cost ceilings are treated as correctness, and which online signals close the loop. Then commit to it: the suite gates the build and a regression blocks release. Mechanics in Evals and LLM-as-Judge.
πŸ’‘ A rubric judge is a second model handed a written scoring guide, used to grade answers that have no single correct version.

6. Now choose models and routing

With a suite defined, model choice becomes a measurement instead of an opinion. Route by difficulty, keep a fallback provider for availability, and justify every choice against the three budgets. See Model Selection and Routing, Tokens and Cost Math, Latency Engineering, Prompt and Semantic Caching.

7. Guardrails, observability, rollout

Input and output validation, structured outputs with a repair path, abstention when evidence is thin, least privilege on any tool the model can reach, and an honest position on prompt injection. Then per-request traces you can replay β€” see Tracing and Observability and Observability β€” and a rollout behind a flag with a canary slice and a stated rollback trigger, which is ordinary deployment reliability applied to a non-deterministic component.
πŸ’‘ A canary slice is a small share of real traffic sent to the new version first, so a problem shows up on one percent of users rather than all of them.


How to Open One of These Questions

Suppose the prompt is: design a system that reads incoming support tickets, summarises them, and routes them to the right team.

A weak opening starts naming things β€” an orchestration framework, a vector store, a model. A strong opening spends its first ninety seconds like this:

β€œBefore I draw anything, three questions. Who sees a mis-route β€” does the ticket land with the wrong team and sit there, or does a human confirm the routing first? That changes how much I need abstention versus accuracy. Second, is this interactive or a background job, because a background job gives me seconds of budget and lets me batch. Third, how many teams and how stable is that list, since routing into a fixed small set is a classification problem I can measure exactly, while free-form routing is not.

Assuming mis-routes are visible to customers as delay and the list of teams is stable, I would frame routing as classification over a closed set, which gives me a reference-based eval for free, and summarisation as the open-ended part that needs a rubric. That split matters because I can gate routing on precision and recall from day one, and I only need a judge for the summary. Let me sketch the path and then say how I would measure each half.”

Nothing exotic happened there. The candidate reframed one half of the task into something with a ground truth, which made it measurable, and committed to a measurement strategy before drawing a box. That is the behaviour the round is looking for. The other half of the signal is negative: notice that no framework or vendor was named, because none was needed yet.

Expect to be steered off the order, and let it happen. Interviewers interrupt to probe the thing they care about, and fighting to finish your sequence reads worse than following them. Answer the question where they put it, then return with one sentence β€” β€œthat covers the retrieval path, so let me come back to how I would measure it” β€” so the eval step still lands before model choice. The order is a checklist you carry, not a script you perform.


Recurring Themes and Answer Shapes

Each of these has a deep dive. In the room, give the shape and the reasoning, then offer depth if the interviewer wants it.

Question Answer shape Depth
How do you know it works? A labelled set from real traffic, tiered checks, retrieval and generation scored apart, deltas against production rather than absolute scores, online signals feeding back. Evals
RAG or fine-tuning? Different problems. Knowledge gaps and freshness mean retrieval. Behaviour, format, and cost reduction on a narrow task mean tuning. They compose - tune the behaviour and retrieve the facts. Fine-Tuning
How do you stop hallucination? You reduce and detect it rather than solve it. Ground answers in retrieved evidence, require citations and verify them against what was retrieved, make abstention a first-class success outcome, and measure a grounding rate. Hallucination and Grounding
How do you handle prompt injection? Admit there is no complete fix, then reduce blast radius - least privilege on tools, an allowlist of actions, no privileged action without confirmation, treat all retrieved text as untrusted input. AI Security
How do you cut cost? Measure where tokens go first. Then shorten context, cache aggressively, route easy requests to smaller models, and consider distillation once you have labelled data - each verified on the same suite. Model Selection, Prompt and Semantic Caching
A provider updates the model - how do you avoid breaking? Pin versions, keep a frozen suite you can re-run against any candidate, hold an abstraction over the provider, canary the new version on a traffic slice, and keep a fallback path. Deploying AI Features
When do you use an agent? When the number of steps genuinely cannot be known in advance. If the path is knowable, write the pipeline - it is cheaper, faster, and testable. Bound any agent with a step budget and idempotent tools. Agent Architectures, Tool Calling
How do you evaluate retrieval separately from generation? Score the retriever on recall at k, precision at k, and rank of first relevant chunk. Score generation against a frozen context so retrieval noise is held constant. The pair localises the regression. Evals, RAG End to End

The pattern across all eight: name the mechanism and the measurement, never the vendor.


The Debugging Round

You are handed a badly-behaving AI system and asked what you would do. The grading is almost entirely about order β€” whether you gather evidence before touching anything.

  1. Reproduce with a trace. Get one concrete failing request and its full trace: input, retrieved chunks with scores, the assembled prompt, model and version, output, latency breakdown, token counts. Without a replayable trace you are guessing. If no such trace exists, say that instrumenting it is step zero.
  2. Check what actually reached the model. Read the assembled prompt. A large share of β€œthe model is wrong” bugs are β€œthe evidence was never in the context” bugs, and no prompt or model change fixes those. Truncation, a silently failing filter, and a stale index all show up here.
  3. Split retrieval from generation. Was the right chunk retrieved at all, and if so was it ranked where the model would read it? Low recall with honest output is a retrieval bug. Good recall with unsupported output is a grounding or prompt bug. These live in different code and have different fixes.
    πŸ’‘ Recall here just asks whether the right passage made it into the retrieved set at all, separately from how well the model then used it.
  4. Read and cluster a sample of failures. Thirty to fifty, one line each on what went wrong, clustered afterwards rather than into categories you decided in advance. Then count and sort. This converts β€œquality is bad” into β€œforty percent of failures are superseded document versions,” which is an actionable statement.
  5. Only now change something. One change, re-run the same suite, keep or revert on the delta. Also check the boring causes before the interesting ones: a recent prompt edit, a provider version change, a reindex, a config difference between environments, or retry behaviour hiding timeouts.

The failure mode the round is testing for is jumping to step five β€” β€œI would try a stronger model and add chain of thought” β€” in the first sentence. It is the most common way to lose this round, and stating your order explicitly before you start is enough to avoid it.


Talking About Your Own Projects

Lead with the measured result, not the architecture. Two sentences on what the system does and what the numbers say, then stop and let the interviewer choose where to go deep.

Have four things ready. Your numbers β€” quality on your suite, p95, cost per request β€” because a candidate who cannot state a cost or latency figure for their own system reads as someone who never operated it. One thing you rejected and why, with the delta that made you reject it. Your largest failure cluster, named and quantified, which is the strongest available proof you did error analysis. And your limitations, offered before you are asked.

Avoid inflation. Claiming a system is production ready when it has served your own test traffic invites exactly the questions that expose it, and interviewers calibrate on precision of language. β€œI tested it single-user; p95 under concurrency is unmeasured” costs you nothing and buys you credibility for every other claim you make.

If the work was done on a team, be exact about your part. β€œI owned the retrieval path and the eval suite; a colleague built the ingestion pipeline” is a stronger sentence than an ambiguous β€œwe,” because it tells the interviewer which half of the system they can go deep on. Vagueness about ownership is usually read as a claim that will not survive questioning.


Red Flags Interviewers Listen For

Red flag What it signals
No evals anywhere in the answer or the project The candidate cannot distinguish improvement from regression, so none of their quality claims mean anything.
Treating the model as magic - β€œthe model will figure it out” No mental model of why these systems fail, therefore no ability to design around it.
Cannot state a cost or latency figure for their own work Never operated the system under a constraint. The two things that most often kill an AI feature were never considered.
Claiming hallucination is solved Either not paying attention or overselling. The credible position is reduce, detect, and abstain.
Name-dropping frameworks without mechanism Tooling familiarity substituting for understanding. Collapses on the first question about how it works.
Naming models before defining quality Optimising a variable with no objective function. Every later tradeoff becomes unjustifiable.
Unbounded agents with privileged tools No blast-radius thinking, which is disqualifying for anything touching real data or money.

None of these require exotic knowledge to avoid. They are avoided by measuring, by reading failures, and by describing mechanisms rather than products.


A Preparation Plan

Sequence the track rather than reading it in an arbitrary order. Each stage builds on the previous one.

  1. Orientation and models first. Topics 1 to 9 β€” an accurate mental model, API mechanics, tokens and cost, prompting, structured outputs. Build one single-call feature. If you are short on time, the Crash Course route on the AI Engineer Track hub is this plus the essentials of retrieval and evals.
  2. Retrieval next, because most systems are retrieval systems. Topics 10 to 15, ending with a working pipeline over your own messy corpus.
  3. Evaluation third, and do not skip it. Topics 16 to 20. This is the stage that turns the rest into engineering, and the stage interviewers probe hardest.
  4. Agents and production fourth. Topics 21 to 24 and 29 to 34, weighted toward whatever the roles you are targeting actually do.
  5. Build the measured project. The full brief is in Portfolio Projects. Baseline, three deliberate changes, recorded deltas. This doubles as your depth-round preparation.
  6. Then the other rounds, in parallel. Keep DSA practice running the whole time, since the coding round is ordinary software engineering. Practise general design at System Design, read /hld/ChatGPT/ as the worked AI design, and shore up the distributed-systems fundamentals that serving and reliability rest on at Concepts β€” caching, rate limiting, message queues, idempotency, circuit breakers, and durable execution all show up in AI system design answers.
  7. Rehearse out loud. Run the seven-step order on three different prompts until the sequence is automatic and you stop reaching for tool names.

The honest timeline is months rather than weeks, and the largest variable is how early you start measuring. Everything else on this track is a knob; the measurement is the gauge.


When to Use

βœ… You are ready to interview when:

❌ Hold off when:


Common Interview Questions

Q1: Design a feature that answers customer questions from our product documentation. Where do you start?

With who sees a wrong answer and what it costs, because a customer-facing answer and an internally reviewed draft justify different designs. Then I would agree the three budgets - a target success rate on a labelled suite, a p95 for an interactive surface, and a cost per request ceiling at expected volume. Documentation questions are a knowledge problem, so retrieval rather than tuning, and I would design parsing, chunking, hybrid retrieval with reranking, and metadata filters for product version and permissions. Then, before touching model choice, the eval strategy. A labelled set from real support questions, retrieval and generation scored separately, and deterministic checks on citations and format. After that a validated rubric judge for helpfulness, with abstention counted as a success when the docs do not answer. Model selection and routing come after that, because only then can I prove a cheaper model is acceptable.

Q2: How do you decide between prompting, retrieval, and fine-tuning?

By asking what is actually missing. If the model lacks facts - private, recent, or specific to a customer - that is retrieval, and no amount of tuning fixes a knowledge gap reliably. If it has the facts but behaves wrong - drifting format, wrong tone, inconsistent on a narrow repeated task - that is tuning, or sometimes just a clearer prompt and a schema constraint. In practice I start with prompting plus retrieval, because the iteration loop is minutes rather than hours. I reach for tuning for one of two concrete reasons: locking in behaviour that prompting keeps losing, or moving a proven task onto a smaller cheaper model. Tuning also needs a labelled dataset, and if I built evals properly I already have the beginnings of one.

Q3: Our assistant occasionally invents policy details. How would you debug it?

I would not change the prompt first. I would get traces for real failing requests and read what actually reached the model, because the most common cause is that the policy text was never in the context - wrong chunk, truncation, a filter that silently excluded it, or a stale index. That splits the problem: if recall is low, it is a retrieval bug and I work on chunking, hybrid search, and reranking. If the right evidence was in context and the output still went beyond it, it is a grounding bug and I work on citation requirements verified against retrieved spans, context ordering, and making abstention an acceptable answer. Then I sample thirty failures, cluster them, and fix the largest cluster - and I would expect superseded document versions to be in there, because near-duplicate policies across years are a classic source of confidently grounded wrong answers.

Q4: The provider is deprecating the model you build on and the replacement behaves differently. What is your plan?

This is a regression test problem, which is why the frozen suite exists. I run the candidate model against the same labelled set and compare per-category, not just in aggregate, because a new version typically improves some clusters and quietly breaks others - format adherence and refusal behaviour are the usual casualties. Prompts often need adjusting rather than the model being worse, so I treat the prompt and the model version as one versioned unit and re-tune before judging. Then canary it on a slice of traffic behind a flag with a defined rollback trigger, watching quality signals and not just error rates. Structurally, I keep an abstraction over the provider and avoid depending on undocumented behaviour, so switching is a config change plus a suite run rather than a rewrite.

Q5: What do you think separates a strong candidate from an average one in this loop?

Raising measurement before being asked. An average answer describes an architecture and names good tools; a strong answer states how quality will be measured, what the latency and cost budgets are, and how a regression would be caught - and then makes every design choice against those. The second difference is talking about failure honestly: what broke, what was rejected and by how much, what is still unmeasured. Those two behaviours signal the same underlying thing, which is that the candidate has actually operated a probabilistic system rather than demoed one. Everything else - retrieval depth, agent design, cost tuning - is easier to teach than the instinct to measure first.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access