Lesson 24 - The Agentic AI Interview
Run it:
python3 -m unittest discover -s tests -t .Concept: The AI Engineer Interview covers the theory and the interview framing, without code.
What you will build
- A repeatable structure for the agent design round, with the eval strategy placed before model choice
- An answer shape for each recurring question, pointing at the lesson that holds the mechanism
- A triage order for the debugging round, starting from the trace rather than from the prompt
- A preparation plan that sequences this course against the design and coding rounds
The idea
There is no single βAI interviewβ. There is a set of rounds, and the mistake that costs the most is preparing for the wrong one - grinding model trivia for a round that is really about failure modes, or preparing architecture diagrams for a round that is really an ordinary coding exercise.
Round types vary by company and by team, so treat the list below as the shapes that recur rather than anyoneβs specific process. What does not vary is the thing being assessed: can this person tell whether the system they built is working? Every question below is a different angle on that, and every strong answer comes back to measurement.
The analogy: you are not being asked whether you can drive. You are being asked whether you would notice the brakes going soft.
The rounds, and what each one grades
An ordinary coding round. Arrays, strings, graphs, the usual. This has not gone away, and βI work on agentsβ does not exempt you. Graded on the same things it always was - correctness, complexity, clear reasoning under time pressure. Keep /dsa in your rotation.
An AI or agent system design round. You are given a product and asked to design the system. Graded on whether you reason about tradeoffs and failure modes, not on whether you know product names. This is the round most people under-prepare, and the one with the most structure available to you - see below.
A practical build or take-home. Wire up a small agent, usually with a scoped tool set, sometimes with a corpus. Graded on whether you handled the boring parts: validation, a budget, error paths, a test. A submission that works on the happy path and falls over on a malformed input scores badly even when the happy path is elegant.
A debugging round. You are handed a misbehaving agent and asked to diagnose it. Graded on your order of operations, not on whether you find the bug - an unstructured search that lands on the answer scores worse than a disciplined search that runs out of time.
A depth conversation about your own project. Graded on whether the numbers you quote hold up when pushed. This is where the capstone earns its place, and where a demo stops being enough.
The agent design round
flowchart LR
T[Clarify the task and who sees a failure] --> Q[Quality bar plus latency and cost budgets]
Q --> W[Workflow or agent and defend it]
W --> TS[Tool surface and authorization]
TS --> EV[Eval strategy]
EV --> M[Model choice]
M --> B[Budgets and termination]
B --> G[Guardrails and containment]
G --> O[Observability]
O --> R[Rollout and rollback]
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class T,Q,W,TS service
class EV async
class M,B,G,O,R data
Clarify the task, and who sees a failure. An agent that drafts something for an expert to approve has a completely different reliability bar from one that emails a customer directly. Ask who is on the other end before you design anything.
Establish the quality bar and the budgets. What pass rate is acceptable, what latency the product can absorb, what a run may cost. Naming numbers early - even as explicit assumptions - is what lets you defend every later choice.
Decide workflow versus agent, and defend it. If the steps are known in advance, a workflow is cheaper, faster, testable and debuggable, and choosing one is a strength rather than a lack of ambition. An agent earns its cost only when the path genuinely depends on what it finds mid-task. Saying βthis should not be an agentβ to an interviewer asking for an agent, with a reason, is one of the strongest moves available.
Design the tool surface and its authorization. Which tools, what each returns, which are mutating, and who is allowed to call them. Authorization belongs to the callerβs identity, never to the modelβs choice - if the model picks a record id, you still check the user may read it.
Then specify the eval strategy, before you talk about the model. This is the single highest-signal move in the round. Say what a golden set looks like for this product, how many cases, how they are tagged, which assertions are deterministic, where a judge is needed and how you would validate it. It works because it reframes the whole conversation from what will you build to how will you know it works - and everything after it becomes concrete. Model choice becomes an experiment you can run rather than an opinion. Budgets become numbers from a distribution. Guardrails become cases that either pass or fail. Almost nobody does this in this order, and the people who do stop sounding like they are describing a plan and start sounding like they have shipped one.
Model choice comes after, and is now cheap to discuss: put the seam in, try the small model on the early steps, measure. Budgets and termination next, with named stop reasons. Guardrails, framed as containment rather than prevention. Observability, with the trace as the bug report. Rollout, with a shadow phase, a canary and a flag you can flip back.
The most common failure
Naming frameworks instead of reasoning about tradeoffs. βI would use LangGraph with a supervisor and CrewAI for the specialistsβ answers nothing: it does not say what the loop does, where it stops, what the tools return, or how you would know it worked. It also invites the follow-up you cannot answer, which is what that framework does for you.
The inverse is the whole reason this course builds everything by hand. When you can describe the loop, the budget and the dispatch path in your own words, naming a framework becomes an implementation detail you mention in passing - βa graph framework gives you this plus persistence, which is worth it once the topology stops being a straight lineβ - and that reads completely differently.
Recurring questions and answer shapes
Each of these has a lesson behind it. Point at the mechanism; do not recite the lesson.
βHow do you know it works?β The question behind most other questions. Tiered evals cheapest-first: deterministic assertions on every commit, reference comparisons where a ground truth exists, a judge where it does not, humans sampled. Trajectory scored alongside outcome, because the right answer after six refund calls is not a working agent. β lesson 10
βWhen would you NOT build an agent?β When the steps are known in advance. Then a workflow is cheaper, faster and debuggable, and you are not paying a model to rediscover a fixed sequence on every request. β lesson 1
βHow do you stop it looping forever?β A budget with named terminal states: step, token and cost ceilings, plus a repeat guard on identical tool calls that fires before dispatch. run() returns rather than raising, so the caller inspects the stop reason. β lesson 8
βHow do you handle prompt injection?β By containment, not prevention. Instructions and data share one channel, so no phrasing makes a documentβs text non-instruction; a context that has read untrusted input loses its mutating tools. Then prove it with an eval case. β lesson 12
βWhy is your agent expensive?β Because input tokens compound: the loop resends the whole transcript every step, so a large tool result is charged again on every subsequent call. Measure from the trace, cut steps first, then trim what tools return. β lesson 21
βHow do you test something non-deterministic?β Put a seam at the model boundary and script it, so the loop, the trace and the suite are deterministic. Above that, assert on properties rather than exact strings, and compare deltas against a baseline rather than absolute scores. β lesson 2
βWhat does your trace capture?β Per span: model id, finish reason, input and output tokens, cost, requested tools, tool outcome, duration, nested under one run. Enough to attribute a regression and to replay the step that failed. β lesson 9
βHow do you evaluate tool selection?β On the trajectory, and on the requested-versus-executed split. used_tools and trajectory_is score what ran; never_used proves a guardrail held; never_requested shows whether the model even tried. β lesson 10
The debugging round
You get a misbehaving agent and a time limit. Say your order of operations out loud, then follow it.
- Read the trace first. Not the prompt. The prompt is where instinct sends you and it is the least informative artefact in the system, because it is the same prompt on the runs that worked.
- Check the stop-reason mix. Free, instant, and it partitions the problem:
STEP_BUDGETclimbing is control flow,REPEATED_CALLmeans a tool result stopped answering the question the model is asking,TOOL_FATALis authorization or credentials. - Split retrieval from generation. Score the retrieval step separately. If the right chunk was never retrieved, no prompt change fixes it, and every minute spent on wording is wasted.
- Read and cluster a sample of failures. Twenty failures by hand, grouped by cause rather than by case. The largest cluster is the only thing worth working on, and you cannot know which one it is without looking.
- Only then change anything. And change one thing, with the suite re-run on both sides.
Stating that order is the answer to the round. It shows you have debugged a non-deterministic system before, because it is not the order anyone reaches by instinct.
Talking about the capstone
Lead with numbers, not architecture. β42 golden cases, pass rate 71% to 86% across three changes, cost per run down 64%, p95 from 9.1s to 5.4sβ earns you the follow-up questions you want. A tour of your module layout does not.
Then be ready for three probes. Which change mattered most and why - have the per-change table. What did you reject - have the rejected change and its numbers, because this is where judgement shows. Where does it still fail - have the limitations, stated before you are asked.
When a number is an assumption, say so. βI used assumed unit prices, so treat the absolute cost as arbitrary and the ratio as realβ is a stronger sentence than a confident figure you cannot source.
Red flags interviewers listen for
- No evals. βIt seemed betterβ after a prompt change. The single loudest signal that the work was demo-shaped.
- No budget. A loop with no ceiling, or a ceiling picked as a round number rather than from a measured distribution.
- Treating the model as magic. Reasoning about what the model βunderstandsβ instead of about what the loop sends, what the tool returns, and where it stops.
- Unable to state a cost or latency figure for your own project. Not knowing the number means you never measured, which means the rest of your claims are guesses.
- Claiming injection is solved. It is contained, not solved. Confidence here reads as inexperience, because the people who have worked on it are the ones who hedge.
- Framework name-dropping without mechanism. Naming a tool is fine; being unable to say what it does for you is not.
A preparation plan
Sequence, not a syllabus. The parts of this course map onto the rounds directly.
- Parts 1 and 2 of this course - the loop, tools, memory, then failure handling, budgets, tracing, evals, guardrails. This is the agent design round and the debugging round. If you only do one thing, make it lessons 8 through 12.
- Part 3, retrieval - embeddings through agentic search. Almost every real product is retrieval-shaped, so expect it in the design round.
- Parts 4 and 5 - planning, multi-agent, durable execution, frameworks, then latency, cost and rollout. This is where senior-level follow-ups live.
- /hld for the design round - agents run on ordinary distributed systems, and the queues, caches and consistency arguments are the same ones.
- /dsa for the coding round - unchanged, still required.
- System design concepts for the fundamentals - idempotency, retry with backoff, circuit breakers, rate limiting, observability. Every one of them appears inside a tool.
- Design ChatGPT as the worked AI system design, end to end.
- The AI Engineer track for the broader vocabulary - tokenization and cost, model selection, inference serving, hallucination and grounding.
Then the capstone from lesson 23, which is what turns all of the above from things you have read into a thing you can be asked about.
Exercise
Run the agent design round on yourself, out loud, timed at 35 minutes, against a product you have not designed before - a support agent over a warranty corpus, or an on-call assistant over runbooks. Record yourself.
Success criterion: on playback, you reached the eval strategy before you named a model, and you can point to the moment you did it. If you named a model first, do it again on a different product.
What to listen for on playback
Five specific things, in order of how much they cost you: 1. **Did you name a model before specifying evals?** The most common ordering mistake. Model choice feels like the concrete decision, so it comes out first, and it turns the rest of the round into opinion because there is no instrument to settle anything with. 2. **Did you ask who sees a failure?** If you never established whether a human approves the output, every reliability claim you made afterwards was ungrounded. 3. **Did you say any numbers?** A pass-rate target, a latency budget, a cost ceiling - even as explicit assumptions. A design round with no numbers cannot be evaluated by the interviewer either. 4. **Did you name a framework before describing the loop?** Count how long you spent on tool names versus on what the loop does and where it stops. 5. **Did you consider *not* building an agent?** Even one sentence. "If these steps turn out to be fixed, this is a workflow and cheaper" is a mark of judgement, and its absence is noticed. Then the tell that matters more than any of them: **could you say how you would know the system was working?** If the answer was vague, that is not a presentation problem to polish. Go back to lesson 10 and build the suite, because the vagueness was real.Checkpoint
Why specify the eval strategy before discussing model choice?
Because it reframes the round from what you would build to how you would know it works, and it makes everything downstream concrete. Model choice becomes an experiment rather than an opinion, budgets become numbers from a distribution, and guardrails become cases that pass or fail.
What is the first thing you read in a debugging round, and why not the prompt?
The trace, then the stop-reason mix. The prompt is identical on the runs that worked, so it carries almost no information, while the stop mix partitions the problem into control flow, tool quality, or authorization in seconds.
Why is βthis should not be an agentβ a strong answer?
Because when the steps are known in advance a workflow is cheaper, faster and debuggable. Recognising that shows you understand what an agent costs rather than reaching for the most impressive-sounding design.
Which red flag is hardest to recover from in a project conversation?
Not being able to state a cost or latency figure for your own system. It establishes that you never measured, which means every other claim in the conversation is a guess.
Theory and interview framing: Become an AI Engineer