Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 14 min read

Lesson 24 - The Agentic AI Interview

Part 6 - Ship Something Lesson 24 of 24

Run it: python3 -m unittest discover -s tests -t . Concept: The AI Engineer Interview covers the theory and the interview framing, without code.


What you will build


The idea

There is no single β€œAI interview”. There is a set of rounds, and the mistake that costs the most is preparing for the wrong one - grinding model trivia for a round that is really about failure modes, or preparing architecture diagrams for a round that is really an ordinary coding exercise.

Round types vary by company and by team, so treat the list below as the shapes that recur rather than anyone’s specific process. What does not vary is the thing being assessed: can this person tell whether the system they built is working? Every question below is a different angle on that, and every strong answer comes back to measurement.

The analogy: you are not being asked whether you can drive. You are being asked whether you would notice the brakes going soft.


The rounds, and what each one grades

An ordinary coding round. Arrays, strings, graphs, the usual. This has not gone away, and β€œI work on agents” does not exempt you. Graded on the same things it always was - correctness, complexity, clear reasoning under time pressure. Keep /dsa in your rotation.

An AI or agent system design round. You are given a product and asked to design the system. Graded on whether you reason about tradeoffs and failure modes, not on whether you know product names. This is the round most people under-prepare, and the one with the most structure available to you - see below.

A practical build or take-home. Wire up a small agent, usually with a scoped tool set, sometimes with a corpus. Graded on whether you handled the boring parts: validation, a budget, error paths, a test. A submission that works on the happy path and falls over on a malformed input scores badly even when the happy path is elegant.

A debugging round. You are handed a misbehaving agent and asked to diagnose it. Graded on your order of operations, not on whether you find the bug - an unstructured search that lands on the answer scores worse than a disciplined search that runs out of time.

A depth conversation about your own project. Graded on whether the numbers you quote hold up when pushed. This is where the capstone earns its place, and where a demo stops being enough.


The agent design round

flowchart LR
    T[Clarify the task and who sees a failure] --> Q[Quality bar plus latency and cost budgets]
    Q --> W[Workflow or agent and defend it]
    W --> TS[Tool surface and authorization]
    TS --> EV[Eval strategy]
    EV --> M[Model choice]
    M --> B[Budgets and termination]
    B --> G[Guardrails and containment]
    G --> O[Observability]
    O --> R[Rollout and rollback]
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
    class T,Q,W,TS service
    class EV async
    class M,B,G,O,R data

Clarify the task, and who sees a failure. An agent that drafts something for an expert to approve has a completely different reliability bar from one that emails a customer directly. Ask who is on the other end before you design anything.

Establish the quality bar and the budgets. What pass rate is acceptable, what latency the product can absorb, what a run may cost. Naming numbers early - even as explicit assumptions - is what lets you defend every later choice.

Decide workflow versus agent, and defend it. If the steps are known in advance, a workflow is cheaper, faster, testable and debuggable, and choosing one is a strength rather than a lack of ambition. An agent earns its cost only when the path genuinely depends on what it finds mid-task. Saying β€œthis should not be an agent” to an interviewer asking for an agent, with a reason, is one of the strongest moves available.

Design the tool surface and its authorization. Which tools, what each returns, which are mutating, and who is allowed to call them. Authorization belongs to the caller’s identity, never to the model’s choice - if the model picks a record id, you still check the user may read it.

Then specify the eval strategy, before you talk about the model. This is the single highest-signal move in the round. Say what a golden set looks like for this product, how many cases, how they are tagged, which assertions are deterministic, where a judge is needed and how you would validate it. It works because it reframes the whole conversation from what will you build to how will you know it works - and everything after it becomes concrete. Model choice becomes an experiment you can run rather than an opinion. Budgets become numbers from a distribution. Guardrails become cases that either pass or fail. Almost nobody does this in this order, and the people who do stop sounding like they are describing a plan and start sounding like they have shipped one.

Model choice comes after, and is now cheap to discuss: put the seam in, try the small model on the early steps, measure. Budgets and termination next, with named stop reasons. Guardrails, framed as containment rather than prevention. Observability, with the trace as the bug report. Rollout, with a shadow phase, a canary and a flag you can flip back.

The most common failure

Naming frameworks instead of reasoning about tradeoffs. β€œI would use LangGraph with a supervisor and CrewAI for the specialists” answers nothing: it does not say what the loop does, where it stops, what the tools return, or how you would know it worked. It also invites the follow-up you cannot answer, which is what that framework does for you.

The inverse is the whole reason this course builds everything by hand. When you can describe the loop, the budget and the dispatch path in your own words, naming a framework becomes an implementation detail you mention in passing - β€œa graph framework gives you this plus persistence, which is worth it once the topology stops being a straight line” - and that reads completely differently.


Recurring questions and answer shapes

Each of these has a lesson behind it. Point at the mechanism; do not recite the lesson.

β€œHow do you know it works?” The question behind most other questions. Tiered evals cheapest-first: deterministic assertions on every commit, reference comparisons where a ground truth exists, a judge where it does not, humans sampled. Trajectory scored alongside outcome, because the right answer after six refund calls is not a working agent. β†’ lesson 10

β€œWhen would you NOT build an agent?” When the steps are known in advance. Then a workflow is cheaper, faster and debuggable, and you are not paying a model to rediscover a fixed sequence on every request. β†’ lesson 1

β€œHow do you stop it looping forever?” A budget with named terminal states: step, token and cost ceilings, plus a repeat guard on identical tool calls that fires before dispatch. run() returns rather than raising, so the caller inspects the stop reason. β†’ lesson 8

β€œHow do you handle prompt injection?” By containment, not prevention. Instructions and data share one channel, so no phrasing makes a document’s text non-instruction; a context that has read untrusted input loses its mutating tools. Then prove it with an eval case. β†’ lesson 12

β€œWhy is your agent expensive?” Because input tokens compound: the loop resends the whole transcript every step, so a large tool result is charged again on every subsequent call. Measure from the trace, cut steps first, then trim what tools return. β†’ lesson 21

β€œHow do you test something non-deterministic?” Put a seam at the model boundary and script it, so the loop, the trace and the suite are deterministic. Above that, assert on properties rather than exact strings, and compare deltas against a baseline rather than absolute scores. β†’ lesson 2

β€œWhat does your trace capture?” Per span: model id, finish reason, input and output tokens, cost, requested tools, tool outcome, duration, nested under one run. Enough to attribute a regression and to replay the step that failed. β†’ lesson 9

β€œHow do you evaluate tool selection?” On the trajectory, and on the requested-versus-executed split. used_tools and trajectory_is score what ran; never_used proves a guardrail held; never_requested shows whether the model even tried. β†’ lesson 10


The debugging round

You get a misbehaving agent and a time limit. Say your order of operations out loud, then follow it.

  1. Read the trace first. Not the prompt. The prompt is where instinct sends you and it is the least informative artefact in the system, because it is the same prompt on the runs that worked.
  2. Check the stop-reason mix. Free, instant, and it partitions the problem: STEP_BUDGET climbing is control flow, REPEATED_CALL means a tool result stopped answering the question the model is asking, TOOL_FATAL is authorization or credentials.
  3. Split retrieval from generation. Score the retrieval step separately. If the right chunk was never retrieved, no prompt change fixes it, and every minute spent on wording is wasted.
  4. Read and cluster a sample of failures. Twenty failures by hand, grouped by cause rather than by case. The largest cluster is the only thing worth working on, and you cannot know which one it is without looking.
  5. Only then change anything. And change one thing, with the suite re-run on both sides.

Stating that order is the answer to the round. It shows you have debugged a non-deterministic system before, because it is not the order anyone reaches by instinct.


Talking about the capstone

Lead with numbers, not architecture. β€œ42 golden cases, pass rate 71% to 86% across three changes, cost per run down 64%, p95 from 9.1s to 5.4s” earns you the follow-up questions you want. A tour of your module layout does not.

Then be ready for three probes. Which change mattered most and why - have the per-change table. What did you reject - have the rejected change and its numbers, because this is where judgement shows. Where does it still fail - have the limitations, stated before you are asked.

When a number is an assumption, say so. β€œI used assumed unit prices, so treat the absolute cost as arbitrary and the ratio as real” is a stronger sentence than a confident figure you cannot source.


Red flags interviewers listen for


A preparation plan

Sequence, not a syllabus. The parts of this course map onto the rounds directly.

  1. Parts 1 and 2 of this course - the loop, tools, memory, then failure handling, budgets, tracing, evals, guardrails. This is the agent design round and the debugging round. If you only do one thing, make it lessons 8 through 12.
  2. Part 3, retrieval - embeddings through agentic search. Almost every real product is retrieval-shaped, so expect it in the design round.
  3. Parts 4 and 5 - planning, multi-agent, durable execution, frameworks, then latency, cost and rollout. This is where senior-level follow-ups live.
  4. /hld for the design round - agents run on ordinary distributed systems, and the queues, caches and consistency arguments are the same ones.
  5. /dsa for the coding round - unchanged, still required.
  6. System design concepts for the fundamentals - idempotency, retry with backoff, circuit breakers, rate limiting, observability. Every one of them appears inside a tool.
  7. Design ChatGPT as the worked AI system design, end to end.
  8. The AI Engineer track for the broader vocabulary - tokenization and cost, model selection, inference serving, hallucination and grounding.

Then the capstone from lesson 23, which is what turns all of the above from things you have read into a thing you can be asked about.


Exercise

Run the agent design round on yourself, out loud, timed at 35 minutes, against a product you have not designed before - a support agent over a warranty corpus, or an on-call assistant over runbooks. Record yourself.

Success criterion: on playback, you reached the eval strategy before you named a model, and you can point to the moment you did it. If you named a model first, do it again on a different product.

What to listen for on playback Five specific things, in order of how much they cost you: 1. **Did you name a model before specifying evals?** The most common ordering mistake. Model choice feels like the concrete decision, so it comes out first, and it turns the rest of the round into opinion because there is no instrument to settle anything with. 2. **Did you ask who sees a failure?** If you never established whether a human approves the output, every reliability claim you made afterwards was ungrounded. 3. **Did you say any numbers?** A pass-rate target, a latency budget, a cost ceiling - even as explicit assumptions. A design round with no numbers cannot be evaluated by the interviewer either. 4. **Did you name a framework before describing the loop?** Count how long you spent on tool names versus on what the loop does and where it stops. 5. **Did you consider *not* building an agent?** Even one sentence. "If these steps turn out to be fixed, this is a workflow and cheaper" is a mark of judgement, and its absence is noticed. Then the tell that matters more than any of them: **could you say how you would know the system was working?** If the answer was vague, that is not a presentation problem to polish. Go back to lesson 10 and build the suite, because the vagueness was real.

Checkpoint

Why specify the eval strategy before discussing model choice?

Because it reframes the round from what you would build to how you would know it works, and it makes everything downstream concrete. Model choice becomes an experiment rather than an opinion, budgets become numbers from a distribution, and guardrails become cases that pass or fail.

What is the first thing you read in a debugging round, and why not the prompt?

The trace, then the stop-reason mix. The prompt is identical on the runs that worked, so it carries almost no information, while the stop mix partitions the problem into control flow, tool quality, or authorization in seconds.

Why is β€œthis should not be an agent” a strong answer?

Because when the steps are known in advance a workflow is cheaper, faster and debuggable. Recognising that shows you understand what an agent costs rather than reaching for the most impressive-sounding design.

Which red flag is hardest to recover from in a project conversation?

Not being able to state a cost or latency figure for your own system. It establishes that you never measured, which means every other claim in the conversation is a guess.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access