Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 23 min read

Evals - Complete Deep Dive

Prerequisites: Prompt Engineering, LLM APIs and SDKs, RAG End to End Used in: LLM-as-Judge, Hallucination and Grounding, Deploying AI Features Build it: Lesson 10 - Agent Evals - Golden Tasks and Trajectories implements this as runnable, tested code you can execute offline.


What Are Evals?

An eval is a repeatable measurement of whether your AI system did its job, run against a fixed set of inputs, producing a number you can compare across versions. Not a vibe check. Not a demo. A measurement.

Say the uncomfortable part first: without evals you are not engineering, you are vibing. Every prompt tweak becomes a coin flip you have no instrument to read. You change a line in a system prompt, try three examples, decide it looks better, and ship. What you actually did was trade an unknown set of improvements for an unknown set of regressions. β€œIt looked better in the three examples I tried” is not a test result β€” it is the exact mechanism by which teams ship regressions into production and discover them from user complaints a week later.

This is why evals sit at the top of the skill list employers screen for. An AI engineer’s real job is making a non-deterministic component behave predictably enough to ship, and evals are the only instrument that makes that possible. Everything else on this track β€” prompting, retrieval, agents, fine-tuning β€” is a knob. Evals are the gauge. Turning knobs without a gauge is not a discipline.

Real-world analogy: Compare grading arithmetic homework to grading essays. Arithmetic has one right answer, so a script can mark it. Essays do not, so schools invented rubrics, multiple graders, and moderation meetings to make subjective grading consistent enough to be fair. LLM output is essays. Evals are the rubric, the graders, and the moderation process β€” the machinery that turns β€œI liked it” into a score two people would agree on.


Why Unit Tests Do Not Transfer

Unit testing rests on two assumptions that LLM output violates outright.

Unit test assumption Reality with an LLM
There is one correct output Many phrasings are equally correct. String equality fails a good answer for using a synonym.
The same input gives the same output Sampling is stochastic. The same prompt run twice can differ, so a passing test can fail on retry with nothing changed.
Failures are binary Quality is a gradient. An answer can be correct but rude, or well-written but subtly unsupported.
A failing test localises the bug A bad answer could come from the retriever, the chunker, the prompt, the model, or the question being genuinely unanswerable.
Coverage is measurable from code paths The input space is natural language. You cannot enumerate it, only sample it.

The consequence is not β€œtests are useless here.” It is that the shape changes: from assert exact equality on every case to measure a distribution of quality over a representative sample, and alert on movement. You stop asking β€œdid it pass” and start asking β€œdid it get worse than the version I already shipped.”

Some parts of the system remain perfectly unit-testable, and you should test them normally: the chunker, the retrieval filter, the JSON parser, the tool dispatcher, the permission predicate. Do not let the non-determinism of one component excuse untested deterministic code around it.


The Eval Loop

Evals are not a phase at the end. They are a loop, and the loop is the product development process.

flowchart LR
    D[Develop - change prompt or retriever or model] --> E[Offline Eval Suite]
    E --> G[Gate - fail the build on a regression]
    G --> S[Ship behind a flag to a slice of traffic]
    S --> O[Observe - traces and online signals]
    O --> M[Mine failures and cluster into categories]
    M --> N[Promote real failures into the golden set]
    N --> E

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class D client
    class E,G service
    class S edge
    class O,M async
    class N data

The loop closes at N. A failure seen in production that never becomes a permanent test case will happen again. Teams that ship reliable AI features are not the ones with the cleverest prompts β€” they are the ones whose golden set grows every week from real breakage.


The Eval Hierarchy

Four tiers, cheapest and most deterministic first. Run them at different cadences, because their costs differ by orders of magnitude.

flowchart TD
    A[Tier 1 - assertions and rule checks - every commit] --> B[Tier 2 - reference based metrics - every pull request]
    B --> C[Tier 3 - model graded evals - nightly and pre release]
    C --> D[Tier 4 - human review - weekly sample and launch gates]

    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef client fill:#f97316,stroke:#c2410c,color:#fff

    class A,B service
    class C async
    class D client

Tier 1 - Assertions and rule-based checks

Deterministic, effectively free, and they run in CI on every commit. Most teams underbuild this tier and then complain that evals are expensive. A surprising share of real production failures are caught here.

def check(response, ctx):
    assert response.status == "ok"
    payload = json.loads(response.text)              # must be parseable at all
    Invoice.model_validate(payload)                  # must match the schema contract
    assert payload["citations"], "no citation emitted"
    assert all(c in ctx.retrieved_ids for c in payload["citations"])
    assert "as an AI language model" not in response.text   # denylist phrases
    assert response.latency_ms < ctx.latency_budget_ms
    assert response.cost_usd < ctx.cost_ceiling_usd

What belongs here: schema validity, required fields present, enum values from an allowlist, a required citation exists and points at a chunk that was actually retrieved, forbidden phrases from a denylist, refusal happened on inputs that must be refused, no PII in the output, tool calls used arguments that type-check, output length inside bounds, and latency and cost budgets treated as correctness. A response that is right but took twelve seconds and cost ten times budget failed.

Tier 2 - Reference-based metrics

For tasks where a ground truth genuinely exists β€” extraction, classification, routing, translation, SQL generation, retrieval. Here you can use exact match, field-level F1, precision and recall, or execution equivalence for generated queries. Cheap, objective, and the right default whenever the task can be framed this way.

Worth reframing tasks to land in this tier when you can. β€œSummarise this contract” has no reference answer. β€œExtract the parties, the term length, and the termination clause” does, and that is most of what the summary was for.

Tier 3 - Model-graded evals

Another model scores the output against a rubric. This is the only practical way to score open-ended dimensions β€” helpfulness, tone, faithfulness to context, instruction adherence β€” at every-commit frequency. It is also the tier where teams fool themselves, because an unvalidated judge produces numbers that feel authoritative and mean nothing. Full treatment in LLM-as-Judge.

Tier 4 - Human review

The gold standard and the expensive one. Humans do three jobs no automated tier can: they set the ground truth the other tiers are calibrated against, they catch failure modes nobody thought to write a check for, and they make the call on genuinely ambiguous cases. Do not scale it to every commit. Scale it to a sampled slice, to launch gates, and to labelling the set that validates your judge.

The tradeoff table

Eval type Relative cost What it catches What it misses
Schema and format assertions Negligible Unparseable output, missing fields, broken tool arguments, contract drift Anything about whether the content is true or useful
Denylist and refusal checks Negligible Leaked boilerplate, banned claims, failure to refuse an unsafe request Subtle policy violations phrased in new words
Latency and cost budgets Negligible Prompt bloat, retry storms, an expensive model creeping into a cheap path Quality of any kind
Reference-based metrics Low Extraction and classification errors, retrieval recall drops, routing mistakes Open-ended quality where no single answer exists
Model-graded rubric Moderate, scales with volume Unfaithful claims, ignored instructions, tone and style regressions Its own biases, and anything outside the rubric you wrote
Human review High, does not scale Novel failure modes, genuine ambiguity, whether users would actually be satisfied Coverage - you will only ever see a sample
Online signals Low once instrumented Real user dissatisfaction, drift, populations your test set never represented Attribution - you rarely know which change caused it

Building a Golden Dataset

The golden set is the asset. Prompts are cheap and disposable; a well-curated eval set is the thing that would take a competitor months to rebuild.

Start from real traffic, not imagination. Invented test cases encode your assumptions about how people will use the system, and those assumptions are wrong in a specific, predictable direction: too well-formed, too polite, too on-topic. Pull cases from logged production requests and from the failures users actually reported. If you have not launched, run a small internal pilot to harvest real inputs before writing a single test case by hand.

Target tens before hundreds. A carefully curated set of 30 to 50 cases that covers your real failure modes is more useful than 500 generated ones that all probe the same easy path. Small sets are also cheap enough to run constantly, which matters more than coverage early on. Grow the set when a new failure category appears, not on a schedule.

Deliberately load it with the hard cases. An eval set drawn uniformly from traffic will be mostly easy, so your score will be high and flat and will not move when you break something. Over-sample the adversarial and the awkward: ambiguous questions, questions the corpus genuinely cannot answer, multi-part questions, questions with a false premise, prompt-injection attempts, very long inputs, non-English inputs, entities with similar names, and anything that has broken before.

Record the expected behaviour, not just the expected string. For many cases the correct outcome is β€œrefuses and explains why” or β€œcites document X” or β€œasks a clarifying question” β€” not a specific paragraph.

- id: refund-policy-0142
  source: production trace 8f2c1a
  question: "Can I still get a refund - I bought it 45 days ago"
  expected_behaviour: decline_and_cite_policy
  must_cite_any: [policy-refunds-v3-section-2]
  must_contain_any: ["30 day", "thirty day"]
  must_not_contain: ["yes you can", "approved"]
  tags: [policy, adversarial, outside-window]
  added_because: "model approved a refund outside the window on a live ticket"

Version it and review changes like code. The set lives in the repo, changes arrive by pull request, and every case records why it exists. An eval set you can quietly edit is an eval set that will be quietly edited to make a bad release look fine.

Keep a held-out slice. Iterate against a development split and hold a portion back, touched only at release gates. Without this you have no defence against tuning the prompt to the test.


Error Analysis - The Highest-Leverage Activity

The single highest-return hour in AI engineering is reading failures. Not looking at the aggregate score. Reading the actual bad outputs, one at a time, with the retrieved context next to them.

The process is mechanical:

  1. Sample failures β€” 30 to 50 is plenty β€” from the eval run and from production traces.
  2. Read each one and write a one-line description of what went wrong, in your own words.
  3. Cluster the descriptions into categories. The categories emerge from the data; do not pre-define them.
  4. Count each cluster and sort by size.
  5. Fix the largest cluster. Re-run. Re-cluster.

This beats chasing an aggregate number for a concrete reason: a score of β€œ72” tells you nothing about what to do next, whereas β€œ41 percent of failures are the retriever returning the wrong version of a policy document” tells you exactly what to build tomorrow. Aggregate scores are for detecting regressions. Error analysis is for making progress.

Typical clusters, roughly in the order teams discover them: retrieval missed the right chunk entirely; the right chunk was retrieved but ranked too low to be read; the model ignored the context and answered from pretraining; the question was genuinely unanswerable and the model bluffed instead of abstaining; output format drifted; the model over-generalised from one document to a claim it does not support; and the question was ambiguous and the model picked an interpretation without saying so.

Notice that the fixes for those clusters live in completely different parts of the system. That is the payoff β€” error analysis routes effort to the component that is actually broken instead of the component you most enjoy tuning.


Split Retrieval Evals From Generation Evals

For any RAG system, a single end-to-end quality score is close to useless diagnostically, because a bad answer has two structurally different causes and the score cannot tell them apart.

Reading Interpretation Where to work
Low retrieval recall, high faithfulness The retriever failed, the generator was honest about what it had Chunking, hybrid search, reranking
High retrieval recall, low faithfulness The evidence was in context and the model went off it Grounding instructions, context ordering, model choice
Both low Retrieval is the binding constraint - fix it first Retrieval, and re-measure before touching the prompt
Both high, users still unhappy Your eval set does not represent real usage Golden set curation, online signals

Score the retriever on its own terms β€” recall at k, precision at k, rank of the first relevant chunk β€” and score generation given a fixed context so generation changes are not confounded by retrieval noise. The mechanics of each scorecard are in RAG End to End.


Regression Suites in CI

An eval that runs when someone remembers to run it is not a gate. Wire it into the build.

Rolling a model or prompt change out behind a flag, with a defined rollback trigger and a canary slice, is ordinary release engineering applied to a non-deterministic component. The same discipline described in Deployment and Reliability applies, with one addition: your canary metric is an eval score, not just an error rate.


Online Signals - Closing the Loop

Offline evals measure what you thought to ask. Production measures what users actually do. You need both, and the second is what keeps the first honest.

Signals worth instrumenting, roughly in increasing order of usefulness:

The loop closes when these signals feed back into the offline set. Every thumbs-down, escalation, and heavily-edited response is a candidate golden case. Triage them weekly, cluster them with the same error analysis process, and promote the representative ones into the versioned set with a note on why. This requires per-request traces you can actually retrieve and replay β€” the inputs, the retrieved context, the prompt, the output, and the model version. See Tracing and Observability for LLM Apps.


Eval Set Rot and Overfitting

Two honest failure modes of the discipline itself.

Your eval set decays. The product changes, the corpus changes, users find new uses, and a set curated six months ago stops representing current traffic. Symptoms: the score is high and never moves, or it stays flat while user complaints rise. Treatment is scheduled review β€” re-sample recent traffic, retire cases about features that no longer exist, and add cases for the parts of the product that did not exist when the set was written.

You will overfit to your own benchmark. Iterate against the same 50 cases for weeks and you will make those 50 cases pass, which is not the same as making the system better. This is Goodhart’s law with a progress bar. Defences: hold out a slice you only touch at release gates, rotate in fresh production cases regularly, be suspicious of a jump that came from a prompt edit rather than an architectural change, and always keep one non-eval check β€” a human reading a fresh sample β€” in the loop.

Neither of these is a reason to skip evals. They are reasons to treat the eval set as a living artifact with an owner, the same way you treat a production runbook.


Bad to Good to Great

Bad - eyeball a few outputs

Change the prompt, run three examples you remember, decide it looks better, ship.

You cannot detect regressions, because you never measured the previous version. You cannot compare two candidate prompts, because the comparison is a memory of a vibe. Your three examples are almost certainly the easy ones, since you wrote them while thinking about the happy path. And there is no artifact β€” the knowledge lives in your head, so the next engineer to touch the prompt starts from zero and re-learns every failure mode by shipping it.

Good - a spreadsheet of fixed test cases

Twenty or so questions with expected answers in a sheet, run manually before a release, results pasted in a column.

This is a genuine improvement and enough to run a small feature. Its limits are specific: it is manual, so it gets skipped under deadline; it is graded by whoever is looking, so the standard drifts between runs; it is not versioned with the code, so you cannot tell which prompt produced which column; it has no cost or latency dimension; and it only contains cases somebody thought of, not cases that actually broke.

Great - a versioned golden set with tiered automated checks in CI, plus tracked online metrics

  1. Golden set in the repo, versioned, every case tagged and annotated with why it exists, sourced from real traffic and real failures, with a held-out slice.
  2. Tier 1 assertions on every commit β€” schema, citations, denylist, refusals, latency and cost budgets.
  3. Tier 2 and sampled Tier 3 on every pull request, reported as a delta against production, with a regression failing the build.
  4. A validated judge for the open-ended dimensions, with measured agreement against human labels β€” see LLM-as-Judge.
  5. Retrieval and generation scored separately so a regression localises to a component.
  6. Full suite plus human review on a sample before release, as a gate rather than a formality.
  7. Online metrics tracked β€” completion, escalation, edit distance, abstention β€” with a weekly triage that promotes real failures into the golden set.
  8. The set itself reviewed on a cadence for rot, with a held-out slice protecting against overfitting.

The difference between Good and Great is not sophistication. It is that Great runs without anyone deciding to run it, and it produces a number that a reviewer who was not in the room can trust.


When to Use

βœ… Build evals when:

❌ Do not over-invest when:


Common Interview Questions

Q1: How do you test something that has no single correct answer?

You stop testing for equality and start measuring a distribution of quality over a representative sample. Decompose the vague notion of quality into properties you can actually check, then match each property to the cheapest tier that can check it: deterministic assertions for schema validity, required citations, refusals, and latency and cost budgets; reference-based metrics for any part you can reframe as extraction or classification; a validated model-graded rubric for open-ended dimensions; human review on a sample for ground truth and novel failures. The unit of comparison is the delta against the version currently in production, not an absolute pass or fail.

Q2: A prompt change improved your eval score. Would you ship it?

Not on that alone. First, is the movement larger than the run-to-run noise of a sampled stochastic system, which means re-running and comparing distributions rather than single numbers. Second, did the aggregate improve because one failure cluster got fixed, or did one cluster improve while another quietly regressed - aggregate scores hide compensating movements, so I check the per-category breakdown. Third, did it hold on the held-out slice, or did I just tune to the development split. Fourth, did latency and cost stay inside budget. If all four hold, I ship behind a flag to a slice of traffic and watch the online signals, because the eval set only represents the usage I thought of.

Q3: What is the first eval you build for a brand-new LLM feature?

The deterministic tier, against roughly thirty real inputs. Schema validity, required fields, a citation that points at a chunk actually retrieved, refusal on the inputs that must be refused, and latency and cost ceilings. It is nearly free, it runs in CI on day one, and it catches a genuinely large share of real production breakage. Only after that do I spend money on a judge, because a model-graded suite sitting on top of output that does not reliably parse is measuring the wrong layer.

Q4: Your end-to-end quality score dropped after a release. How do you find out why?

Split the score before interpreting it. Retrieval metrics and generation metrics fail for different reasons, so I check whether recall at k moved, and separately re-score generation against a frozen context so retrieval noise is held constant. Low recall with honest generation points at chunking, the embedding model, or the index. Good recall with unfaithful output points at the prompt, context assembly, or the model. Then I stop looking at numbers, sample thirty failures, read them with their retrieved context, and cluster them - the largest cluster usually names the regression outright, and traces make each case replayable.

Q5: How do you keep an eval set from going stale or getting gamed?

Treat it as a living artifact with an owner. Keep it in version control so changes arrive by pull request and nobody can quietly delete an inconvenient case. Hold out a slice that is only run at release gates, so week-to-week iteration cannot tune against it. Feed it continuously from production - thumbs-down, escalations, heavily-edited accepted output - so it tracks how the product is actually used rather than how it was used at launch. Review it on a cadence to retire cases for removed features. And watch for the tell: a score that is high and never moves means the set stopped being hard, not that the system got good.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access