Evals - Complete Deep Dive
Prerequisites: Prompt Engineering, LLM APIs and SDKs, RAG End to End Used in: LLM-as-Judge, Hallucination and Grounding, Deploying AI Features Build it: Lesson 10 - Agent Evals - Golden Tasks and Trajectories implements this as runnable, tested code you can execute offline.
What Are Evals?
An eval is a repeatable measurement of whether your AI system did its job, run against a fixed set of inputs, producing a number you can compare across versions. Not a vibe check. Not a demo. A measurement.
Say the uncomfortable part first: without evals you are not engineering, you are vibing. Every prompt tweak becomes a coin flip you have no instrument to read. You change a line in a system prompt, try three examples, decide it looks better, and ship. What you actually did was trade an unknown set of improvements for an unknown set of regressions. βIt looked better in the three examples I triedβ is not a test result β it is the exact mechanism by which teams ship regressions into production and discover them from user complaints a week later.
This is why evals sit at the top of the skill list employers screen for. An AI engineerβs real job is making a non-deterministic component behave predictably enough to ship, and evals are the only instrument that makes that possible. Everything else on this track β prompting, retrieval, agents, fine-tuning β is a knob. Evals are the gauge. Turning knobs without a gauge is not a discipline.
Real-world analogy: Compare grading arithmetic homework to grading essays. Arithmetic has one right answer, so a script can mark it. Essays do not, so schools invented rubrics, multiple graders, and moderation meetings to make subjective grading consistent enough to be fair. LLM output is essays. Evals are the rubric, the graders, and the moderation process β the machinery that turns βI liked itβ into a score two people would agree on.
Why Unit Tests Do Not Transfer
Unit testing rests on two assumptions that LLM output violates outright.
| Unit test assumption | Reality with an LLM |
|---|---|
| There is one correct output | Many phrasings are equally correct. String equality fails a good answer for using a synonym. |
| The same input gives the same output | Sampling is stochastic. The same prompt run twice can differ, so a passing test can fail on retry with nothing changed. |
| Failures are binary | Quality is a gradient. An answer can be correct but rude, or well-written but subtly unsupported. |
| A failing test localises the bug | A bad answer could come from the retriever, the chunker, the prompt, the model, or the question being genuinely unanswerable. |
| Coverage is measurable from code paths | The input space is natural language. You cannot enumerate it, only sample it. |
The consequence is not βtests are useless here.β It is that the shape changes: from assert exact equality on every case to measure a distribution of quality over a representative sample, and alert on movement. You stop asking βdid it passβ and start asking βdid it get worse than the version I already shipped.β
Some parts of the system remain perfectly unit-testable, and you should test them normally: the chunker, the retrieval filter, the JSON parser, the tool dispatcher, the permission predicate. Do not let the non-determinism of one component excuse untested deterministic code around it.
The Eval Loop
Evals are not a phase at the end. They are a loop, and the loop is the product development process.
flowchart LR
D[Develop - change prompt or retriever or model] --> E[Offline Eval Suite]
E --> G[Gate - fail the build on a regression]
G --> S[Ship behind a flag to a slice of traffic]
S --> O[Observe - traces and online signals]
O --> M[Mine failures and cluster into categories]
M --> N[Promote real failures into the golden set]
N --> E
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class D client
class E,G service
class S edge
class O,M async
class N data
The loop closes at N. A failure seen in production that never becomes a permanent test case will happen again. Teams that ship reliable AI features are not the ones with the cleverest prompts β they are the ones whose golden set grows every week from real breakage.
The Eval Hierarchy
Four tiers, cheapest and most deterministic first. Run them at different cadences, because their costs differ by orders of magnitude.
flowchart TD
A[Tier 1 - assertions and rule checks - every commit] --> B[Tier 2 - reference based metrics - every pull request]
B --> C[Tier 3 - model graded evals - nightly and pre release]
C --> D[Tier 4 - human review - weekly sample and launch gates]
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef client fill:#f97316,stroke:#c2410c,color:#fff
class A,B service
class C async
class D client
Tier 1 - Assertions and rule-based checks
Deterministic, effectively free, and they run in CI on every commit. Most teams underbuild this tier and then complain that evals are expensive. A surprising share of real production failures are caught here.
def check(response, ctx):
assert response.status == "ok"
payload = json.loads(response.text) # must be parseable at all
Invoice.model_validate(payload) # must match the schema contract
assert payload["citations"], "no citation emitted"
assert all(c in ctx.retrieved_ids for c in payload["citations"])
assert "as an AI language model" not in response.text # denylist phrases
assert response.latency_ms < ctx.latency_budget_ms
assert response.cost_usd < ctx.cost_ceiling_usd
What belongs here: schema validity, required fields present, enum values from an allowlist, a required citation exists and points at a chunk that was actually retrieved, forbidden phrases from a denylist, refusal happened on inputs that must be refused, no PII in the output, tool calls used arguments that type-check, output length inside bounds, and latency and cost budgets treated as correctness. A response that is right but took twelve seconds and cost ten times budget failed.
Tier 2 - Reference-based metrics
For tasks where a ground truth genuinely exists β extraction, classification, routing, translation, SQL generation, retrieval. Here you can use exact match, field-level F1, precision and recall, or execution equivalence for generated queries. Cheap, objective, and the right default whenever the task can be framed this way.
Worth reframing tasks to land in this tier when you can. βSummarise this contractβ has no reference answer. βExtract the parties, the term length, and the termination clauseβ does, and that is most of what the summary was for.
Tier 3 - Model-graded evals
Another model scores the output against a rubric. This is the only practical way to score open-ended dimensions β helpfulness, tone, faithfulness to context, instruction adherence β at every-commit frequency. It is also the tier where teams fool themselves, because an unvalidated judge produces numbers that feel authoritative and mean nothing. Full treatment in LLM-as-Judge.
Tier 4 - Human review
The gold standard and the expensive one. Humans do three jobs no automated tier can: they set the ground truth the other tiers are calibrated against, they catch failure modes nobody thought to write a check for, and they make the call on genuinely ambiguous cases. Do not scale it to every commit. Scale it to a sampled slice, to launch gates, and to labelling the set that validates your judge.
The tradeoff table
| Eval type | Relative cost | What it catches | What it misses |
|---|---|---|---|
| Schema and format assertions | Negligible | Unparseable output, missing fields, broken tool arguments, contract drift | Anything about whether the content is true or useful |
| Denylist and refusal checks | Negligible | Leaked boilerplate, banned claims, failure to refuse an unsafe request | Subtle policy violations phrased in new words |
| Latency and cost budgets | Negligible | Prompt bloat, retry storms, an expensive model creeping into a cheap path | Quality of any kind |
| Reference-based metrics | Low | Extraction and classification errors, retrieval recall drops, routing mistakes | Open-ended quality where no single answer exists |
| Model-graded rubric | Moderate, scales with volume | Unfaithful claims, ignored instructions, tone and style regressions | Its own biases, and anything outside the rubric you wrote |
| Human review | High, does not scale | Novel failure modes, genuine ambiguity, whether users would actually be satisfied | Coverage - you will only ever see a sample |
| Online signals | Low once instrumented | Real user dissatisfaction, drift, populations your test set never represented | Attribution - you rarely know which change caused it |
Building a Golden Dataset
The golden set is the asset. Prompts are cheap and disposable; a well-curated eval set is the thing that would take a competitor months to rebuild.
Start from real traffic, not imagination. Invented test cases encode your assumptions about how people will use the system, and those assumptions are wrong in a specific, predictable direction: too well-formed, too polite, too on-topic. Pull cases from logged production requests and from the failures users actually reported. If you have not launched, run a small internal pilot to harvest real inputs before writing a single test case by hand.
Target tens before hundreds. A carefully curated set of 30 to 50 cases that covers your real failure modes is more useful than 500 generated ones that all probe the same easy path. Small sets are also cheap enough to run constantly, which matters more than coverage early on. Grow the set when a new failure category appears, not on a schedule.
Deliberately load it with the hard cases. An eval set drawn uniformly from traffic will be mostly easy, so your score will be high and flat and will not move when you break something. Over-sample the adversarial and the awkward: ambiguous questions, questions the corpus genuinely cannot answer, multi-part questions, questions with a false premise, prompt-injection attempts, very long inputs, non-English inputs, entities with similar names, and anything that has broken before.
Record the expected behaviour, not just the expected string. For many cases the correct outcome is βrefuses and explains whyβ or βcites document Xβ or βasks a clarifying questionβ β not a specific paragraph.
- id: refund-policy-0142
source: production trace 8f2c1a
question: "Can I still get a refund - I bought it 45 days ago"
expected_behaviour: decline_and_cite_policy
must_cite_any: [policy-refunds-v3-section-2]
must_contain_any: ["30 day", "thirty day"]
must_not_contain: ["yes you can", "approved"]
tags: [policy, adversarial, outside-window]
added_because: "model approved a refund outside the window on a live ticket"
Version it and review changes like code. The set lives in the repo, changes arrive by pull request, and every case records why it exists. An eval set you can quietly edit is an eval set that will be quietly edited to make a bad release look fine.
Keep a held-out slice. Iterate against a development split and hold a portion back, touched only at release gates. Without this you have no defence against tuning the prompt to the test.
Error Analysis - The Highest-Leverage Activity
The single highest-return hour in AI engineering is reading failures. Not looking at the aggregate score. Reading the actual bad outputs, one at a time, with the retrieved context next to them.
The process is mechanical:
- Sample failures β 30 to 50 is plenty β from the eval run and from production traces.
- Read each one and write a one-line description of what went wrong, in your own words.
- Cluster the descriptions into categories. The categories emerge from the data; do not pre-define them.
- Count each cluster and sort by size.
- Fix the largest cluster. Re-run. Re-cluster.
This beats chasing an aggregate number for a concrete reason: a score of β72β tells you nothing about what to do next, whereas β41 percent of failures are the retriever returning the wrong version of a policy documentβ tells you exactly what to build tomorrow. Aggregate scores are for detecting regressions. Error analysis is for making progress.
Typical clusters, roughly in the order teams discover them: retrieval missed the right chunk entirely; the right chunk was retrieved but ranked too low to be read; the model ignored the context and answered from pretraining; the question was genuinely unanswerable and the model bluffed instead of abstaining; output format drifted; the model over-generalised from one document to a claim it does not support; and the question was ambiguous and the model picked an interpretation without saying so.
Notice that the fixes for those clusters live in completely different parts of the system. That is the payoff β error analysis routes effort to the component that is actually broken instead of the component you most enjoy tuning.
Split Retrieval Evals From Generation Evals
For any RAG system, a single end-to-end quality score is close to useless diagnostically, because a bad answer has two structurally different causes and the score cannot tell them apart.
| Reading | Interpretation | Where to work |
|---|---|---|
| Low retrieval recall, high faithfulness | The retriever failed, the generator was honest about what it had | Chunking, hybrid search, reranking |
| High retrieval recall, low faithfulness | The evidence was in context and the model went off it | Grounding instructions, context ordering, model choice |
| Both low | Retrieval is the binding constraint - fix it first | Retrieval, and re-measure before touching the prompt |
| Both high, users still unhappy | Your eval set does not represent real usage | Golden set curation, online signals |
Score the retriever on its own terms β recall at k, precision at k, rank of the first relevant chunk β and score generation given a fixed context so generation changes are not confounded by retrieval noise. The mechanics of each scorecard are in RAG End to End.
Regression Suites in CI
An eval that runs when someone remembers to run it is not a gate. Wire it into the build.
- Tier 1 on every commit. Deterministic, fast, no model call needed for most of it. Treat a failure exactly like a failing unit test.
- Tier 2 and a sampled Tier 3 on every pull request. Report the delta against the current production version, not the absolute score. Humans cannot judge whether 0.81 is good; everyone can judge that it used to be 0.86.
- An eval drop is a build failure. This is the cultural line that separates teams who ship confidently from teams who ship hopefully. If a score regression only produces a warning, it will be ignored under deadline.
- Pin what you can. Fixed seeds where the provider supports them, temperature zero for eval runs, a pinned model identifier, a frozen retrieval index snapshot. You want the only moving part to be your change β and on genuinely noisy dimensions, run each case several times and compare distributions rather than blocking a release on one unlucky sample.
- Budget the suite. If the full run is too slow or expensive for a pull request, run a fast subset per pull request and the full suite nightly and before release.
Rolling a model or prompt change out behind a flag, with a defined rollback trigger and a canary slice, is ordinary release engineering applied to a non-deterministic component. The same discipline described in Deployment and Reliability applies, with one addition: your canary metric is an eval score, not just an error rate.
Online Signals - Closing the Loop
Offline evals measure what you thought to ask. Production measures what users actually do. You need both, and the second is what keeps the first honest.
Signals worth instrumenting, roughly in increasing order of usefulness:
- Explicit feedback β thumbs up and down. Sparse and biased toward the annoyed, but the negatives are high-quality failure leads.
- Task completion β did the user get the thing done. The closest proxy for whether the feature works at all.
- Escalation rate β how often a user abandons the AI path for a human, a search box, or a support ticket. Hard to argue with.
- Edit distance on accepted output β when users can edit before accepting, how much they change is a continuous quality signal that needs no survey. Heavy editing on a βsuccessfulβ response is a failure the thumbs never recorded.
- Regeneration and rephrasing β a user immediately retrying is telling you the first answer was useless.
- Abstention and refusal rates β zero abstentions on a retrieval system means it is bluffing. A spike means retrieval broke.
The loop closes when these signals feed back into the offline set. Every thumbs-down, escalation, and heavily-edited response is a candidate golden case. Triage them weekly, cluster them with the same error analysis process, and promote the representative ones into the versioned set with a note on why. This requires per-request traces you can actually retrieve and replay β the inputs, the retrieved context, the prompt, the output, and the model version. See Tracing and Observability for LLM Apps.
Eval Set Rot and Overfitting
Two honest failure modes of the discipline itself.
Your eval set decays. The product changes, the corpus changes, users find new uses, and a set curated six months ago stops representing current traffic. Symptoms: the score is high and never moves, or it stays flat while user complaints rise. Treatment is scheduled review β re-sample recent traffic, retire cases about features that no longer exist, and add cases for the parts of the product that did not exist when the set was written.
You will overfit to your own benchmark. Iterate against the same 50 cases for weeks and you will make those 50 cases pass, which is not the same as making the system better. This is Goodhartβs law with a progress bar. Defences: hold out a slice you only touch at release gates, rotate in fresh production cases regularly, be suspicious of a jump that came from a prompt edit rather than an architectural change, and always keep one non-eval check β a human reading a fresh sample β in the loop.
Neither of these is a reason to skip evals. They are reasons to treat the eval set as a living artifact with an owner, the same way you treat a production runbook.
Bad to Good to Great
Bad - eyeball a few outputs
Change the prompt, run three examples you remember, decide it looks better, ship.
You cannot detect regressions, because you never measured the previous version. You cannot compare two candidate prompts, because the comparison is a memory of a vibe. Your three examples are almost certainly the easy ones, since you wrote them while thinking about the happy path. And there is no artifact β the knowledge lives in your head, so the next engineer to touch the prompt starts from zero and re-learns every failure mode by shipping it.
Good - a spreadsheet of fixed test cases
Twenty or so questions with expected answers in a sheet, run manually before a release, results pasted in a column.
This is a genuine improvement and enough to run a small feature. Its limits are specific: it is manual, so it gets skipped under deadline; it is graded by whoever is looking, so the standard drifts between runs; it is not versioned with the code, so you cannot tell which prompt produced which column; it has no cost or latency dimension; and it only contains cases somebody thought of, not cases that actually broke.
Great - a versioned golden set with tiered automated checks in CI, plus tracked online metrics
- Golden set in the repo, versioned, every case tagged and annotated with why it exists, sourced from real traffic and real failures, with a held-out slice.
- Tier 1 assertions on every commit β schema, citations, denylist, refusals, latency and cost budgets.
- Tier 2 and sampled Tier 3 on every pull request, reported as a delta against production, with a regression failing the build.
- A validated judge for the open-ended dimensions, with measured agreement against human labels β see LLM-as-Judge.
- Retrieval and generation scored separately so a regression localises to a component.
- Full suite plus human review on a sample before release, as a gate rather than a formality.
- Online metrics tracked β completion, escalation, edit distance, abstention β with a weekly triage that promotes real failures into the golden set.
- The set itself reviewed on a cadence for rot, with a held-out slice protecting against overfitting.
The difference between Good and Great is not sophistication. It is that Great runs without anyone deciding to run it, and it produces a number that a reviewer who was not in the room can trust.
When to Use
β Build evals when:
- You are about to change a prompt, model, retriever, or chunking strategy and want to know what it cost you
- Anything LLM-generated is reaching users, at any volume
- More than one person can modify the prompts
- You are choosing between models and need a defensible basis beyond preference
- Quality complaints exist but nobody can say whether the system got worse
β Do not over-invest when:
- You are still exploring whether the feature is worth building β a handful of manual cases is the right size for a spike
- The task is fully deterministic and unit tests already cover it
- You have not defined what βgoodβ means for the task, in which case write that definition first
- You would be building a model-graded suite before validating the judge, which measures nothing
Common Interview Questions
Q1: How do you test something that has no single correct answer?
You stop testing for equality and start measuring a distribution of quality over a representative sample. Decompose the vague notion of quality into properties you can actually check, then match each property to the cheapest tier that can check it: deterministic assertions for schema validity, required citations, refusals, and latency and cost budgets; reference-based metrics for any part you can reframe as extraction or classification; a validated model-graded rubric for open-ended dimensions; human review on a sample for ground truth and novel failures. The unit of comparison is the delta against the version currently in production, not an absolute pass or fail.
Q2: A prompt change improved your eval score. Would you ship it?
Not on that alone. First, is the movement larger than the run-to-run noise of a sampled stochastic system, which means re-running and comparing distributions rather than single numbers. Second, did the aggregate improve because one failure cluster got fixed, or did one cluster improve while another quietly regressed - aggregate scores hide compensating movements, so I check the per-category breakdown. Third, did it hold on the held-out slice, or did I just tune to the development split. Fourth, did latency and cost stay inside budget. If all four hold, I ship behind a flag to a slice of traffic and watch the online signals, because the eval set only represents the usage I thought of.
Q3: What is the first eval you build for a brand-new LLM feature?
The deterministic tier, against roughly thirty real inputs. Schema validity, required fields, a citation that points at a chunk actually retrieved, refusal on the inputs that must be refused, and latency and cost ceilings. It is nearly free, it runs in CI on day one, and it catches a genuinely large share of real production breakage. Only after that do I spend money on a judge, because a model-graded suite sitting on top of output that does not reliably parse is measuring the wrong layer.
Q4: Your end-to-end quality score dropped after a release. How do you find out why?
Split the score before interpreting it. Retrieval metrics and generation metrics fail for different reasons, so I check whether recall at k moved, and separately re-score generation against a frozen context so retrieval noise is held constant. Low recall with honest generation points at chunking, the embedding model, or the index. Good recall with unfaithful output points at the prompt, context assembly, or the model. Then I stop looking at numbers, sample thirty failures, read them with their retrieved context, and cluster them - the largest cluster usually names the regression outright, and traces make each case replayable.
Q5: How do you keep an eval set from going stale or getting gamed?
Treat it as a living artifact with an owner. Keep it in version control so changes arrive by pull request and nobody can quietly delete an inconvenient case. Hold out a slice that is only run at release gates, so week-to-week iteration cannot tune against it. Feed it continuously from production - thumbs-down, escalations, heavily-edited accepted output - so it tracks how the product is actually used rather than how it was used at launch. Review it on a cadence to retire cases for removed features. And watch for the tell: a score that is high and never moves means the set stopped being hard, not that the system got good.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts