LLM-as-Judge - Complete Deep Dive
Prerequisites: Evals, Prompt Engineering, Structured Outputs Used in: Hallucination and Grounding, RAG End to End, The AI Engineer Interview Build it: Lesson 11 - LLM-as-Judge - Rubrics and Kappa implements this as runnable, tested code you can execute offline.
What is LLM-as-Judge?
LLM-as-judge is using a language model to grade the output of another language model against a written rubric. You hand the judge the input, the output, sometimes a reference answer or the retrieved context, and a set of criteria, and it returns a verdict you can aggregate into a score.
It exists because of a gap in the eval hierarchy. Deterministic assertions handle schema, citations, refusals, and budgets. Reference-based metrics handle anything you can frame as extraction or classification. Neither can tell you whether an answer was helpful, whether it stayed faithful to the provided context, or whether it followed the instruction it was given β and those are exactly the dimensions users care about. Humans can judge them, but human review cannot run on every commit.
Real-world analogy: A newly hired grader on a teaching team. Fast, tireless, cheap, and able to mark a thousand essays overnight. Also unproven. No sensible department hands a new grader the final exams on day one. They give them a marking scheme, have them grade a batch that experienced markers have already graded, compare the two, and only let them mark independently once their marks line up. A judge you have not put through that process is a grader whose standards nobody has ever checked.
That comparison step is the entire subject of this page, because it is the step almost everyone skips.
Why Model-Graded Evaluation Is Necessary
Three arguments, in order of force.
Most interesting quality dimensions have no reference answer. βIs this summary useful to a support agentβ cannot be expressed as string equality against a gold answer, because ten good summaries are all different. Writing a gold answer for every case is itself the expensive human work you were trying to avoid.
Human review does not scale to the cadence of development. You may ship several prompt or retrieval changes a week. Each needs a before-and-after read on the same case set. Human graders cannot turn that around per pull request, so in practice the choice is not βjudge versus humansβ β it is βjudge versus no measurement at all on these dimensions.β
Judges scale to dimensions you would never staff. Tone consistency, instruction adherence, whether the answer addressed all parts of a multi-part question, whether it abstained when it should have. Each is a rubric away from being measured continuously.
Why It Is Dangerous Unvalidated
A judge produces a number with a decimal point, and numbers with decimal points are socially persuasive. That is the trap. An unvalidated judge is an unmeasured instrument: it emits confident scores whose relationship to the quality you care about is unknown. You have not solved your trust problem, you have moved it down one layer β from βdo I trust this outputβ to βdo I trust this grader,β and the second question is now invisible because it looks like data.
The concrete ways this hurts:
- The judge is systematically lenient on your worst failure mode, so your score climbs while users get worse answers.
- It scores length rather than substance, so your team learns to write padded prompts.
- It scores its own familyβs output higher, so your model selection is quietly rigged.
- The score sits in a dashboard for months, and every decision made against it inherits an error nobody measured.
The Biases a Judge Exhibits
These are structural, well-documented behaviours of the technique β not bugs in a particular model. Design around them.
| Bias | How it shows up | Mitigation |
|---|---|---|
| Position bias | In a pairwise comparison, whichever candidate is presented first wins more often than chance. Present the same text as both candidates and the judge still picks a winner. | Run both orderings and average. Treat order-inconsistent pairs as ties rather than signal. |
| Verbosity bias | Longer, more hedged, more heavily bulleted answers score higher whether or not they are more correct. Padding reads as thoroughness. | Add a rubric criterion that explicitly penalises unsupported padding. Compare candidates of similar length. Track answer length alongside the score so you can see the correlation. |
| Self-preference | Output that resembles the judgeβs own style scores higher. A judge grading a generator from its own family is not a neutral referee. | Use a different model family as judge. When comparing two generators, check the result holds with a second judge from a third family. |
| Score clustering | On a 1-to-10 scale almost everything lands on 7 or 8, so the metric has no resolution and real regressions hide inside the clump. | Binary verdicts, or a 3-point scale at most. Pairwise comparison instead of absolute scoring. Anchor every point with an example. |
| Sycophancy to the prompt | Hint in the prompt that the answer is the improved version, and the judge finds reasons to agree. | Neutral wording. Never reveal which candidate is the new version, the baseline, or yours. |
| Formatting bias | Markdown headings, bold text, and a confident tone are read as quality signals independent of content. | Tie every rubric criterion to an observable content property. Normalise formatting before judging where the task allows. |
| Leniency drift | The rubric did not change but scores moved, because the judge model version changed underneath you. | Pin the judge model identifier. Re-audit agreement on a fixed labelled set after any judge change, and treat a judge upgrade as a migration, not a patch. |
Two are worth dwelling on. Position bias is measurable in a single afternoon β feed the judge identical pairs and count how often it declares a winner. If it does that at all, your pairwise numbers need order-swapping before anyone acts on them. And verbosity bias interacts badly with optimisation pressure: if your team is iterating to raise a judge score that secretly rewards length, you will converge on longer answers and call it progress.
Techniques That Make a Judge Usable
1. A concrete rubric, not βrate the qualityβ
βRate this answer from 1 to 10 on qualityβ produces a number that means whatever the judge happened to weight that run. Replace it with criteria a careful human could check by pointing at the text. Observable beats evaluative: βadds no dates or figures absent from the contextβ is checkable; βis accurateβ is an opinion.
2. Binary or small discrete scales
Ask a question with a defensible boundary. Faithfulness is naturally binary β either every claim traces to the context or one does not. Where you need gradation, three points with named anchors beats ten points without them. Fine-grained scales feel more informative and deliver less, because the judge cannot reliably distinguish a 6 from a 7 and neither can your reviewers.
3. Pairwise comparison instead of absolute scoring
βWhich of these two answers better satisfies the criteriaβ is a far easier question than βscore this answer,β and it maps directly onto the decision you actually face: is the new version better than what is in production. Absolute scores drift between judge versions; relative preferences are more stable. The cost is that you need a baseline to compare against β keep the current production output for every golden case.
4. Require a written justification before the verdict
Force the judge to state its reasoning, quoting the decisive span, and only then emit the verdict on the final line. Two benefits: the reasoning constrains the verdict rather than rationalising it after the fact, and the justifications are the artifact you read when you audit the judge. A judge whose verdicts you cannot inspect is unauditable. Parse the verdict from a fixed final-line format, or use a schema-constrained response so the grade is machine-readable β see Structured Outputs.
5. Swap candidate order and average
For every pairwise comparison, run it twice with the candidates swapped. Agreement across both orders is a real preference. Disagreement means position bias decided the outcome, and the honest reading is a tie. This doubles judge cost and is almost always worth it.
6. Few-shot examples of graded output
Include two or three already-graded examples in the judge prompt, ideally ones near the decision boundary. This anchors the scale far more effectively than adjectives do, and it is how you transfer your teamβs actual standard β the one that lives in reviewersβ heads β into the instrument. Include at least one example that fails for a subtle reason.
A Concrete Rubric
Faithfulness to retrieved context, written the way a usable rubric looks:
name: answer_faithfulness
version: 4
judge_model: pinned-different-family-than-generator
scale: binary # supported or not_supported
inputs: [question, retrieved_context, answer]
instructions: |
Decide whether the answer is supported by the provided context.
Judge support by the context ONLY. Do not use outside knowledge, and do not
reward good writing. A claim is supported if a reader could point at a span
of the context that states it.
criteria:
- id: claims_traceable
text: "Every factual claim in the answer traces to a span of the context."
- id: no_added_specifics
text: "No numbers, dates, names, or conditions appear that are absent from the context."
- id: no_overreach
text: "A single example in the context is not generalised into a rule."
- id: citations_correct
text: "Each cited chunk id actually contains the claim attached to it."
- id: no_padding_credit
text: "Length, formatting, and confident tone carry no weight."
output_format: |
Write at most three sentences of justification, quoting the decisive span of
context. Then, on the final line only, write exactly one of:
VERDICT: supported
VERDICT: not_supported
If any criterion fails, the verdict is not_supported.
anchors:
- answer: "The refund window is 30 days from delivery."
context_span: "Refunds are accepted within 30 days of delivery."
verdict: supported
- answer: "The refund window is 30 days, or 45 days for damaged goods."
context_span: "Refunds are accepted within 30 days of delivery."
verdict: not_supported
reason: "the 45 day damaged goods clause does not appear in the context"
- answer: "Refunds are generally straightforward and customers are usually happy."
context_span: "Refunds are accepted within 30 days of delivery."
verdict: not_supported
reason: "fluent, on topic, and makes a claim the context does not support"
Note what the rubric does not contain: the word βquality,β any 1-to-10 scale, and any hint about which system produced the answer.
Validating the Judge Against Human Labels
This is the non-negotiable step, and the one almost everyone skips. Until you have measured agreement between your judge and human labels on the same cases, your judge score is decoration.
flowchart LR
S[Sample cases from the golden set] --> H[Human labels - the reference standard]
S --> J[Judge labels - rubric version v4]
H --> C[Compare - agreement and disagreement list]
J --> C
C --> R[Read every disagreement - ambiguous rubric or wrong judge]
R --> F[Revise rubric wording or few shot anchors]
F --> J
C --> A[Agreement clears your bar - judge approved for CI]
A --> G[Run the judge on every pull request]
G --> P[Re audit later against fresh human labels]
P --> S
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class S data
class H client
class J,C service
class R,F async
class A,G edge
class P async
How to run it. Sample somewhere in the range of a hundred cases from your golden set, spanning good, bad, and borderline output. Have humans label them against the same rubric the judge gets β if your graders need extra context to decide, your rubric is underspecified and the judge is guessing too. Then run the judge on the identical cases and compare.
Measure human-human agreement first. Have two people independently label an overlapping slice. Their agreement is the ceiling: if your own experts disagree often, the task is genuinely ambiguous and you cannot expect the judge to beat them. Teams that skip this step end up chasing a level of judge accuracy that no grader, human or model, could reach.
Use chance-corrected agreement, not raw percentage. On a skewed set where most output is fine, a judge that says βgoodβ every time scores high raw agreement and is worthless. Cohenβs kappa and similar chance-corrected statistics account for that. Inspect the confusion structure too β a judge that never produces a false pass but sometimes flags a good answer is much safer for a release gate than the reverse, and a single agreement number hides that asymmetry.
Read the disagreements individually. This is where the value is. Each one resolves into either βthe rubric was ambiguousβ (fix the rubric, and note your human labels were also unreliable there) or βthe judge is wrong in a specific wayβ (add a few-shot anchor covering it). A handful of rounds of this usually moves agreement more than switching judge models.
Re-audit on a cadence and after every change. New judge model, new rubric version, a shift in what the product does β each invalidates prior calibration. Keep the labelled validation set as a fixed artifact so re-auditing is cheap, and version the rubric so a score is always attributable to a specific instrument.
Choose a Judge Model That Is Not the Model Under Test
Self-preference makes this structural rather than stylistic. A judge grading output from its own family has a thumb on the scale, and the effect is invisible in the aggregate score. Use a different family β GPT, Claude, Gemini, Llama, Qwen, and Mistral families are all plausible judges β and when a decision matters, confirm the conclusion holds with a second judge from a third family. Disagreement between two judges is itself a useful signal: those cases are the ones worth sending to a human.
A related rule: a judge does not need to be the strongest model available, it needs to be the one with the best measured agreement on your rubric. A cheaper model with anchored few-shot examples and a tight binary rubric frequently agrees with humans better than a bigger model handed a vague scale. Measure, do not assume.
Controlling Cost
Judge calls are model calls: they cost money, add latency, and can become the most expensive part of your pipeline if you grade everything.
- Sample rather than grade exhaustively. A random slice of production traffic gives you a trend line for a fraction of the price. Reserve full-set grading for release gates.
- Screen with the cheap tiers first. Anything a deterministic assertion can catch should never reach a judge. Run schema, citation, and budget checks first and only judge what survives.
- Cascade. A small cheap judge handles the clear cases; only borderline verdicts escalate to the expensive judge or a human. This is the same tiering logic as Model Selection and Routing.
- Cache verdicts keyed on the rubric version plus a hash of the inputs and the output. Re-running an unchanged suite should cost nearly nothing. See Caching.
- Grade one dimension per call when dimensions conflict. Bundling six criteria into one verdict is cheaper but muddier; split the ones you actually gate on.
- Keep judge cost in the eval budget explicitly, so a suite that quietly triples in price shows up as a number rather than a surprise invoice.
Offline Grading Versus Online Guardrails
The same rubric gets deployed in two places with completely different constraints, and conflating them causes real incidents.
Offline grading runs in CI or a nightly job. There is no latency budget, so you can afford order swapping, multiple judges, and long justifications. Failure means a slow build, not a slow user.
An online guardrail runs inside the userβs request path β checking a response for unsupported claims before it renders. Now the judge call lands directly on your P95, roughly doubling the model latency of the request, and every judge failure becomes a user-visible failure. Three consequences follow. You need a timeout and an explicit decision about what happens when the judge times out; failing open ships an ungraded answer and failing closed shows an error, and which is correct depends on the stakes, so choose deliberately rather than by default. You need the cheapest judge that clears your agreement bar, not the best one. And you should prefer a deterministic check wherever one exists, because verifying that a citation span actually contains the claim is faster, free, and more reliable than asking a model whether it does β see Hallucination, Grounding, and Citations and Latency Engineering.
For most systems the right answer is neither extreme: grade a sample asynchronously after the response has been returned, feed the verdicts into your dashboards and failure triage, and reserve in-path judging for the narrow set of outputs where a bad answer is genuinely costly. The asynchronous path needs the trace to be retrievable after the fact, which is a tracing and observability requirement, not a judge requirement.
Bad to Good to Great
Bad - ask a model to rate quality from 1 to 10
No rubric, absolute scale, one call, whatever model is already wired up.
Scores cluster in a narrow band, so regressions vanish inside the clump. The number is uninterpretable, because nobody can say what separates a 7 from an 8. It drifts when the judge model updates. And it has never been compared to a human judgement, so its relationship to user-visible quality is unknown.
Good - a written rubric with a small scale and justifications
Named criteria, a binary or 3-point verdict, reasoning before the verdict, a pinned judge model from a different family.
This is usable for day-to-day iteration and a genuine step up. Its remaining gap is the important one: you still do not know whether the judge agrees with your teamβs standard, so you cannot defend the score to anyone who did not write it. Position and verbosity bias are also still live if you compare candidates in a fixed order.
Great - a validated, versioned instrument
- Observable criteria in a versioned rubric, stored in the repo and reviewed like code.
- Binary or small-scale verdicts, with pairwise comparison against production output where the decision is βbetter or worse.β
- Order swapped and averaged on every pairwise call, with inconsistent pairs recorded as ties.
- Justification required before the verdict, parsed from a fixed schema so it is machine-readable and human-auditable.
- Few-shot anchors near the boundary, including at least one subtle failure.
- Agreement measured against human labels with a chance-corrected statistic, human-human agreement established as the ceiling, and the confusion structure inspected for asymmetry.
- A judge from a different family than the generator, with critical conclusions confirmed by a second judge.
- Cost controlled by sampling, cheap-tier screening, cascading, and verdict caching.
- Re-audited on a cadence and after every judge or rubric change, against a preserved labelled set.
When to Use
β Use a judge when:
- The dimension you care about has no reference answer β faithfulness, helpfulness, tone, instruction adherence
- You need the measurement on every pull request, faster than humans can turn around
- You are comparing two candidate versions and can grade them pairwise against a baseline
- You have human labels, or can get a hundred of them, to validate against
- You want to triage a large volume of traffic down to the cases a human should actually read
β Do not use a judge when:
- A deterministic check would do β schema validity, a required citation, a refusal, a latency budget
- A reference answer exists and exact match or field-level scoring applies
- You have not validated it against human labels, in which case you are generating numbers, not measurements
- The stakes are high and irreversible, where the judge screens but a human decides
- The judge is the same model as the generator and nothing else corroborates the result
- You cannot state the criteria clearly enough for a human grader to apply them
Common Interview Questions
Q1: How do you know your judge is any good?
You measure agreement against human labels on the same cases, using the same rubric, and you do it before trusting a single score. Sample around a hundred cases spanning good, bad, and borderline output. Establish human-human agreement on an overlapping slice first, because that is the ceiling - if two experts disagree often, the task is ambiguous and no judge can beat them. Report a chance-corrected statistic rather than raw percentage agreement, since a lenient judge scores high raw agreement on a mostly-good set. Then read every disagreement individually: each one is either an ambiguous rubric to fix or a specific judge error to anchor with a few-shot example.
Q2: Your judge prefers whichever answer is listed first. What do you do?
That is position bias, and it is a property of the technique rather than a bug to patch. First confirm and quantify it by feeding identical pairs and counting how often a winner is declared - if that number is meaningfully above chance, every pairwise result you have collected is contaminated. The fix is to run each comparison in both orders and only count a preference when both orders agree, treating inconsistent pairs as ties. It doubles cost and is worth it. Keeping the inconsistency rate as a monitored metric is also useful, because a jump means the rubric has become harder to apply.
Q3: Why use pairwise comparison instead of scoring each answer out of ten?
Because the easier question gets the more reliable answer, and it matches the decision you actually make. Absolute scoring asks the judge to hold a stable internal standard across runs, which it does not do - scores cluster in a narrow band and drift when the judge model updates. Pairwise asks which of two candidates better satisfies stated criteria, which is a local comparison with a defensible answer, and it maps directly onto shipping: is the new version better than production. The cost is needing a baseline output stored for every golden case, plus order swapping to cancel position bias.
Q4: Can a model grade its own output?
It can, and you should expect the result to be flattering. Self-preference means output resembling the judgeβs own style scores higher, so a model grading its own family is not a neutral referee, and the bias does not announce itself in the aggregate. Use a different family as judge, and when you are choosing between two generators, confirm the conclusion holds with a second judge from a third family. There is one legitimate self-grading pattern: a model checking its own output against a mechanical criterion, such as whether a cited chunk actually contains the claim. That is verification, not preference, and it is better implemented as a deterministic check than as a judge at all.
Q5: A judge on every case is too expensive. How do you cut the cost without losing the signal?
Tier and sample. Deterministic assertions run first and for free, so nothing a schema or citation check can catch ever reaches a judge. Then cascade - a small cheap judge resolves the clear cases and only borderline verdicts escalate to the expensive judge or a human. Cache verdicts keyed on rubric version plus input and output hashes, so re-running an unchanged suite costs almost nothing. For continuous production monitoring, grade a random sample rather than all traffic, because you need a trend line, not a verdict on every request; reserve full-set grading for release gates. Finally, track judge spend as a line item in the eval budget so cost growth is visible.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts