Lesson 11 - LLM-as-Judge - Rubrics and Kappa
Code:
agentic-course/agentic/evals.pyTests:agentic-course/tests/test_memory_evals_guard.pyRun it:python3 -m unittest tests.test_memory_evals_guard -vConcept: LLM-as-Judge covers the theory and the interview framing, without code.
What you will build
- A
Judgethat grades an answer against a written rubric and returns aVerdictβ binary, with the justification attached. as_check(), so a judged criterion drops into the sameSuiteas your deterministic checks from lesson 10.agreement(), which measures the judge against human labels and reports Cohenβs kappa, not raw agreement.- A demonstration of the exact trap kappa exists to catch: a judge that says pass to everything, scoring 0.9 agreement and 0.0 kappa on a skewed set.
The idea
You have pushed everything you can into deterministic assertions. Schema validity, required facts, forbidden phrases, tool trajectory, step count, cost. What is left is the part users actually complain about: was the answer helpful, did it stay faithful to the retrieved context, did it follow the constraint it was given, did it explain or just assert. None of those has a reference answer to compare against, and no regex reaches them.
Humans can grade them. Humans cannot grade them on every commit. So you use a model as the grader, and that is a genuinely good idea β it is fast, tireless, cheap enough to run on a few hundred cases per pull request, and consistent in a way a tired reviewer is not.
Here is the trap. An unvalidated judge is an unmeasured instrument. It emits numbers with three decimal places whose relationship to the quality you care about is unknown. You have not solved your trust problem; you have moved it one layer down, where it is much harder to see, because it now arrives in the shape of data. βFaithfulness 0.87, up from 0.84β reads like a measurement and can be pure instrument noise.
The analogy: a lab reporting temperatures from a thermometer nobody ever put in ice water. The readings are precise, internally consistent, and plot beautifully over time. Whether they correspond to temperature is a separate question, and it is answered by calibration, not by looking at the readings.
Calibration is agreement(). It is the step almost everyone skips, and skipping it means every judged number downstream is decorative.
The biases, and what you do about each
A judge is a language model, so it inherits every bias a language model has β now aimed at your release gate.
- Position bias. In a pairwise comparison, the option presented first wins more often than it should. Mitigation: run both orders and only count a win when the judge agrees with itself. Disagreement is a tie, which is information rather than noise.
- Verbosity bias. Longer answers score higher independent of content. Mitigation: make length a rubric criterion in the direction you want, and sanity-check by grading a deliberately padded version of a passing answer. If the padded one still passes, the rubric is measuring effort.
- Self-preference. A model favours text that resembles its own output. Mitigation: use a different model family as judge from the one under test. This is the cheapest single improvement available, and it costs you nothing but a second provider.
- Score clustering. Ask for 1-10 and you get 7 and 8, occasionally 6. Real regressions move a 7.6 to a 7.4 and disappear into the clump. Mitigation: do not ask for a scale. Ask a question that has a defensible boundary.
- Sycophancy. Framing leaks. βThis answer was written by our improved system β grade itβ produces higher scores. Mitigation: strip all provenance from the judge prompt. The judge sees the task and the answer, never which variant produced it.
- Formatting bias. Headers, bullets and bold text read as authoritative. Mitigation: state explicitly in the rubric that presentation is not being graded, and test with a plain-prose version of a passing answer.
- Leniency drift. Your judge model gets updated under you, and the same suite grades softer than it did last month. Mitigation: pin the judge model, version the rubric (
Judge(model, rubric, version=2)), and re-run the agreement study whenever either changes. A rubric edit is an instrument change, not a copy edit.
The design of the real Judge
JUDGE_TEMPLATE = """You are grading one answer against a rubric.
Rubric:
{rubric}
Question:
{task}
Answer:
{answer}
Write at most two sentences of justification, quoting the decisive part of the
answer. Then on the final line write exactly one of:
VERDICT: pass
VERDICT: fail
"""
Three decisions are encoded there, and each one is deliberate.
Binary verdict, not a 1-10 score. Scores cluster, and a clustered score cannot detect a regression. A binary verdict also forces the rubric to be specific enough to have a boundary β writing βpass if the answer states the refund window as a number of daysβ is real work, and it is work that makes the criterion auditable. βRate the helpfulness from 1 to 10β avoids that work and buys a number nobody can defend. If you need gradation, use several binary criteria and count how many passed.
Justification before the verdict. The ordering is load-bearing in both directions. Written first, the reasoning constrains the verdict; written after, it rationalises a verdict already chosen. And because the justification is retained on the Verdict, a human can audit a disagreement in ten seconds instead of re-grading the case from scratch.
An unparseable verdict is a failure.
def grade(self, task: str, answer: str) -> Verdict:
prompt = JUDGE_TEMPLATE.format(rubric=self.rubric, task=task, answer=answer)
text = self.model.complete(
[Message(role="user", content=prompt)], max_output_tokens=300
).text
m = _VERDICT.search(text)
if not m:
# An unparseable verdict is a failure of the instrument, not a pass.
return Verdict(False, "judge returned no parseable verdict", text)
justification = _VERDICT.split(text)[0].strip()
return Verdict(m.group(1).lower() == "pass", justification, text)
test_unparseable_verdict_is_treated_as_a_failure pins that. The alternative β defaulting to pass β means a judge that silently stops emitting the verdict line after a provider update turns your entire judged suite green. A broken instrument must not read as a pass. Failing closed makes the breakage visible as a wall of failures, which is loud, ugly, and correct.
The raw text is kept on the Verdict too, so when parsing fails you can see what the judge actually said instead of guessing.
Plugging a judge into the suite
def as_check(self) -> Check:
def check(run: Run) -> tuple[bool, str]:
task = next(
(m.content for m in reversed(run.messages) if m.role == "user"), ""
)
v = self.grade(task, run.output)
return v.passed, v.justification[:200]
check.__name__ = f"judge(v{self.version})"
return check
A judged criterion becomes an ordinary Check, so it sits in the same Case as answered() and max_cost():
judge = Judge(FakeModel(["States 30 days.\nVERDICT: pass"]), RUBRIC, version=2)
report = Suite([Case(id="refund-window", task="How long do I have to refund?",
checks=[judge.as_check()], tags=("judged",))]).run(
lambda: Agent(FakeModel(["You have 30 days from delivery."]))
)
print(report.summary())
1/1 passed (100%)
mean steps 1.0 total cost 0.000000
by tag: judged=100%
Two details earn their place. The check name carries the rubric version, so a failure report says which judge failed the case β essential once a rubric has been revised. And the justification is truncated into the failure detail, so a red suite explains itself without a second query.
Tag judged cases with judged. by_tag() then separates deterministic pass rate from judged pass rate, which matters because they have different error bars and you should never average them into one headline.
agreement() - the step that makes any of this real
def agreement(self, labelled: list[tuple[str, str, bool]]) -> dict[str, float]:
tp = tn = fp = fn = 0
for task, answer, human in labelled:
judged = self.grade(task, answer).passed
...
observed = (tp + tn) / n
p_judge_pass = (tp + fp) / n
p_human_pass = (tp + fn) / n
expected = p_judge_pass * p_human_pass + (1 - p_judge_pass) * (1 - p_human_pass)
kappa = 0.0 if expected >= 1.0 else (observed - expected) / (1 - expected)
You hand it (task, answer, human_verdict) triples β a set a human has already labelled β and it grades the same items and compares. A hundred items, drawn from real traffic, deliberately including the hard and ambiguous ones, is enough to be informative.
Why kappa and not raw agreement
Raw agreement is the fraction the judge and the human got the same. On a realistic eval set most answers are fine, so a judge that says pass to everything agrees with the human on every good answer for free. Kappa corrects for that: it subtracts the agreement you would expect by chance given each raterβs pass rate, then normalises by the headroom that remains.
Work it through on the skewed set in test_always_pass_judge_has_zero_kappa_despite_high_agreement β nine human passes, one human fail, a judge that always says pass:
observed = 9/10 = 0.9
p_judge_pass = 10/10 = 1.0
p_human_pass = 9/10 = 0.9
expected = 1.0 * 0.9 + 0.0 * 0.1 = 0.9
kappa = (0.9 - 0.9) / (1 - 0.9) = 0.0
0.9 agreement. 0.0 kappa. The judge has zero skill: everything it got right, it got right by always guessing the majority class. Ship on the strength of 0.9 and you have installed a gate that cannot fail anything, on a suite whose whole purpose is to fail things. Kappa reports the truth in one number.
Rough reading: below about 0.4 the judge is not usable; 0.4 to 0.6 is usable for trend watching with human spot checks; above 0.6 is reasonable for a release gate on that criterion. Treat those as conventions rather than laws, and remember the number applies to one rubric on one distribution of answers β it does not transfer to a new criterion or a new corpus.
The confusion asymmetry
agreement() also returns the counts, and the source explains why:
"false_pass": fp, # judge said pass, human said fail - the dangerous one
"false_fail": fn,
For a release gate these are not interchangeable. A false pass lets a real regression through, silently, with a green build behind it. A false fail costs an engineer twenty minutes reading a justification and overriding it. Same contribution to kappa, wildly different blast radius.
So read the counts, not just the summary statistic. A judge with kappa 0.55 and zero false passes is a better gate than a judge with kappa 0.65 that lets two regressions through, and no single agreement number will tell you which one you have. If false passes are non-zero on a criterion that gates deploys, tighten the rubric on exactly the cases that produced them.
One more structural safeguard: use a different model family as judge than the one under test. Self-preference is real, and the failure it produces is the worst-shaped one available β a judge that is systematically lenient toward precisely the outputs it should be scrutinising. Different family, pinned version, agreement study on record.
Exercise
Build an always-pass judge, run agreement() against a skewed labelled set, and observe high raw agreement with zero kappa. Then write one sentence explaining what the metric saved you from.
Success criterion: python3 exercise_judge.py prints agreement: 0.9 with kappa: 0.0 for the lenient judge, and a contrasting kappa of 1.0 for a judge that actually reads the answer.
Worked solution
```python from agentic import FakeModel from agentic.evals import Judge RUBRIC = ( "Pass only if the answer states the refund window as a number of days " "and does not invent a policy detail." ) # Skewed on purpose: nine good answers, one bad. This is what real eval sets # look like, and it is exactly where raw agreement flatters a lenient judge. LABELLED = [("How long do I have to refund?", "You have 30 days from delivery.", True)] * 9 + [ ("How long do I have to refund?", "Refunds are never possible.", False) ] # 1. The lenient judge: says pass to everything. always_pass = Judge(FakeModel(["Looks fine.\nVERDICT: pass"] * 10), RUBRIC, version=1) print("always-pass judge: ", always_pass.agreement(LABELLED)) # 2. A judge that actually reads the answer. def grade(msgs): answer = msgs[-1].content.split("Answer:\n", 1)[-1] if "day" in answer.lower(): return "The answer names a window in days.\nVERDICT: pass" return "The answer gives no window in days.\nVERDICT: fail" reader = Judge(FakeModel([grade] * 10), RUBRIC, version=2) print("discriminating judge:", reader.agreement(LABELLED)) ``` Real output: ```text always-pass judge: {'n': 10, 'agreement': 0.9, 'kappa': 0.0, 'false_pass': 1, 'false_fail': 0} discriminating judge: {'n': 10, 'agreement': 1.0, 'kappa': 1.0, 'false_pass': 0, 'false_fail': 0} ``` One sentence: **kappa saved me from shipping a gate that could not fail anything, because raw agreement of 0.9 was measuring the skew of my labelled set rather than any skill in the judge.** Note the `false_pass: 1` on the lenient judge as well. That single count is the regression it would have waved through, and it is visible even before you look at kappa. The callable reply is worth understanding: `FakeModel` hands the scripted function the messages the judge sent, so splitting on `Answer:\n` isolates the answer from the rubric. Without that split, the word "days" in the rubric itself would make every grade a pass β a small mechanical version of the leakage that makes real judge prompts misbehave.Checkpoint
Why is model-graded evaluation necessary at all?
The dimensions users care about most β helpfulness, faithfulness to the retrieved context, instruction adherence β have no reference answer, so no deterministic check reaches them, and human review cannot run on every commit.
What is wrong with an unvalidated judge, stated precisely?
It is an unmeasured instrument that emits confident numbers with an unknown relationship to real quality. The trust problem has not been solved, only moved one layer down where it now looks like data.
Why a binary verdict instead of a 1-10 score?
Scores cluster into a narrow band and regressions hide inside the clump. A binary verdict also forces the rubric to define a defensible boundary, which is what makes a disagreement auditable.
Why does the judge return the justification before the verdict?
So the reasoning constrains the verdict rather than rationalising one already chosen, and so a human can audit a disagreement from the recorded justification instead of re-grading the case.
A judge scores 0.9 raw agreement against human labels. Why might it still be worthless?
If the labelled set is mostly passes, a judge that always says pass reaches 0.9 by never discriminating. Kappa corrects for chance agreement and reports 0.0, and the
false_passcount shows the regression it would have let through.
Theory and interview framing: Become an AI Engineer