Lesson 22 - Deployment and Rollout for Agents
Code:
agentic-course/agentic/evals.pyTests:agentic-course/tests/test_memory_evals_guard.pyRun it:python3 -m unittest tests.test_memory_evals_guard -vConcept: Deploying and Rolling Out AI Features covers the theory and the interview framing, without code.
What you will build
- A
Releaserecord versioning everything that can change behaviour, stamped onto every requestβs trace - An eval gate on the real
Report.regressed_against(baseline), so a regression fails the build instead of reaching users - A delta report that compares against production rather than reporting an absolute score
- The stop-reason mix as your earliest rollout signal, read off a real report
The idea
Shipping code has a comforting property: the binary you deployed is the binary that runs, until you deploy another one. Agents break that in two directions at once.
Your systemβs behaviour can change with no deploy on your side. A provider updates the model behind the alias you called, and your prompt now lands differently. Your repository is untouched, your pipeline is green, and your pass rate moved.
π‘ An alias is a model name that always points at the providerβs current version, so the thing behind it can change without you doing anything. That is not a hypothetical to design around later, it is the normal operating condition of a system built on someone elseβs model.
And an identical deploy can behave differently run to run. Sampling is stochastic, retrieval ties break differently as the index grows, a tool times out at 9.8 seconds today and 3 seconds tomorrow. So βit worked when I tested itβ is a weaker statement than it is for ordinary code.
The analogy: you have shipped a dependency you did not pin, cannot see the changelog for, and cannot roll back. Everything below responds to that - pin what you can, measure what you cannot, and make sure you can always get back to a combination that worked.
The release identity
An agent release is not a commit. It is a combination, and every part of it can move behaviour independently:
@dataclass(frozen=True)
class Release:
system_prompt_version: str # the commonest quality change, invisible in a code diff
model_id: str # a snapshot id; a floating alias is a deploy you did not make
tool_schema_version: str # descriptions are prompt text, so this is behaviour
index_version: str # same query, different corpus, different answer
embedding_model: str # changing it invalidates every vector you stored
eval_suite_version: int # scores from different suites are not comparable
judge_rubric_version: int # a rubric edit moves the score with the agent unchanged
Record it on every request, not only at deploy time, because a canary means two releases are live at once and an aggregate metric cannot say which one produced a given output. The tracer already nests, so wrapping the run is enough:
with tracer.span("request", "run", **asdict(RELEASE)) as req:
run = agent.run(task)
req.set(stop=run.stop.value, cost=run.cost)
{"span_id": "ee293f5078c4", "parent_id": null, "name": "request", "kind": "run", "attributes": {"system_prompt_version": "support-v7", "model_id": "fake-2026-03-11", "tool_schema_version": "tools-v3", "index_version": "policy-2026-03-09", "embedding_model": "hashing-256", "eval_suite_version": 1, "judge_rubric_version": 2, "stop": "answered", "cost": 0.0175}}
The point is blunt: an incident is unresolvable if you cannot say which combination produced an output. A user reports a wrong answer from last Tuesday. Without the release stamp you cannot tell whether the prompt had shipped yet, whether the index had been rebuilt, or whether the model id was the one you think it was, so every remediation you propose is a guess. With it, the first question of the investigation is answered before you start.
Eval gates in CI
The gate is a comparison, not a threshold. Run the same suite - same version, since scores from different suites are not comparable - against the release in production and against the candidate, then ask which cases got worse:
baseline = SUITE.run(build_production_agent)
candidate = SUITE.run(build_candidate_agent)
candidate.regressed_against(baseline) # ['refund-in-window']
def regressed_against(self, baseline: "Report") -> list[str]:
was = {r.case.id for r in baseline.results if r.passed}
now = {r.case.id for r in self.results if r.passed}
return sorted(was - now)
That is the whole mechanism: case ids that passed before and fail now. test_regression_against_a_baseline pins the asymmetry that makes it useful - bad.regressed_against(good) returns ["a"] while good.regressed_against(bad) returns [], so a case that was already failing does not block a change that happens not to fix it.
Wire it into CI with an exit code:
def gate(candidate_report, baseline_report) -> int:
regressed = candidate_report.regressed_against(baseline_report)
if regressed:
print(f"FAIL {len(regressed)} case(s) regressed: {regressed}")
return 1
print("PASS no regressions")
return 0
sys.exit(gate(candidate, baseline))
Store the baseline report as an artefact keyed by the release it came from, so the comparison is always against what is serving users and not against the last green build.
Report the delta, not the absolute
The report the gate produced on a real candidate, where a tool description edit made the model keep searching instead of answering:
pass rate 100% -> 67%
mean steps 1.3 -> 2.0
total cost 0.1105 -> 0.1490
stop mix {'answered': 3} -> {'answered': 2, 'step_budget_exhausted': 1}
tag adversarial 100% -> 100%
tag happy 100% -> 0%
tag policy 100% -> 50%
tag security 100% -> 100%
regressed ['refund-in-window']
Every line is a pair. That is deliberate, and it is the single change that makes eval results usable in a team. Nobody can judge whether 0.81 is good. Everybody can judge that it used to be 0.86. An absolute pass rate invites a debate about whether the suite is too hard; a delta against the version currently serving users invites a decision about whether to ship.
by_tag() is in the report for the reason its docstring gives: the aggregate hides compensating movement, and a candidate can hold pass rate flat while security regresses and an easy category improves. Here the aggregate moved and the breakdown localised it to happy and policy with security untouched, which tells you where to look before you open a single trace.
Progressive rollout, adapted for agents
flowchart LR
CAND[Candidate release built] --> SHADOW[Shadow mode - output discarded]
SHADOW --> CMP[Offline compare against production]
CMP --> CANARY[Canary on a traffic slice]
CANARY --> WATCH[Watch stop mix and cost per run]
WATCH --> FULL[Full rollout]
WATCH --> FLAG[Flag flip back to previous release]
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class CAND,CANARY,FULL service
class SHADOW,CMP async
class WATCH,FLAG data
Shadow mode runs the candidate on real traffic and throws its output away. The user sees only productionβs answer, so a bad candidate cannot hurt anyone. And you get the one thing an offline suite cannot give you: the real input distribution, including the malformed, multilingual and adversarial inputs nobody thought to write a golden case for.
π‘ A golden case is one test input paired with the answer you have agreed is correct, written down ahead of time so a later change can be scored against it.
Two agent-specific cautions. You pay twice for every shadowed request, so sample rather than shadowing everything. And shadow mode is only safe if the candidateβs mutating tools are disabled or pointed at a sandbox - a shadow run that issues a real refund is not a shadow run. Registry(tools, allowed={...}) enforces that, and scoped() gives you the narrowed view.
Canary on a traffic slice is where the candidateβs output actually reaches users. Route by a stable hash of the user id rather than per request, so one person does not get the new agent on turn 1 and the old one on turn 2 mid-conversation.
π‘ A canary is a deliberately small share of live traffic, sized so that a bad release is something you notice rather than something that lands on everybody at once. Flag-based rollback then means the switch is configuration read at request time, not a redeploy: when a canary goes wrong you want the next request on the old release, not the next deploy. The general vocabulary - blue-green, ring, bake time - is in deployment reliability, and the model-serving version in AI deployment.
What to watch during a rollout
The stop-reason mix, first. The agent-specific signal, and it leads everything else:
baseline {'answered': 3} candidate {'answered': 2, 'step_budget_exhausted': 1}
Budget stops appearing on cases that used to pass is the earliest regression signal you get. It fires before pass rate moves much, because a run that hits STEP_BUDGET still returns a string and a lenient check may still accept it. And it localises the problem immediately. STEP_BUDGET climbing means the agent is taking more steps than it used to. REPEATED_CALL climbing means a tool result stopped answering the question the model is asking, and TOOL_FATAL climbing is an authorization or credentials problem rather than a prompt problem. A pass-rate drop says something is wrong; the stop mix says which kind of wrong. The full table is in lesson 8. Then, in order:
- Pass rate, by tag - the aggregate plus the breakdown, never the aggregate alone
- p95 latency and cost per run -
report.cost()andmean_steps(). A candidate that adds a step adds a round trip and a bigger context, so these move together, and a quality win that triples cost per resolution is a business decision rather than an engineering one - Tool error rate -
run.blocked_tools()growing means arguments are getting worse or a dependency is degrading - Abstention rate - the one people forget. A candidate that stopped saying βI could not find thatβ has usually started bluffing, and bluffing scores well on a lenient rubric.
abstained()is a check for exactly this, and a retrieval agent with a zero abstention rate is guessing.
π‘ Abstention is the agent answering βI could not find thatβ instead of inventing something, and it counts as a correct answer when the documents really do not say.
A tool description change is a prompt change
This one catches experienced teams, because of how the diff looks:
@tool(description="Search company policy and runbooks. Use for questions about rules.",
q="A short search phrase, not a full sentence.")
def search_docs(q: str) -> str:
...
Editing that string looks like tidying a docstring. It is not. The description and the parameter docs are serialised into to_schema() and sent to the model on every single call, which makes them prompt text going through a code review that does not treat them as prompt text. Rewording βuse for questions about rulesβ is a change to tool selection, and the failure is not an exception - it is the model quietly picking a different tool, which surfaces as a step-count increase or a REPEATED_CALL and gets blamed on the model. So version tool schemas, put them in the release identity, and gate a description edit on the eval suite exactly as you would gate a system prompt edit. The candidate in the delta report above is this scenario, and the gate caught it.
Rollback that is genuinely possible
βWe can roll backβ is usually a claim about the application container only. For an agent the rollback has to restore the whole combination, so the previous prompt, model pin, tool schemas and index snapshot all have to remain deployable together. The prompt lives somewhere you can fetch an old version from rather than being edited in place. The model pin requires that you pinned a snapshot at all - you cannot roll back a floating alias, because there is nothing to roll back to. The index snapshot is retained rather than overwritten, because a rebuilt index is not reversible by redeploying code. And if the embedding model changed, the old vectors are not comparable to the new ones, so re-embedding the corpus is the only path back.
π‘ An embedding model turns text into a list of numbers, and two different ones produce numbers you cannot compare, which is why swapping it means processing the whole corpus again.
The test is not a document. It is whether you can name the previous releaseβs seven fields right now and fetch all seven artefacts. If the index is overwritten in place or the model was called through an alias, your rollback story has a hole in it, and you will find the hole during an incident.
Exercise
Record a baseline report, regress the agent deliberately, and gate on regressed_against returning a non-empty list. Save as tests/test_exercise_gate.py inside agentic-course/. Success criterion: python3 -m unittest tests.test_exercise_gate -v reports OK, with one test asserting an identical release produces [] and one asserting the regressed release is named and stopped with Stop.STEP_BUDGET.
Worked solution
```python import unittest from agentic import Agent, Budget, FakeModel, Registry, Stop, tool, tool_call from agentic.evals import Case, Suite, answered, contains, max_steps, used_tools PRICES = dict(price_per_1k_input=0.5, price_per_1k_output=1.5) # assumed unit prices @tool(description="Fetch the refund policy section.", topic="A policy topic.") def fetch_policy(topic: str) -> str: return "Refunds: within 30 days." SUITE = Suite(name="support-agent", version=1, cases=[ Case(id="refund-in-window", task="Can I refund order 4471?", tags=("policy",), why="the core question the agent exists to answer", checks=[answered(), used_tools("fetch_policy"), contains("30"), max_steps(4)]), ]) def reply_good(msgs): if not any(m.role == "tool" for m in msgs): return tool_call("fetch_policy", {"topic": "refunds"}) return "Yes - order 4471 is inside the 30 day refund window." def reply_regressed(msgs): # A description edit made the model keep searching instead of answering. return tool_call("fetch_policy", {"topic": f"refunds {len(msgs)}"}) def factory(reply): def build(): return Agent(FakeModel([reply] * 12), Registry([fetch_policy]), system="You are a support agent.", budget=Budget(max_steps=4, **PRICES)) return build class TestRegressionGate(unittest.TestCase): def test_gate_is_green_on_an_identical_release(self): baseline = SUITE.run(factory(reply_good)) self.assertEqual(baseline.pass_rate, 1.0) self.assertEqual(SUITE.run(factory(reply_good)).regressed_against(baseline), []) def test_gate_catches_a_deliberate_regression(self): baseline = SUITE.run(factory(reply_good)) candidate = SUITE.run(factory(reply_regressed)) self.assertEqual(candidate.regressed_against(baseline), ["refund-in-window"]) self.assertIs(candidate.results[0].run.stop, Stop.STEP_BUDGET) # measured: pass rate 100% -> 0% ``` Two details are load-bearing. `build_agent` is a **callable** rather than an agent instance, so each case gets a fresh agent and a fresh `FakeModel` - a shared agent between cases is how an eval suite starts lying about which case caused what. And the last assertion is the one worth copying into your own harness. `regressed_against` *names* the regression, but the reason lives in the stop reason: `STEP_BUDGET`, not a wrong answer. Two candidates with identical pass-rate drops, one stopping `ANSWERED` and one `STEP_BUDGET`, need completely different investigations - a quality problem versus a control-flow problem.Checkpoint
Why can an agentβs behaviour change with no deploy on your side?
Because the model is a dependency you do not control. A provider updating the model behind a floating alias changes how your prompt lands while your repository, your pipeline and your config stay exactly as they were.
Why report a delta instead of an absolute pass rate, and what does by_tag() add?
An absolute number is not interpretable - nobody can say whether 0.81 is good, so the conversation becomes a debate about the suite, while everyone can act on βit used to be 0.86β.
by_tag()then localises the movement, because an aggregate can stay flat while security regresses and an easy category improves.
Which signal moves earliest in a bad rollout, and why that one?
The stop-reason mix. Budget stops appearing on cases that used to pass show up before pass rate moves much. A truncated run still returns a string that a lenient check may accept. The stop reason also tells you which kind of failure it is.
Why does a tool description edit need an eval gate even though it looks like a docstring change?
Because the description is serialised into the schema and sent to the model on every call, which makes it prompt text. Rewording it changes tool selection, and the symptom is a quiet step-count increase rather than an error.
What makes an agent rollback incomplete even when the code rolls back cleanly?
The prompt, model pin, tool schemas and index snapshot have to remain deployable together. A floating model alias has nothing to roll back to, and an index overwritten in place is not restored by redeploying code - especially if the embedding model changed, since the old vectors are no longer comparable.
Theory and interview framing: Become an AI Engineer