Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 16 min read

Lesson 22 - Deployment and Rollout for Agents

Code: agentic-course/agentic/evals.py Tests: agentic-course/tests/test_memory_evals_guard.py Run it: python3 -m unittest tests.test_memory_evals_guard -v Concept: Deploying and Rolling Out AI Features covers the theory and the interview framing, without code.


What you will build


The idea

Shipping code has a comforting property: the binary you deployed is the binary that runs, until you deploy another one. Agents break that in two directions at once.

Your system’s behaviour can change with no deploy on your side. A provider updates the model behind the alias you called, and your prompt now lands differently. Your repository is untouched, your pipeline is green, and your pass rate moved.
πŸ’‘ An alias is a model name that always points at the provider’s current version, so the thing behind it can change without you doing anything. That is not a hypothetical to design around later, it is the normal operating condition of a system built on someone else’s model.

And an identical deploy can behave differently run to run. Sampling is stochastic, retrieval ties break differently as the index grows, a tool times out at 9.8 seconds today and 3 seconds tomorrow. So β€œit worked when I tested it” is a weaker statement than it is for ordinary code.

The analogy: you have shipped a dependency you did not pin, cannot see the changelog for, and cannot roll back. Everything below responds to that - pin what you can, measure what you cannot, and make sure you can always get back to a combination that worked.


The release identity

An agent release is not a commit. It is a combination, and every part of it can move behaviour independently:

@dataclass(frozen=True)
class Release:
    system_prompt_version: str     # the commonest quality change, invisible in a code diff
    model_id: str                  # a snapshot id; a floating alias is a deploy you did not make
    tool_schema_version: str       # descriptions are prompt text, so this is behaviour
    index_version: str             # same query, different corpus, different answer
    embedding_model: str           # changing it invalidates every vector you stored
    eval_suite_version: int        # scores from different suites are not comparable
    judge_rubric_version: int      # a rubric edit moves the score with the agent unchanged

Record it on every request, not only at deploy time, because a canary means two releases are live at once and an aggregate metric cannot say which one produced a given output. The tracer already nests, so wrapping the run is enough:

with tracer.span("request", "run", **asdict(RELEASE)) as req:
    run = agent.run(task)
    req.set(stop=run.stop.value, cost=run.cost)
{"span_id": "ee293f5078c4", "parent_id": null, "name": "request", "kind": "run", "attributes": {"system_prompt_version": "support-v7", "model_id": "fake-2026-03-11", "tool_schema_version": "tools-v3", "index_version": "policy-2026-03-09", "embedding_model": "hashing-256", "eval_suite_version": 1, "judge_rubric_version": 2, "stop": "answered", "cost": 0.0175}}

The point is blunt: an incident is unresolvable if you cannot say which combination produced an output. A user reports a wrong answer from last Tuesday. Without the release stamp you cannot tell whether the prompt had shipped yet, whether the index had been rebuilt, or whether the model id was the one you think it was, so every remediation you propose is a guess. With it, the first question of the investigation is answered before you start.


Eval gates in CI

The gate is a comparison, not a threshold. Run the same suite - same version, since scores from different suites are not comparable - against the release in production and against the candidate, then ask which cases got worse:

baseline = SUITE.run(build_production_agent)
candidate = SUITE.run(build_candidate_agent)
candidate.regressed_against(baseline)      # ['refund-in-window']

def regressed_against(self, baseline: "Report") -> list[str]:
    was = {r.case.id for r in baseline.results if r.passed}
    now = {r.case.id for r in self.results if r.passed}
    return sorted(was - now)

That is the whole mechanism: case ids that passed before and fail now. test_regression_against_a_baseline pins the asymmetry that makes it useful - bad.regressed_against(good) returns ["a"] while good.regressed_against(bad) returns [], so a case that was already failing does not block a change that happens not to fix it.

Wire it into CI with an exit code:

def gate(candidate_report, baseline_report) -> int:
    regressed = candidate_report.regressed_against(baseline_report)
    if regressed:
        print(f"FAIL {len(regressed)} case(s) regressed: {regressed}")
        return 1
    print("PASS no regressions")
    return 0

sys.exit(gate(candidate, baseline))

Store the baseline report as an artefact keyed by the release it came from, so the comparison is always against what is serving users and not against the last green build.

Report the delta, not the absolute

The report the gate produced on a real candidate, where a tool description edit made the model keep searching instead of answering:

pass rate   100% -> 67%
mean steps  1.3 -> 2.0
total cost  0.1105 -> 0.1490
stop mix    {'answered': 3} -> {'answered': 2, 'step_budget_exhausted': 1}
  tag adversarial  100% -> 100%
  tag happy        100% -> 0%
  tag policy       100% -> 50%
  tag security     100% -> 100%
regressed   ['refund-in-window']

Every line is a pair. That is deliberate, and it is the single change that makes eval results usable in a team. Nobody can judge whether 0.81 is good. Everybody can judge that it used to be 0.86. An absolute pass rate invites a debate about whether the suite is too hard; a delta against the version currently serving users invites a decision about whether to ship.

by_tag() is in the report for the reason its docstring gives: the aggregate hides compensating movement, and a candidate can hold pass rate flat while security regresses and an easy category improves. Here the aggregate moved and the breakdown localised it to happy and policy with security untouched, which tells you where to look before you open a single trace.

Progressive rollout, adapted for agents

flowchart LR
    CAND[Candidate release built] --> SHADOW[Shadow mode - output discarded]
    SHADOW --> CMP[Offline compare against production]
    CMP --> CANARY[Canary on a traffic slice]
    CANARY --> WATCH[Watch stop mix and cost per run]
    WATCH --> FULL[Full rollout]
    WATCH --> FLAG[Flag flip back to previous release]
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
    class CAND,CANARY,FULL service
    class SHADOW,CMP async
    class WATCH,FLAG data

Shadow mode runs the candidate on real traffic and throws its output away. The user sees only production’s answer, so a bad candidate cannot hurt anyone. And you get the one thing an offline suite cannot give you: the real input distribution, including the malformed, multilingual and adversarial inputs nobody thought to write a golden case for.
πŸ’‘ A golden case is one test input paired with the answer you have agreed is correct, written down ahead of time so a later change can be scored against it.
Two agent-specific cautions. You pay twice for every shadowed request, so sample rather than shadowing everything. And shadow mode is only safe if the candidate’s mutating tools are disabled or pointed at a sandbox - a shadow run that issues a real refund is not a shadow run. Registry(tools, allowed={...}) enforces that, and scoped() gives you the narrowed view.

Canary on a traffic slice is where the candidate’s output actually reaches users. Route by a stable hash of the user id rather than per request, so one person does not get the new agent on turn 1 and the old one on turn 2 mid-conversation.
πŸ’‘ A canary is a deliberately small share of live traffic, sized so that a bad release is something you notice rather than something that lands on everybody at once. Flag-based rollback then means the switch is configuration read at request time, not a redeploy: when a canary goes wrong you want the next request on the old release, not the next deploy. The general vocabulary - blue-green, ring, bake time - is in deployment reliability, and the model-serving version in AI deployment.


What to watch during a rollout

The stop-reason mix, first. The agent-specific signal, and it leads everything else:

baseline   {'answered': 3}          candidate   {'answered': 2, 'step_budget_exhausted': 1}

Budget stops appearing on cases that used to pass is the earliest regression signal you get. It fires before pass rate moves much, because a run that hits STEP_BUDGET still returns a string and a lenient check may still accept it. And it localises the problem immediately. STEP_BUDGET climbing means the agent is taking more steps than it used to. REPEATED_CALL climbing means a tool result stopped answering the question the model is asking, and TOOL_FATAL climbing is an authorization or credentials problem rather than a prompt problem. A pass-rate drop says something is wrong; the stop mix says which kind of wrong. The full table is in lesson 8. Then, in order:

A tool description change is a prompt change

This one catches experienced teams, because of how the diff looks:

@tool(description="Search company policy and runbooks. Use for questions about rules.",
      q="A short search phrase, not a full sentence.")
def search_docs(q: str) -> str:
    ...

Editing that string looks like tidying a docstring. It is not. The description and the parameter docs are serialised into to_schema() and sent to the model on every single call, which makes them prompt text going through a code review that does not treat them as prompt text. Rewording β€œuse for questions about rules” is a change to tool selection, and the failure is not an exception - it is the model quietly picking a different tool, which surfaces as a step-count increase or a REPEATED_CALL and gets blamed on the model. So version tool schemas, put them in the release identity, and gate a description edit on the eval suite exactly as you would gate a system prompt edit. The candidate in the delta report above is this scenario, and the gate caught it.

Rollback that is genuinely possible

β€œWe can roll back” is usually a claim about the application container only. For an agent the rollback has to restore the whole combination, so the previous prompt, model pin, tool schemas and index snapshot all have to remain deployable together. The prompt lives somewhere you can fetch an old version from rather than being edited in place. The model pin requires that you pinned a snapshot at all - you cannot roll back a floating alias, because there is nothing to roll back to. The index snapshot is retained rather than overwritten, because a rebuilt index is not reversible by redeploying code. And if the embedding model changed, the old vectors are not comparable to the new ones, so re-embedding the corpus is the only path back.
πŸ’‘ An embedding model turns text into a list of numbers, and two different ones produce numbers you cannot compare, which is why swapping it means processing the whole corpus again.

The test is not a document. It is whether you can name the previous release’s seven fields right now and fetch all seven artefacts. If the index is overwritten in place or the model was called through an alias, your rollback story has a hole in it, and you will find the hole during an incident.


Exercise

Record a baseline report, regress the agent deliberately, and gate on regressed_against returning a non-empty list. Save as tests/test_exercise_gate.py inside agentic-course/. Success criterion: python3 -m unittest tests.test_exercise_gate -v reports OK, with one test asserting an identical release produces [] and one asserting the regressed release is named and stopped with Stop.STEP_BUDGET.

Worked solution ```python import unittest from agentic import Agent, Budget, FakeModel, Registry, Stop, tool, tool_call from agentic.evals import Case, Suite, answered, contains, max_steps, used_tools PRICES = dict(price_per_1k_input=0.5, price_per_1k_output=1.5) # assumed unit prices @tool(description="Fetch the refund policy section.", topic="A policy topic.") def fetch_policy(topic: str) -> str: return "Refunds: within 30 days." SUITE = Suite(name="support-agent", version=1, cases=[ Case(id="refund-in-window", task="Can I refund order 4471?", tags=("policy",), why="the core question the agent exists to answer", checks=[answered(), used_tools("fetch_policy"), contains("30"), max_steps(4)]), ]) def reply_good(msgs): if not any(m.role == "tool" for m in msgs): return tool_call("fetch_policy", {"topic": "refunds"}) return "Yes - order 4471 is inside the 30 day refund window." def reply_regressed(msgs): # A description edit made the model keep searching instead of answering. return tool_call("fetch_policy", {"topic": f"refunds {len(msgs)}"}) def factory(reply): def build(): return Agent(FakeModel([reply] * 12), Registry([fetch_policy]), system="You are a support agent.", budget=Budget(max_steps=4, **PRICES)) return build class TestRegressionGate(unittest.TestCase): def test_gate_is_green_on_an_identical_release(self): baseline = SUITE.run(factory(reply_good)) self.assertEqual(baseline.pass_rate, 1.0) self.assertEqual(SUITE.run(factory(reply_good)).regressed_against(baseline), []) def test_gate_catches_a_deliberate_regression(self): baseline = SUITE.run(factory(reply_good)) candidate = SUITE.run(factory(reply_regressed)) self.assertEqual(candidate.regressed_against(baseline), ["refund-in-window"]) self.assertIs(candidate.results[0].run.stop, Stop.STEP_BUDGET) # measured: pass rate 100% -> 0% ``` Two details are load-bearing. `build_agent` is a **callable** rather than an agent instance, so each case gets a fresh agent and a fresh `FakeModel` - a shared agent between cases is how an eval suite starts lying about which case caused what. And the last assertion is the one worth copying into your own harness. `regressed_against` *names* the regression, but the reason lives in the stop reason: `STEP_BUDGET`, not a wrong answer. Two candidates with identical pass-rate drops, one stopping `ANSWERED` and one `STEP_BUDGET`, need completely different investigations - a quality problem versus a control-flow problem.

Checkpoint

Why can an agent’s behaviour change with no deploy on your side?

Because the model is a dependency you do not control. A provider updating the model behind a floating alias changes how your prompt lands while your repository, your pipeline and your config stay exactly as they were.

Why report a delta instead of an absolute pass rate, and what does by_tag() add?

An absolute number is not interpretable - nobody can say whether 0.81 is good, so the conversation becomes a debate about the suite, while everyone can act on β€œit used to be 0.86”. by_tag() then localises the movement, because an aggregate can stay flat while security regresses and an easy category improves.

Which signal moves earliest in a bad rollout, and why that one?

The stop-reason mix. Budget stops appearing on cases that used to pass show up before pass rate moves much. A truncated run still returns a string that a lenient check may accept. The stop reason also tells you which kind of failure it is.

Why does a tool description edit need an eval gate even though it looks like a docstring change?

Because the description is serialised into the schema and sent to the model on every call, which makes it prompt text. Rewording it changes tool selection, and the symptom is a quiet step-count increase rather than an error.

What makes an agent rollback incomplete even when the code rolls back cleanly?

The prompt, model pin, tool schemas and index snapshot have to remain deployable together. A floating model alias has nothing to roll back to, and an index overwritten in place is not restored by redeploying code - especially if the embedding model changed, since the old vectors are no longer comparable.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access