Deploying and Rolling Out AI Features - Complete Deep Dive
Prerequisites: Evals, Deployment and Reliability, LLM Observability Used in: Multimodal, Portfolio Projects, The AI Engineer Interview Build it: Lesson 22 - Deployment and Rollout for Agents implements this as runnable, tested code you can execute offline.
What Makes Deploying AI Different?
Deploying an AI feature is ordinary release engineering with two of its foundational assumptions removed. Everything in Deployment and Reliability still applies β flags, canaries, health checks, rollback. What breaks is the pair of beliefs that make those tools feel safe: that your system only changes when you deploy, and that the same code does the same thing twice.
Asymmetry one: behaviour can change with no deploy on your side. Your provider updates a model, deprecates a snapshot, adjusts a safety filter, or shifts default parameters, and your untouched, unreleased, fully-tested system starts behaving differently in production. No commit, no pipeline run, no changelog entry of yours. Nothing in a conventional deployment process is watching for this, because conventional deployment processes assume the cause of a change is a change.
Asymmetry two: an identical deploy can behave differently run to run. Sampling is stochastic. The same build, the same prompt, the same input, twice β two outputs. So βit worked when I tested itβ is a weaker claim than usual, and a single failing observation is not automatically a regression.
Those two facts point at the same two controls, and they carry the whole practice: version pinning, so behaviour only changes when you decide it changes, and eval gates, so you can tell a real regression from noise. Every other technique on this page hangs off those.
Real-world analogy: a restaurant that changes its menu but not its suppliers. Normally the dish changes when the chef changes the recipe. Here the supplier can substitute an ingredient without telling you, and the same recipe cooked twice comes out slightly differently. The fix is not a better recipe. It is naming the exact supplier and batch you accept, and tasting every dish against a standard before it leaves the kitchen.
Release Identity - What Must Be Versioned Together
A deploy of an AI feature is not a code artifact. It is a combination, and the combination is the unit you ship, gate, and roll back.
| Component | Why it belongs in the release | What breaks without it |
|---|---|---|
| Prompt version | The prompt is program logic | You cannot reproduce an output or bisect a quality drop |
| Model id and pinned snapshot | A family name is not a version - pin the specific snapshot | Behaviour changes with no deploy and you cannot prove it did |
| Sampling parameters | Temperature and top_p change the output distribution | Two environments quietly disagree about determinism |
| Retrieval index version | The same question over a different corpus is a different question | A retrieval regression looks like a model regression |
| Embedding model version | It determines the vector space the index lives in | Old and new vectors silently become incomparable |
| Tool schema version | Available tools and argument shapes change behaviour | Tool errors spike and nobody knows what moved |
| Rubric or judge version | It defines what your score means | Eval scores are not comparable across time |
Record all of it on every request as a single release identity, alongside the trace. Then your logs answer the only question that matters during an incident: which combination produced this output. An incident is genuinely unresolvable without that β βquality dropped on Tuesdayβ with no release identity leaves you guessing across seven independently-moving parts, and the guess is usually wrong because the change was in the one you did not touch. The trace plumbing for this is in LLM Observability.
Two corollaries worth stating outright. Never point production at an unpinned model alias, because that is opting into an upgrade on someone elseβs schedule. And treat prompts and config as deploys, not as rows you edit in a live database. A prompt is the most behaviour-dense code in the system; editing it in a console with no review, no version, no eval run, and no rollback path is shipping straight to production. Keep prompts in version control, ship them through the pipeline, and let a flag control which version is active β see Prompt Engineering.
Eval Gates Are the Deployment Gate
For ordinary services the gate is βtests pass and the health check is greenβ. For an AI feature the gate is an eval score that did not regress, and the mechanism is the same: a failing gate fails the build.
- Deterministic checks on every commit β schema validity, required citations, refusals on inputs that must be refused, latency and cost ceilings. Cheap enough to be unconditional.
- The scored suite on every pull request, reported as a delta against the release currently in production. Absolute numbers are unreadable; deltas are not.
- A regression fails the build. If it only warns, it gets ignored under deadline. This is a cultural line more than a technical one.
- Hold noise constant. Pin the model snapshot, temperature zero for eval runs, frozen index snapshot, fixed seeds where supported. On genuinely noisy dimensions, run each case several times and compare distributions rather than blocking on one unlucky sample.
- Gate the combination, not the diff. A prompt change is evaluated against the model pin and index version it will actually ship with.
The full construction of the suite β tiers, golden set, held-out slice β is Evals. What matters here is that it is wired into the pipeline rather than run when someone remembers.
Shadow Mode
Run the candidate path on real production traffic without using its output. The user is served by the current release; the candidateβs response is recorded and compared offline.
flowchart LR
U[User request] --> R[Router]
R --> P[Primary path - current release]
P --> O[Response served to user]
R --> S[Shadow path - candidate release]
S --> D[Output discarded]
P --> L[Trace store keyed by release identity]
S --> L
L --> C[Offline comparison and eval scoring]
C --> G[Promote or reject the candidate]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class U,O client
class R edge
class P,S service
class D,C,G async
class L data
Shadow mode buys the thing offline evals cannot: real traffic distribution with zero user risk. Your golden set contains the inputs you thought of; production contains the ones you did not. You get a paired comparison on identical inputs, which is a far stronger signal than comparing two scores from two different samples.
What it costs and constrains:
- You pay twice for every shadowed request, so sample a slice rather than mirroring everything.
- Side effects must be suppressed. The shadow path must not write to primary stores, send notifications, charge anything, or call mutating tools. Route it at stubs or a sandboxed tool layer, or you will discover the hard way that βthe output was discardedβ did not mean βnothing happenedβ.
- It cannot measure user reaction. Nobody saw the output, so acceptance, edits, and escalation are unavailable. Shadow mode answers βis it different and is it better by our rubricβ, not βdo users prefer itβ.
- Latency is not free if the shadow call shares a connection pool or a rate limit budget with the primary path. Isolate the quota.
Canary, A/B, and Instant Rollback
Once shadow mode says the candidate is plausible, expose it to a small slice of real users.
flowchart LR
T[Traffic] --> F[Feature flag split]
F -->|majority| A[Stable release]
F -->|small slice| B[Canary release]
A --> M[Metrics grouped by release identity]
B --> M
M --> G[Rollback guard comparing canary against stable]
G -->|inside thresholds| W[Widen the slice]
G -->|threshold breach| K[Flip flag back to stable]
W --> F
K --> F
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class T client
class F edge
class A,B service
class M data
class G,W,K async
Canary widens in steps, with the guard evaluated at each step and a defined threshold that triggers an automatic flip back. Keep the assignment sticky per user, because a user whose assistant changes personality mid-conversation experiences a bug even when both releases are fine.
A/B goes further and measures an online quality metric β task completion, acceptance rate, edit distance on accepted output, escalation to a human. This is the only way to answer βdo users actually prefer itβ, and it is how you catch the candidate that scores better on your rubric and worse with people, which happens often enough to be worth expecting. It needs enough traffic and enough time to be meaningful, so do not read an A/B on an afternoon of data.
Flag-based instant rollback is the non-negotiable part. Rollback must be a configuration flip that takes effect in seconds, not a pipeline run. For an AI feature that means the flag selects a whole release identity β prompt version, model pin, index snapshot, tool schema β and not just a code path. Upstream of all this, a circuit breaker on the provider call with a defined degraded mode handles the failure you cannot roll back from, because the outage is not yours.
What to Watch During a Rollout
Segment every one of these by release identity, or you are averaging the canary into the stable population and hiding exactly what you are looking for.
| Signal | What a move means |
|---|---|
| Eval score on live traffic samples | The direct quality read, and the one worth alerting on |
| Schema validation failure rate | The candidate stopped honouring the output contract - see Structured Outputs |
| Abstention rate | A spike suggests retrieval broke; a collapse to zero suggests the model started bluffing |
| Escalation rate to a human | Users voting with their feet, and hard to argue with |
| p95 time to first token and total latency | A quality win that blows the latency budget is not a win |
| Cost per request | Prompt bloat, retry storms, or an expensive model creeping into a cheap path |
| Tool error rate and retry rate | A tool schema or argument-shape regression |
Alert on ratios, not counts. Absolute counts move with traffic, so a 3am dip in errors is just fewer users and a Monday spike is just Monday. Failures per thousand requests, abstentions as a share of answers, escalations as a share of sessions, cost per successful request β these are comparable across the canary slice and the stable population, which is the entire point of running them side by side.
Pick a small number of these as the automatic rollback trigger before the rollout starts, and write down the threshold. A trigger chosen during an incident is a judgement call made under pressure.
The Embedding Model Migration
This is the single most under-planned change in RAG systems, and it deserves its own plan every time.
Changing the embedding model changes the vector space itself. Old vectors and new vectors are not comparable, so there is no partial state in which a query embedded by the new model can search an index built by the old one and return anything meaningful. A flag flip does not work. This is not a deploy, it is a migration: the entire corpus must be re-embedded and re-indexed. See Embeddings and Vector Databases.
flowchart LR
I[Ingest pipeline] --> E1[Embedding model v1]
I --> E2[Embedding model v2]
B[Backfill job over the full corpus] --> E2
E1 --> X1[Index v1 - serving]
E2 --> X2[Index v2 - shadow]
Q[Query] --> R1[Retriever on index v1]
R1 --> A[Answer served]
Q --> R2[Shadow retriever on index v2]
R2 --> C[Offline recall comparison]
C --> P[Cut over then retire index v1]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class Q,A client
class R1,R2 edge
class I,E1,E2 service
class X1,X2 data
class B,C,P async
The sequence:
- Stand up a second index. Do not migrate in place; you need both live to compare and to roll back.
- Dual-write from ingest. Every new or updated document is embedded by both models so the new index does not fall behind while the backfill runs.
- Backfill the corpus into the new index as a resumable, rate-limited batch job. On a large corpus this runs for a long time and costs real money in embedding calls β budget both.
- Dual-read for comparison. Queries serve from the old index while the new index answers in shadow. Compare retrieval metrics on the same queries, then generation quality on the frozen golden set.
- Cut over behind a flag once the new index is complete and measurably at least as good, canarying the read path the same way you would any other change.
- Keep the old index deployable until confidence is high. Retire it deliberately, not as cleanup.
Also invalidate what depends on the old space: any semantic cache keyed on old embeddings must be rebuilt rather than reused, since its stored vectors belong to a space that no longer exists β see Prompt and Semantic Caching.
Change Type, Blast Radius, Gate, Rollback
| Change | Blast radius | Required gate | Rollback path |
|---|---|---|---|
| Prompt wording tweak | All traffic on that path, immediately | Deterministic checks plus scored suite delta | Flip prompt version flag |
| Prompt structural rewrite or new tool instructions | All traffic, behaviour shift can be large | Full suite plus shadow mode plus canary | Flip prompt version flag |
| Sampling parameter change | All traffic, mostly variance and cost | Scored suite run several times to see distribution | Config flip |
| Model snapshot bump inside a family | All traffic, behaviour shift often subtle | Full suite plus shadow comparison plus canary | Repin the previous snapshot |
| Model swap across families or providers | All traffic, prompt may not transfer at all | Re-tune the prompt then full suite then A/B - see Model Selection | Flip router back, keep old client path deployable |
| Tool schema change | Every call using that tool, plus caches keyed on it | Deterministic argument checks plus canary on tool error rate | Version the schema and serve the old one |
| Re-index with the same embedding model | Retrieval quality and freshness | Retrieval metrics on a fixed query set | Serve the previous index snapshot |
| Embedding model change | Whole retrieval layer, no partial valid state | Dual-write plus backfill plus dual-read comparison | Point reads at the old index while it still exists |
| Chunking or parsing strategy change | Whole retrieval layer, requires re-index | Retrieval metrics plus generation on frozen context | Previous index snapshot |
| Rubric or judge version change | No user impact, but all scores shift meaning | Re-score a known release to rebaseline | Revert rubric version and re-score |
The pattern in the right-hand column: rollback is only real if the whole previous combination is still deployable together. Rolling the prompt back onto a model snapshot that was retired, or onto an index that was migrated in place, is not a rollback. Before shipping, ask the mechanical question β if this goes wrong in ten minutes, can I restore the previous prompt, model pin, and index snapshot as a set, right now. If any one of the three is gone, the rollout does not start.
Bad to Good to Great
Bad - edit the prompt in production and watch
No version, no eval, unpinned model alias, no release identity on requests. It works until quality drops and nobody can say what changed, at which point the only available response is more prompt edits, each one an unmeasured bet.
Good - prompts in version control, evals in CI, flag-controlled rollout
Prompts and config ship through the pipeline, a scored suite gates the build, a pinned model snapshot, a flag for instant rollback, and an alert on error rate. Enough for many teams. The gaps are usually no shadow stage, no release identity on individual requests, and no plan for an embedding migration.
Great - release identity plus staged exposure plus a rehearsed rollback
- Every request stamped with the full release identity β prompt, model pin, index, embedding model, tool schema, rubric.
- Eval gates in CI, reported as a delta against production, a regression failing the build.
- Shadow mode on a real traffic slice with side effects suppressed, compared offline on paired inputs.
- Canary with a written rollback trigger, then A/B on an online quality metric before full exposure.
- Ratio-based alerting segmented by release, on quality, schema failures, abstention, escalation, latency, cost, and tool errors.
- Embedding and index migrations run as dual-write plus backfill plus dual-read, never as a flag flip.
- Rollback rehearsed, not assumed β the previous prompt, model pin, and index snapshot all remain deployable together, and someone has actually done it once.
When to Use
β Apply this discipline when:
- Anything LLM-generated reaches users, at any volume
- You depend on a hosted model whose behaviour can change without your involvement
- More than one person can change a prompt, a model choice, or an index
- You are changing the model, the embedding model, or the chunking strategy
- An incident would require you to say which combination produced a specific output
β Do not over-invest when:
- You are on an internal prototype with a handful of consenting users, where version control and a pinned model are the right stopping point
- The feature is a one-off batch job whose output a human reviews before anything acts on it
- You have no evals yet, in which case build those first β a canary with nothing to measure is just a slower way to ship a regression
Common Interview Questions
Q1: What is different about deploying an LLM feature compared to deploying a normal service?
Two assumptions break. First, your behaviour can change with no deploy on your side, because the provider updated a model or a filter β so nothing in a conventional pipeline is watching the thing most likely to change. Second, an identical deploy can behave differently run to run, because sampling is stochastic, so a single bad observation is not automatically a regression and βit worked in testingβ is a weaker claim. The two load-bearing controls that follow are version pinning, so behaviour only moves when I decide it moves, and eval gates, so I can distinguish a real regression from noise. Everything else β shadow, canary, flags β is ordinary release engineering applied on top of those.
Q2: Quality dropped yesterday. Walk me through the investigation.
First I check what the release identity says, because the answer is usually in there: prompt version, model id and pinned snapshot, index version, embedding model version, tool schema version, rubric version. If any moved, that is the prime suspect and I compare scored samples across the two identities. If none moved, the suspects are a provider-side change, a corpus change that arrived through ingest rather than a deploy, or drift in the input distribution. Then I split retrieval from generation so the regression localises to a component, and I read actual failures rather than staring at the aggregate. If I cannot answer the first question at all β which combination produced these outputs β the investigation is already lost, and the fix is instrumentation before anything else.
Q3: How do you roll out a model version change safely?
Pin both the old and new snapshots so the change is explicit, then run the full eval suite against the candidate with everything else held constant. Next shadow mode on a slice of real traffic with side effects suppressed, which gives a paired comparison on the inputs my golden set never contained. Then a canary on a small sticky slice with a written rollback threshold on quality, schema failure rate, abstention, latency, and cost per request, alerting on ratios so the comparison against the stable population is valid. Widen in steps, and if user preference matters, A/B on an online metric like acceptance or escalation before full exposure. Throughout, the previous snapshot stays pinned and deployable so rollback is a flag flip.
Q4: You need to change the embedding model in a RAG system. What is the plan?
It is a migration, not a deploy, because the new model defines a different vector space and old vectors are not comparable with new ones β there is no partially-migrated state that returns meaningful results, so a flag flip is not available. I stand up a second index rather than migrating in place, dual-write from ingest so the new index does not fall behind, and run a resumable rate-limited backfill over the whole corpus, budgeting the embedding cost and the wall-clock time honestly. While it fills I dual-read, serving from the old index and shadowing the new one, comparing retrieval metrics on the same queries and then generation quality on the frozen golden set. Cut over behind a flag once it is complete and measurably at least as good, keep the old index deployable until confidence is high, and rebuild any semantic cache whose stored vectors belong to the old space.
Q5: Should prompt changes go through the deploy pipeline or a live config store?
The pipeline, with a flag selecting which version is active. A prompt is the most behaviour-dense code in the system, so editing it live is shipping unreviewed logic straight to production with no version, no eval run, no trace attribution, and no rollback. Version-controlled prompts get review, a gate, a version id stamped on every request, and a bisectable history. A config store is a reasonable delivery mechanism for the active version pointer β that is what makes rollback a fast flip rather than a pipeline run β but the content itself must be a reviewed, evaluated artifact. The practical test is whether I can name the prompt version that produced any given output from last week; live editing fails that test immediately.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts