AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

A Runbook for Shipping a New Model Version Without Downtime

Swapping in a new model version carries a different kind of risk than a normal code deploy. The failure mode is rarely a crash. It is a model that runs fine but answers worse, drifts on tone, or gets slower under the same load, and none of that shows up in a health check.

The fix is the same discipline good deploy practice already uses for code, adapted for the fact that quality here is fuzzier than a pass or fail test: shadow the new version before it serves anyone, ramp it in behind a small canary, and define your rollback trigger before you start, not after something looks wrong.

Shadowing Before Anything Is User-Facing

Run the new version alongside the current one on a copy of real traffic, without returning its answers to users, before it serves a single live request. This catches the failures that never show up in a staging test: latency under your actual concurrency, memory behavior with your actual prompt lengths, and answers that differ from the current version in ways worth reviewing by hand. Shadowing is cheap compared to what it prevents, since the only cost is running a second copy of inference you throw away, and it is the single biggest gap between teams that migrate model versions calmly and teams that find problems in production.

Choosing a Canary Size That Actually Tells You Something

A canary that is too small will not surface a real problem before it reaches everyone, and a canary that is too large defeats the point of having one. Size it to your traffic: enough real requests in a fixed window to catch a meaningful shift in latency or error rate, small enough that a bad version only affects a limited slice of users while you are watching closely. Hold that slice for long enough to see your actual traffic pattern, not just a quiet hour, before deciding to expand it.

Setting Rollback Triggers Before You Start, Not After

Decide what would make you roll back before the migration begins: a latency threshold, an error rate threshold, or a defined process for reviewing a sample of answers for quality. Writing these down ahead of time removes the moment, common in a live migration, where something looks a little off and nobody wants to be the one to call it. If you need a human review step for answer quality rather than a pure metric, decide who does that review and how quickly, so the canary window is not held open indefinitely waiting on a decision nobody owns.

Running Both Versions During the Cutover

Plan to serve both the old and new version simultaneously for the full length of your canary and ramp, which means your infrastructure needs headroom for two versions running at once, not just the one you are moving toward. This is also the window to keep any caching layer aware of which version produced which cached answer, so a cache hit does not quietly serve an old version's response after you believe the migration is complete. Decommission the old version only once the new one has carried full traffic long enough to prove itself, not the moment the ramp finishes.

What to Check After the Migration Is Done

Once the new version is fully live, confirm three things before calling the migration finished: no caller is still receiving cached answers from the retired version, your monitoring and any audit logging correctly attribute new requests to the new version rather than a stale label, and whoever owns customer communication knows the change happened in case a customer notices a shift in answers and asks about it.

A Worked Example: Moving a Support Assistant to a New Version

Say a support assistant is moving from one model version to a newer release of the same family. Shadowing surfaces that the new version answers refund policy questions slightly differently, close enough that automated checks pass but different enough to matter. Because rollback triggers included a manual review step for a sample of answers, not just latency and error rate, this gets caught during the canary window rather than after a customer complains. The team adjusts the prompt template, reruns the shadow test, and only then expands the canary. Without that manual review trigger, the same issue would likely have reached full traffic before anyone noticed.

The migration sequence in short:

  1. Write down rollback triggers covering latency, error rate, and a sample-based answer quality review before any traffic moves.
  2. Shadow the new version on a copy of real traffic and review the places where its answers differ from the current version.
  3. Send a small canary slice of live traffic, sized to reveal a real shift, and hold it through a full cycle of your usual traffic pattern.
  4. Ramp up gradually while both versions keep running, with capacity for two versions and a cache that knows which version produced each answer.
  5. After full cutover, confirm no cached answers remain from the retired version and that monitoring and audit logs label the new version correctly.
Executive Capability Standard

What Good Looks Like

A new model version is shadowed on real traffic, ramped through a sized canary with predefined rollback triggers, and only retires the prior version after proving itself under full load.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map out your current model deployment process and identify where shadowing or a canary step is missing today.
2. Do Manually:Run a shadow test on a copy of real traffic for your next version change and review the results by hand before any live traffic moves.
3. Delegate:Assign an engineer to own the migration runbook, including rollback triggers, for every model version change going forward.
4. Automate:Automate canary traffic splitting and rollback triggers so a bad version is pulled back without waiting on someone to notice manually.
5. Buy:Bring in fractional platform engineering to build a repeatable shadow-canary-cutover pipeline if version changes are currently ad hoc.

How to Get Started

Frequently Asked Questions

How long should we run a canary before expanding it to full traffic?

Long enough to see your real traffic pattern, not just a quiet period. If your traffic varies meaningfully by time of day or day of week, hold the canary through at least one full cycle of that pattern before deciding to expand it.

What should trigger an automatic rollback versus a manual review?

Clear metric thresholds like latency or error rate are good candidates for automatic rollback. Answer quality is usually fuzzier and needs a human reviewer with authority to call a rollback, decided before the migration starts rather than during it.

Do we need to keep both model versions running the whole time?

Yes, for the length of the canary and ramp. Decommissioning the old version early removes your fallback exactly when you might still need it, and your capacity plan should account for running both versions at once during that window.

What is the most common mistake teams make migrating model versions?

Skipping the shadow phase and finding out about latency or quality problems only once real users are affected. Shadowing on real traffic before anyone sees the new version's answers catches most of what a staging test misses.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides