A Runbook for Shipping a New Model Version Without Downtime
Swapping in a new model version carries a different kind of risk than a normal code deploy. The failure mode is rarely a crash. It is a model that runs fine but answers worse, drifts on tone, or gets slower under the same load, and none of that shows up in a health check.
The fix is the same discipline good deploy practice already uses for code, adapted for the fact that quality here is fuzzier than a pass or fail test: shadow the new version before it serves anyone, ramp it in behind a small canary, and define your rollback trigger before you start, not after something looks wrong.
Shadowing Before Anything Is User-Facing
Run the new version alongside the current one on a copy of real traffic, without returning its answers to users, before it serves a single live request. This catches the failures that never show up in a staging test: latency under your actual concurrency, memory behavior with your actual prompt lengths, and answers that differ from the current version in ways worth reviewing by hand. Shadowing is cheap compared to what it prevents, since the only cost is running a second copy of inference you throw away, and it is the single biggest gap between teams that migrate model versions calmly and teams that find problems in production.
Choosing a Canary Size That Actually Tells You Something
A canary that is too small will not surface a real problem before it reaches everyone, and a canary that is too large defeats the point of having one. Size it to your traffic: enough real requests in a fixed window to catch a meaningful shift in latency or error rate, small enough that a bad version only affects a limited slice of users while you are watching closely. Hold that slice for long enough to see your actual traffic pattern, not just a quiet hour, before deciding to expand it.
Setting Rollback Triggers Before You Start, Not After
Decide what would make you roll back before the migration begins: a latency threshold, an error rate threshold, or a defined process for reviewing a sample of answers for quality. Writing these down ahead of time removes the moment, common in a live migration, where something looks a little off and nobody wants to be the one to call it. If you need a human review step for answer quality rather than a pure metric, decide who does that review and how quickly, so the canary window is not held open indefinitely waiting on a decision nobody owns.
Running Both Versions During the Cutover
Plan to serve both the old and new version simultaneously for the full length of your canary and ramp, which means your infrastructure needs headroom for two versions running at once, not just the one you are moving toward. This is also the window to keep any caching layer aware of which version produced which cached answer, so a cache hit does not quietly serve an old version's response after you believe the migration is complete. Decommission the old version only once the new one has carried full traffic long enough to prove itself, not the moment the ramp finishes.
What to Check After the Migration Is Done
Once the new version is fully live, confirm three things before calling the migration finished: no caller is still receiving cached answers from the retired version, your monitoring and any audit logging correctly attribute new requests to the new version rather than a stale label, and whoever owns customer communication knows the change happened in case a customer notices a shift in answers and asks about it.
A Worked Example: Moving a Support Assistant to a New Version
Say a support assistant is moving from one model version to a newer release of the same family. Shadowing surfaces that the new version answers refund policy questions slightly differently, close enough that automated checks pass but different enough to matter. Because rollback triggers included a manual review step for a sample of answers, not just latency and error rate, this gets caught during the canary window rather than after a customer complains. The team adjusts the prompt template, reruns the shadow test, and only then expands the canary. Without that manual review trigger, the same issue would likely have reached full traffic before anyone noticed.
The migration sequence in short:
- Write down rollback triggers covering latency, error rate, and a sample-based answer quality review before any traffic moves.
- Shadow the new version on a copy of real traffic and review the places where its answers differ from the current version.
- Send a small canary slice of live traffic, sized to reveal a real shift, and hold it through a full cycle of your usual traffic pattern.
- Ramp up gradually while both versions keep running, with capacity for two versions and a cache that knows which version produced each answer.
- After full cutover, confirm no cached answers remain from the retired version and that monitoring and audit logs label the new version correctly.
What Good Looks Like
A new model version is shadowed on real traffic, ramped through a sized canary with predefined rollback triggers, and only retires the prior version after proving itself under full load.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How long should we run a canary before expanding it to full traffic?
Long enough to see your real traffic pattern, not just a quiet period. If your traffic varies meaningfully by time of day or day of week, hold the canary through at least one full cycle of that pattern before deciding to expand it.
What should trigger an automatic rollback versus a manual review?
Clear metric thresholds like latency or error rate are good candidates for automatic rollback. Answer quality is usually fuzzier and needs a human reviewer with authority to call a rollback, decided before the migration starts rather than during it.
Do we need to keep both model versions running the whole time?
Yes, for the length of the canary and ramp. Decommissioning the old version early removes your fallback exactly when you might still need it, and your capacity plan should account for running both versions at once during that window.
What is the most common mistake teams make migrating model versions?
Skipping the shadow phase and finding out about latency or quality problems only once real users are affected. Shadowing on real traffic before anyone sees the new version's answers catches most of what a staging test misses.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Zero Trust for Machines Calling Your Model Endpoints
Why internal services calling your inference endpoints still need identity checks, and where CrowdStrike-style posture checks and Tenable-style scanning fit.
A Rollout Checklist for Swapping Models in Production
A rollout checklist for swapping AI models in production: evaluation gates, canary and shadow traffic, and fast rollback paths.
Why Inference Latency Creeps Up After You Ship
Where AI model-serving latency actually hides: tokenization, queueing, batching windows, and network hops, plus a worked example fix.
A Runbook for Version Migrations Your Customers Never Notice
The sequencing that keeps a version migration from becoming an outage: compatibility windows, rollout order, and what to check before you remove the old path.
A Runbook for Shipping Breaking API Changes Without Downtime
A step-by-step approach to shipping a breaking API or schema change without a maintenance window, built around parallel versions.
A Runbook for Zero-Downtime Schema Migrations on a Live Database
A step-by-step runbook for running schema migrations against a production database without an outage window, including the rollback checkpoints.