A Rollout Checklist for Swapping Models in Production
Deploying a new code build and deploying a new model version aren't the same operation, even though most CI/CD pipelines treat them identically. A model swap changes your product's actual behavior, not just its implementation, and a bug can look like a quality regression instead of a crash. That's harder to catch in a health check.
If your rollout playbook for production deployment is copied from your regular app deploys, it's missing the checks that matter most for a model change.
Why model deploys need a different playbook than code deploys
A broken code deploy usually throws errors, fails a health check, and gets rolled back automatically within minutes. A broken model deploy often returns valid-looking responses that are just worse: less accurate, more prone to a specific failure mode, or subtly off for one segment of traffic.
That means your standard health check, does the endpoint return a success code, isn't enough. You need an evaluation gate that runs before traffic increases, checking the new model against a fixed test set and comparing its outputs to the version it's replacing. A model that passes every infrastructure check and still fails your evaluation gate should block the rollout the same way a failed test suite would.
Canary, shadow, and blue-green for a model swap
Three rollout patterns work well for models, for different reasons:
- Shadow traffic: send a copy of production requests to the new model without serving its responses, so you can compare outputs before anyone is affected.
- Canary: serve the new model to a small, fixed slice of real traffic and watch both infrastructure metrics and output quality before expanding it.
- Blue-green: run both versions fully in parallel and switch traffic at the load balancer, which gives you an instant rollback but doubles your GPU cost during the cutover.
Most teams should default to shadow traffic first, then canary. Blue-green is worth the extra cost mainly for a model change large enough that you don't trust a canary sample to catch the risk.
What to check before you flip traffic
Beyond the evaluation gate, confirm four things before a model swap goes past canary: the new model's latency profile under real load, not just a benchmark script; that your prompt templates and any tool-calling schema still match what the new model expects; that logging and tracing tag requests with the model version, so a regression is traceable after the fact; and that you have a fast rollback path that doesn't require a full redeploy, usually a router flag or a weighted traffic split you can flip in seconds.
Skipping the schema check is the one that bites teams most often: a new model version can change how it formats tool calls or structured output even when the prompt hasn't changed.
What deploy frequency has to do with any of this
Teams that deploy often tend to deploy safer, not despite the frequency but because of it: smaller changes are easier to evaluate and easier to roll back. DORA's research clusters engineering organizations by deploy frequency, from top-performing teams deploying on demand down to the slowest cluster deploying between once a month and twice a year1.
Model deploys are a good place to apply the same logic. A smaller, more frequent model update is easier to gate and easier to clear when something goes wrong than a rare, large one that changes several things about the model's behavior at once.
Pitfalls that turn a routine swap into an incident
- Rolling back the model but forgetting to roll back the prompt template that was tuned for its quirks, which breaks the old model too.
- Running the evaluation gate against a stale test set that doesn't reflect current traffic patterns.
- Treating a fine-tune of the same base model as low-risk by default; a fine-tune can shift behavior more than a version bump.
- Skipping shadow traffic for a small change, which is exactly the kind of change most likely to slip through a canary sample unnoticed.
The common thread: every one of these skips a check that felt optional under deadline pressure.
What Good Looks Like
A production-ready deploy process for model swaps means every rollout passes an evaluation gate before traffic increases, runs through shadow or canary traffic first, and can be rolled back with a config change rather than a redeploy.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need a full evaluation suite before every model deploy, even small ones?
You need some evaluation gate before every deploy, but its size can scale with the change. A prompt tweak might only need a handful of regression cases checked automatically. A new base model or a retrain deserves a fuller test set and a canary period. What you shouldn't do is skip the gate entirely because a change looks small; that's usually when a regression slips through.
What's the fastest way to roll back a bad model deploy?
A weighted traffic split at your router or load balancer that you can flip without a redeploy. If your only rollback path is redeploying the previous container image, you've added minutes of exposure to every incident. Keep the previous model version warm and reachable so a rollback is a config change, not a deploy.
How do we know if a model regression is from the model itself or something else in the pipeline?
Tag every request and log entry with the model version that handled it, along with prompt template version and any retrieval or tool results involved. When a regression shows up, filter by model version first. If the same version behaves differently across two time windows, look at what changed around it, not the model.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Build or Buy: Canary Deployments for a Small Team
What a canary release actually needs to catch problems, how far you can get with a load balancer alone, and when a managed platform earns its keep.
Why Inference Latency Creeps Up After You Ship
Where AI model-serving latency actually hides: tokenization, queueing, batching windows, and network hops, plus a worked example fix.
Zero Trust for Machines Calling Your Model Endpoints
Why internal services calling your inference endpoints still need identity checks, and where CrowdStrike-style posture checks and Tenable-style scanning fit.
What a Real Security Audit of Model Serving Should Cover
A practical checklist for auditing AI model serving and inference: endpoint access, weight security, prompt logging, and patch timelines.
What SOC 2 Actually Expects From a Model-Serving Team
What SOC 2 expects from a team serving AI models: how change, access, patch, and vendor controls apply, and the evidence to have ready.
Catching Broken Tool-Calling Schemas Before They Reach Production
How to build contract tests for AI model serving that catch schema and tool-calling drift, including provider-side changes.