The Runbook for a Version Migration Nobody Notices
A version migration goes unnoticed when you run it as a sequence of small, reversible steps instead of one big cutover. Run old and new in parallel, agree on rollback triggers in advance, and watch user facing metrics rather than the deploy dashboard. Done that way, users keep working and the only trace is a changelog entry weeks later.
This applies whether you're moving a database to a new major version, upgrading a language runtime across a service fleet, or migrating an API to a new contract version.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How do you run old and new versions in parallel?
The core technique behind a zero downtime migration is running both versions side by side long enough to prove the new one behaves correctly under real traffic, before the old version goes away. For a database, that means dual writes to old and new schemas with reads still coming from the old one. For a service upgrade, that means a canary population on the new version while the majority stays on the old one. The point isn't speed, it's having a safe, cheap way to compare behavior before you're committed.
Skipping this step to save a week is how a subtle behavior change in the new version reaches every user at once instead of the small canary group you were watching closely.
Make every step reversible on its own
Break the migration into steps small enough that any single one can be rolled back without touching the others. Add the new schema or new code path without removing the old one. Start writing to both, still reading from the old one. Switch reads to the new path once writes have been verified for a full traffic cycle, typically at least one business week so you catch anything that only happens on a specific weekday or during a batch job. Only remove the old path once the new one has run alone, in production, long enough that rolling back would mean redoing work rather than flipping a flag.
If a step can't be undone in minutes, it's not a step, it's the whole migration disguised as one.
A reversible migration usually moves in this order:
- Add the new schema or code path alongside the old one, without removing anything from the old path.
- Start writing to both paths while reads still come from the old one, so nothing depends on the new path yet.
- Verify the new writes across a full traffic cycle before you trust them for reads.
- Switch reads to the new path once those writes are confirmed correct, keeping the old path available for rollback.
- Remove the old path only after giving it a real expiration date and an owner.
When should you decide the rollback trigger?
Write down, before the migration begins, exactly what error rate, latency threshold, or specific failure would trigger a rollback, and who has the authority to call it. Deciding this during an active incident means someone is doing math under pressure instead of checking a number against a threshold that was already agreed. The trigger should be conservative early in the rollout, when you're watching a small percentage of traffic, and can loosen slightly once more of the fleet has proven stable on the new version.
This single document, agreed before the first line of the migration runs, is usually the difference between a rollback that takes five minutes and one that takes an hour of debate first.
Watch the metric that actually reflects users, not just the deploy
A deploy dashboard turning green tells you the new version started. It doesn't tell you it's correct. Watch application level signals, error rate on the specific endpoints touched by the migration, latency at the percentile that matters for your product, and any business metric downstream of the migrated component, like checkout completion if you migrated the payments service. Deployment frequency and lead time matter for how fast you can iterate, but they don't tell you whether this specific migration behaved: the highest performing engineering organizations manage on-demand deployment, essentially daily releases1, and that speed is only worth having if what ships is actually correct.
Set the dashboard up before you start, not after something looks wrong.
Give the old path a real expiration date
The most common failure after a successful migration isn't technical, it's that the old code path never actually gets removed. It sits there, unmaintained, until the next security patch or dependency bump forces someone to touch it again, at which point nobody remembers why it's still there. Set a calendar date for removing the old path when you start the migration, not a vague someday, and treat that removal as part of the migration's definition of done rather than a separate future project.
A migration that's mostly done and stuck there for a year is worse than one that hasn't started, because it looks finished on every dashboard that matters except the one that shows dead code.
Communicate the migration like a change, not a secret
Even an invisible migration touches other teams: support needs to know a behavior might shift slightly during the parallel run, and whoever owns the on call rotation needs to know a rollback trigger exists and what it looks like when it fires. Write a short, plain description of what's changing, when, and how to tell if something's wrong, and share it before the migration starts rather than after someone notices a symptom and has to guess whether it's related. This costs fifteen minutes and saves the alternative: a support ticket escalated as a mystery incident that turns out to be the migration working exactly as planned.
The goal isn't a big announcement, it's making sure nobody outside engineering is caught off guard by something that was, by design, supposed to be unremarkable.
What Good Looks Like
A good version migration runs old and new in parallel, breaks the cutover into steps that are each independently reversible, and has a written rollback trigger agreed before it starts.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How long should we run old and new in parallel?
Long enough to see a full business cycle, at minimum one week, so you catch behavior that only shows up on a specific weekday, during a batch job, or at a monthly boundary. For anything touching billing or reporting, run through at least one full billing cycle before removing the old path, since that's often where subtle discrepancies actually surface.
What's the biggest mistake teams make on these migrations?
Treating the cutover as the finish line instead of the midpoint. Teams plan the switch carefully and then rush or skip decommissioning the old path, which is exactly the part that causes problems months later when nobody remembers it's still there. Put decommissioning on the same plan and timeline as the cutover itself, with its own owner and deadline.
Do we need a maintenance window for a well planned migration?
For most application level and database version migrations, no, that's the entire point of running old and new in parallel with reversible steps. Some infrastructure level changes, like certain storage engine swaps, genuinely can't avoid a brief window. Decide this explicitly during planning rather than defaulting to a window because that's how migrations were always done before.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
A Runbook for Shipping Breaking API Changes Without Downtime
A step-by-step approach to shipping a breaking API or schema change without a maintenance window, built around parallel versions.
A Runbook for Zero-Downtime Schema Migrations on a Live Database
A step-by-step runbook for running schema migrations against a production database without an outage window, including the rollback checkpoints.
Shipping API Version Migrations Without a Maintenance Window
A step-by-step approach to migrating API versions and running database or schema changes without a maintenance window or breaking existing clients.
Build or Buy for Verifying Every Device That Connects?
How to split device identity from device posture checking, what building either one in house actually costs, and where a platform earns its keep instead.
Running Schema and Version Migrations Without an Outage
A step-by-step approach to running database schema and version migrations without downtime, including the rollback decision most teams put off.
A Runbook for Version Migrations Your Customers Never Notice
The sequencing that keeps a version migration from becoming an outage: compatibility windows, rollout order, and what to check before you remove the old path.