Cloud FinOps & Infrastructure ScalingPlaybook4 min readUpdated September 2026

The Runbook for a Version Migration Nobody Notices

A version migration goes unnoticed when you run it as a sequence of small, reversible steps instead of one big cutover. Run old and new in parallel, agree on rollback triggers in advance, and watch user facing metrics rather than the deploy dashboard. Done that way, users keep working and the only trace is a changelog entry weeks later.

This applies whether you're moving a database to a new major version, upgrading a language runtime across a service fleet, or migrating an API to a new contract version.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you run old and new versions in parallel?

The core technique behind a zero downtime migration is running both versions side by side long enough to prove the new one behaves correctly under real traffic, before the old version goes away. For a database, that means dual writes to old and new schemas with reads still coming from the old one. For a service upgrade, that means a canary population on the new version while the majority stays on the old one. The point isn't speed, it's having a safe, cheap way to compare behavior before you're committed.

Skipping this step to save a week is how a subtle behavior change in the new version reaches every user at once instead of the small canary group you were watching closely.

Make every step reversible on its own

Break the migration into steps small enough that any single one can be rolled back without touching the others. Add the new schema or new code path without removing the old one. Start writing to both, still reading from the old one. Switch reads to the new path once writes have been verified for a full traffic cycle, typically at least one business week so you catch anything that only happens on a specific weekday or during a batch job. Only remove the old path once the new one has run alone, in production, long enough that rolling back would mean redoing work rather than flipping a flag.

If a step can't be undone in minutes, it's not a step, it's the whole migration disguised as one.

A reversible migration usually moves in this order:

  1. Add the new schema or code path alongside the old one, without removing anything from the old path.
  2. Start writing to both paths while reads still come from the old one, so nothing depends on the new path yet.
  3. Verify the new writes across a full traffic cycle before you trust them for reads.
  4. Switch reads to the new path once those writes are confirmed correct, keeping the old path available for rollback.
  5. Remove the old path only after giving it a real expiration date and an owner.

When should you decide the rollback trigger?

Write down, before the migration begins, exactly what error rate, latency threshold, or specific failure would trigger a rollback, and who has the authority to call it. Deciding this during an active incident means someone is doing math under pressure instead of checking a number against a threshold that was already agreed. The trigger should be conservative early in the rollout, when you're watching a small percentage of traffic, and can loosen slightly once more of the fleet has proven stable on the new version.

This single document, agreed before the first line of the migration runs, is usually the difference between a rollback that takes five minutes and one that takes an hour of debate first.

Watch the metric that actually reflects users, not just the deploy

A deploy dashboard turning green tells you the new version started. It doesn't tell you it's correct. Watch application level signals, error rate on the specific endpoints touched by the migration, latency at the percentile that matters for your product, and any business metric downstream of the migrated component, like checkout completion if you migrated the payments service. Deployment frequency and lead time matter for how fast you can iterate, but they don't tell you whether this specific migration behaved: the highest performing engineering organizations manage on-demand deployment, essentially daily releases1, and that speed is only worth having if what ships is actually correct.

Set the dashboard up before you start, not after something looks wrong.

Give the old path a real expiration date

The most common failure after a successful migration isn't technical, it's that the old code path never actually gets removed. It sits there, unmaintained, until the next security patch or dependency bump forces someone to touch it again, at which point nobody remembers why it's still there. Set a calendar date for removing the old path when you start the migration, not a vague someday, and treat that removal as part of the migration's definition of done rather than a separate future project.

A migration that's mostly done and stuck there for a year is worse than one that hasn't started, because it looks finished on every dashboard that matters except the one that shows dead code.

Communicate the migration like a change, not a secret

Even an invisible migration touches other teams: support needs to know a behavior might shift slightly during the parallel run, and whoever owns the on call rotation needs to know a rollback trigger exists and what it looks like when it fires. Write a short, plain description of what's changing, when, and how to tell if something's wrong, and share it before the migration starts rather than after someone notices a symptom and has to guess whether it's related. This costs fifteen minutes and saves the alternative: a support ticket escalated as a mystery incident that turns out to be the migration working exactly as planned.

The goal isn't a big announcement, it's making sure nobody outside engineering is caught off guard by something that was, by design, supposed to be unremarkable.

Executive Capability Standard

What Good Looks Like

A good version migration runs old and new in parallel, breaks the cutover into steps that are each independently reversible, and has a written rollback trigger agreed before it starts.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read through the last migration your team ran and note where it deviated from a parallel run, reversible steps, or a written rollback trigger.
2. Do Manually:Plan the next migration on paper first, listing each step and whether it's reversible on its own, before writing any migration code.
3. Delegate:Assign a specific engineer to own the rollback decision during the migration window, separate from whoever is executing the steps, so judgment isn't rushed.
4. Automate:Automate the traffic split between old and new versions so the percentage can be adjusted without a deploy, rather than hardcoding the canary size.
5. Buy:Bring in outside infrastructure expertise for a migration touching your primary datastore if nobody on the team has run one at this scale before.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

ClickUp

tracking each migration step, its owner, and its rollback trigger in a tool like ClickUp keeps the plan visible instead of living only in one engineer's head

Visit ClickUp→

Frequently Asked Questions

How long should we run old and new in parallel?

Long enough to see a full business cycle, at minimum one week, so you catch behavior that only shows up on a specific weekday, during a batch job, or at a monthly boundary. For anything touching billing or reporting, run through at least one full billing cycle before removing the old path, since that's often where subtle discrepancies actually surface.

What's the biggest mistake teams make on these migrations?

Treating the cutover as the finish line instead of the midpoint. Teams plan the switch carefully and then rush or skip decommissioning the old path, which is exactly the part that causes problems months later when nobody remembers it's still there. Put decommissioning on the same plan and timeline as the cutover itself, with its own owner and deadline.

Do we need a maintenance window for a well planned migration?

For most application level and database version migrations, no, that's the entire point of running old and new in parallel with reversible steps. Some infrastructure level changes, like certain storage engine swaps, genuinely can't avoid a brief window. Decide this explicitly during planning rather than defaulting to a window because that's how migrations were always done before.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides