A Runbook for Version Migrations Your Customers Never Notice
Most "zero-downtime" migrations fail for the same reason: something during the rollout assumes only one version is live at a time, when in practice both versions are live simultaneously for the whole rollout window. Fixing that assumption, not finding a cleverer deployment tool, is what actually gets you through a migration cleanly.
Assume both versions run at once, because they will
During any rolling deployment, old and new code serve traffic at the same time, sometimes for minutes and sometimes for hours if the rollout is gradual or gets paused. A schema change, an API contract change, or a message format change that only one of the two versions understands will break requests during that overlap window even if both versions work fine in isolation. Design every migration around the overlap, not around the moment before or after it.
How do you sequence database changes to be additive?
Add new columns or tables before any code reads or writes them. Deploy code that writes to both the old and new structure while only reading from the old one. Backfill historical data. Deploy code that reads from the new structure. Only after all of that is confirmed working do you remove the old structure. Skipping steps to save time is exactly how a migration turns into an incident, because it collapses the window in which you can safely roll back.
Follow this order for a database change:
- Add new columns or tables before any code reads or writes them.
- Deploy code that writes to both the old and new structure while still reading only from the old one.
- Backfill all historical data into the new structure.
- Deploy code that reads from the new structure.
- Remove the old structure only after everything above is confirmed working.
Version your API contracts, not just your code
If a message format, event schema, or API response shape is changing, give it an explicit version identifier the consumer checks, rather than relying on every consumer being redeployed at the same instant. This matters even more for internal services and background workers than for public APIs, because internal consumers are exactly the ones people forget still exist. Say a queue consumer written two years ago is still parsing the old event shape; an unversioned schema change breaks it silently.
Keep a rollback path that doesn't depend on the new path working
A rollback plan that assumes the new version's write path is available to reverse is not a real rollback plan, because the reason you are rolling back is often that the new path is broken. Keep the old read and write path functional and reachable until you have confidence in the new one across a full traffic cycle, including your highest-traffic period, not just a quiet afternoon. Practice the rollback itself before you need it for real, the same way you would practice a restore from backup, because a rollback procedure nobody has ever actually executed tends to reveal its missing steps exactly when you can least afford to discover them.
A useful test of a rollback plan is to ask what has to be true for it to work. For example, if reversing a release requires the new write path to accept a reverse operation, and the new path is the thing that broke, the plan fails exactly when needed. Write the rollback as steps that use only the old, known-good path, then run them in a staging environment before the migration begins. A common mistake is treating the rollback as documentation instead of a rehearsed procedure. Rehearsal turns missing steps into a quick fix on a quiet afternoon instead of a surprise during an incident.
When should you remove the old code path?
The most common failure after a successful migration is that the old path never actually gets removed, because removing it feels riskier than leaving it. That leftover path becomes technical debt and, eventually, a security or compliance gap nobody remembers exists, since old code paths tend to stop receiving the same scrutiny in reviews once attention has moved on to whatever shipped next. Put a specific removal date on the calendar when you start the migration, tied to a condition (a full week of clean metrics on the new path, for instance) rather than leaving it open-ended, and treat missing that date as a signal to ask why, not as something to quietly push back again.
Communicate the migration window to everyone who depends on it
A migration that is technically well-sequenced can still cause confusion if other teams don't know it's happening. Tell downstream consumers, support teams, and anyone who might be debugging an unrelated issue during the window that both versions are live and what the expected differences are. A support engineer who doesn't know a migration is in progress can spend hours chasing a ghost, treating an expected transitional behavior as a new bug, when a two-line notice beforehand would have saved that time entirely.
What Good Looks Like
A safe version migration sequences database and contract changes to be additive first, runs old and new versions in parallel through a full traffic cycle, keeps a rollback path that does not depend on the new version working, and has an explicit, condition-based date for removing the old path.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How long should we run both versions in parallel?
Long enough to see a full traffic cycle on the new path, including your highest-traffic period, with no elevated error rate or rollback need. For most services that is at least one full week, and longer for anything with monthly or seasonal traffic patterns you would otherwise miss.
What's the biggest cause of migrations that don't actually run without downtime?
Treating the deployment as instantaneous instead of as a window where old and new code both handle live traffic. Any change that isn't safe for both versions to make at once, such as a non-additive schema change, will surface exactly during that window.
Do internal services need the same version discipline as public APIs?
Yes, and often more, because an internal consumer that nobody remembers still exists is more likely to be running old code with no one watching it closely. Treat internal event and message schemas as versioned contracts, not implementation details.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
A Runbook for Shipping Breaking API Changes Without Downtime
A step-by-step approach to shipping a breaking API or schema change without a maintenance window, built around parallel versions.
A Runbook for Shipping a New Model Version Without Downtime
A step-by-step way to roll a new model version into production: shadow traffic first, a small canary, clear rollback triggers, and a real cutover.
A Runbook for Zero-Downtime Schema Migrations on a Live Database
A step-by-step runbook for running schema migrations against a production database without an outage window, including the rollback checkpoints.
Shipping API Version Migrations Without a Maintenance Window
A step-by-step approach to migrating API versions and running database or schema changes without a maintenance window or breaking existing clients.
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.