A Runbook for Shipping Breaking API Changes Without Downtime
Zero-downtime migrations aren't a single trick, they're a sequence: you run the old and new versions side by side long enough to move every caller, then retire the old one. Skip a step and you either break a client mid-request or you carry two versions forever because no one finishes the cutover.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Step 1: Separate the Schema Change From the Behavior Change
The most common mistake is shipping a new column and a new code path that depends on it in the same deploy. Split it: add the column first, deploy code that writes to both old and new fields, backfill existing rows, then only after that's confirmed complete, deploy code that reads from the new field.
Each step is independently safe to roll back, because at every point along the way the old code path still works against the current state of the data. If something looks wrong after step two, you roll back to step one without losing writes; if you'd combined the steps, a rollback could mean losing data written under the new assumption.
Step 2: Version the API Contract, Not Just the Code
If external or internal clients call your API directly, a breaking change needs its own version in the URL or header, served alongside the old version, not a flag that flips for everyone at once. This is the difference between a rollout you control and an outage a client discovers first.
Keep the old version's behavior frozen once the new one ships. It's tempting to patch small things in the old version while both are live, but every change you make there is a change you have to verify twice, once for each version, for as long as both stay in service.
Step 3: Track Who's Still on the Old Version
You can't retire a version safely if you don't know who's still calling it. Log the version on every request and build a simple dashboard of callers still on the deprecated path before you set a sunset date.
This step is where most migrations stall, not because it's technically hard, but because no one owns chasing down the last few callers. Name an owner for that outreach explicitly, the same way you'd name an owner for the code itself, or the migration sits at ninety percent done indefinitely.
Step 4: Rehearse the Cutover Against Your Deploy Cadence
How you sequence a migration depends on how often you ship. A team that deploys on demand can run each step as its own small release and back out quickly if something looks wrong; a team on a slower release cadence needs to bundle steps more carefully because each release is a bigger commitment, and deployment frequency is a real constraint on how you plan a migration, not just a metric to look good on1.
Know which team you are before you write the rollout schedule. If your releases are infrequent, consider whether a temporary increase in deploy cadence, just for the migration window, is worth the coordination cost, because smaller, more frequent steps are genuinely safer than one large one.
Step 5: Set a Hard Sunset Date and Enforce It
Old versions that never get a real sunset date become permanent maintenance burden. Once the caller dashboard shows the old version at zero or near-zero for a full week, set a date, notify anyone left, and remove the old code path.
Leaving it "just in case" for months is how a two-version migration turns into a three-version mess, where every future change has to be tested against paths nobody actually uses anymore. Removing dead code is part of the migration, not an optional cleanup step you get to later.
What to Do When a Migration Goes Wrong Mid-Rollout
Even with all five steps followed, something can still surface only under real production load. The rollback plan for each step matters more than the forward plan, because a rollback under pressure is not the moment to be improvising which flag to flip or which deploy to revert.
Write the rollback command or the specific revert steps down before you start the migration, not after something breaks. A migration plan without a written rollback for each step isn't actually a plan, it's an optimistic forward-only script.
The full sequence, in order:
- Split the schema change from the behavior change: add the new column, write to both fields, backfill, and only then read from the new one.
- Give the breaking change its own API version in the URL or header and serve it alongside the old version.
- Log the version on every request and watch a dashboard of callers still on the deprecated path.
- Rehearse the cutover against your deploy cadence, with a written rollback plan for each step.
- Set a hard sunset date once the old version sits at or near zero for a full week, then remove the old code path.
What Good Looks Like
A safe migration separates schema changes from behavior changes, tracks who's still on the old path, and sets a real sunset date instead of running two versions indefinitely.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How long should we run two API versions side by side?
Long enough for every real caller to migrate, which you can only know by tracking version usage per request. For an internal API that might be a week; for an external API with third-party integrators it's often a month or more. Set the sunset date from actual usage data, not a fixed calendar guess.
What's the biggest risk in a database migration done without downtime?
Reading from a field before the backfill that populates it has actually finished. Always confirm the backfill is complete and verified before deploying any code that depends on the new data being present for every row.
Do we need feature flags for this, or is API versioning enough?
They solve different problems. API versioning controls what contract a client sees; feature flags control what code path runs internally. Most zero-downtime migrations use both: a flag to control the internal rollout, and a version to control what external clients experience.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
A Runbook for Zero-Downtime Schema Migrations on a Live Database
A step-by-step runbook for running schema migrations against a production database without an outage window, including the rollback checkpoints.
Shipping API Version Migrations Without a Maintenance Window
A step-by-step approach to migrating API versions and running database or schema changes without a maintenance window or breaking existing clients.
How to Upgrade a Major Dependency Without a Maintenance Window
Zero-downtime version migrations depend on running two versions in production at once, not a well-timed maintenance window. Here is the pattern that works.
The Runbook for a Version Migration Nobody Notices
A step by step approach to migrating a service or database to a new major version without a maintenance window, and what to check before you start.
Running a Schema Migration on a Live Event Pipeline
A practical runbook for changing a live event pipeline's schema or message format without dropping data or breaking downstream consumers.
Running Schema and Version Migrations Without an Outage
A step-by-step approach to running database schema and version migrations without downtime, including the rollback decision most teams put off.