Running Schema and Version Migrations Without an Outage
Most migration outages don't come from a bad idea, they come from a good idea run as one big irreversible step. The schema change was correct. The new code path was correct. The problem was doing both at once, with no way back once something started going wrong.
This runbook walks through the sequence that avoids that: separate the schema change from the code deploy, make every step reversible, and decide your rollback trigger before you need it.
How do you separate a schema change from the code deploy?
Add the new column, table, or field first, in a deploy that doesn't depend on it yet. Backfill the data next, on its own schedule, so a slow backfill never blocks a release. Only after both of those are done and verified does the code that reads or writes the new shape go out.
This expand-then-contract sequence turns one risky migration into three smaller, independently reversible steps, each of which is easy to verify before moving to the next.
A common shortcut is skipping straight from adding the new field to switching the application over to it, on the theory that the backfill will finish in time. It usually does, until a slow query or a locked table stalls it mid-run, and now the new code is reading a field that's only partially populated.
Make Every Step Reversible Before You Run It Forward
Add columns instead of renaming them, and keep the old field readable until every reader, including any batch jobs, exports, or downstream services, has moved off it. A rename can't be undone cleanly once something has already written to the new name; an addition can simply be ignored if you need to back out.
Write the backfill so it can be re-run safely if it's interrupted partway through, rather than assuming it will complete in one uninterrupted pass.
Roll Out Behind a Flag, Not a Big-Bang Release
Ship the code that reads the new shape behind a flag while the write path still populates both old and new fields. Flip a small slice of traffic first, watch for errors, and only remove the old path once everything downstream has been confirmed to expect the new one.
How small a step this needs to be tracks your deploy frequency: a team that can release on demand keeps each migration step tiny, while a team stuck releasing only once every 30 days or so ends up bundling bigger, riskier changes into each release by default, simply because that's the next chance they'll get1.
If your team can't yet deploy that often, the flag-based approach still helps: it lets you decouple turning on the new behavior from the less frequent code releases, so you're not stuck choosing between shipping the migration and waiting a month for the next window.
Watch the Signals That Actually Catch Problems
Error rate and write latency on the affected tables tell you more during a cutover than raw CPU or memory ever will. A migration can look perfectly healthy on infrastructure dashboards while quietly writing data in the wrong shape, because CPU doesn't know the difference between a correct write and a corrupt one.
A small reconciliation check, comparing old and new fields for a sample of rows, catches that kind of silent corruption before customers do, and it's worth writing before the migration starts, not after something looks wrong.
How do you decide the rollback trigger before you start?
Pick a specific, measurable trigger in advance: a defined error rate threshold, or a data integrity check failing on more than a small sample. Write it down before the migration begins.
A rollback decision made mid-incident, by a stressed team staring at ambiguous dashboards, is where good migrations turn into bad outages. Deciding the trigger in advance removes that judgment call from the moment when judgment is least reliable.
Make sure the rollback itself has actually been tested, not just written down. A rollback plan that has never been exercised outside of a real incident is a guess with good intentions, and the first time you run it shouldn't be the first time you find out it doesn't quite work.
A safe migration run follows this order:
- Add the new column, table, or field in a deploy that does not depend on it yet.
- Backfill the data on its own schedule, written so it can be re-run safely if it is interrupted.
- Ship the code that reads the new shape behind a flag while writes still populate both the old and new fields.
- Move a small slice of traffic first, watching error rate and write latency on the affected tables plus a reconciliation check of old and new fields.
- Remove the old path only after every reader has moved off it, and roll back if the trigger you wrote down in advance fires.
What Good Looks Like
Good migration practice means schema changes, data backfills, and code changes ship as separate, individually reversible steps behind a flag, with a rollback trigger defined before the migration starts rather than decided mid-incident.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Can we run a zero-downtime migration without feature flags?
It's possible for very simple, additive changes, but flags are what let you separate deploying new code from turning on new behavior. Without that separation, you're back to a big-bang release where the schema change and the code change succeed or fail together, which is exactly the risk this approach is meant to avoid.
What's the most common mistake teams make with migrations?
Renaming a column or table in place instead of adding the new shape alongside the old one. A rename can't be reversed cleanly once anything has written to the new name, while an addition can simply be left unused if something goes wrong, which is why expand-then-contract avoids renames almost entirely.
How long should we keep the old schema around after a migration?
Until every reader of the old field, including batch jobs, exports, and any downstream service, has been confirmed to use the new one instead. That's rarely as fast as engineers expect, since forgotten batch jobs and reports are usually the last consumers anyone checks.
What should actually trigger an automatic rollback during a migration?
A specific, pre-defined signal: a fixed error rate threshold being crossed, or a data integrity check failing on more than a small sample of rows. Deciding this before the migration starts matters more than the exact numbers, because it removes the decision from the middle of an incident.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
A Runbook for Shipping Breaking API Changes Without Downtime
A step-by-step approach to shipping a breaking API or schema change without a maintenance window, built around parallel versions.
A Runbook for Zero-Downtime Schema Migrations on a Live Database
A step-by-step runbook for running schema migrations against a production database without an outage window, including the rollback checkpoints.
Shipping API Version Migrations Without a Maintenance Window
A step-by-step approach to migrating API versions and running database or schema changes without a maintenance window or breaking existing clients.
The Runbook for a Version Migration Nobody Notices
A step by step approach to migrating a service or database to a new major version without a maintenance window, and what to check before you start.
A Runbook for Version Migrations Your Customers Never Notice
The sequencing that keeps a version migration from becoming an outage: compatibility windows, rollout order, and what to check before you remove the old path.