Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Running Schema and Version Migrations Without an Outage

Most migration outages don't come from a bad idea, they come from a good idea run as one big irreversible step. The schema change was correct. The new code path was correct. The problem was doing both at once, with no way back once something started going wrong.

This runbook walks through the sequence that avoids that: separate the schema change from the code deploy, make every step reversible, and decide your rollback trigger before you need it.

How do you separate a schema change from the code deploy?

Add the new column, table, or field first, in a deploy that doesn't depend on it yet. Backfill the data next, on its own schedule, so a slow backfill never blocks a release. Only after both of those are done and verified does the code that reads or writes the new shape go out.

This expand-then-contract sequence turns one risky migration into three smaller, independently reversible steps, each of which is easy to verify before moving to the next.

A common shortcut is skipping straight from adding the new field to switching the application over to it, on the theory that the backfill will finish in time. It usually does, until a slow query or a locked table stalls it mid-run, and now the new code is reading a field that's only partially populated.

Make Every Step Reversible Before You Run It Forward

Add columns instead of renaming them, and keep the old field readable until every reader, including any batch jobs, exports, or downstream services, has moved off it. A rename can't be undone cleanly once something has already written to the new name; an addition can simply be ignored if you need to back out.

Write the backfill so it can be re-run safely if it's interrupted partway through, rather than assuming it will complete in one uninterrupted pass.

Roll Out Behind a Flag, Not a Big-Bang Release

Ship the code that reads the new shape behind a flag while the write path still populates both old and new fields. Flip a small slice of traffic first, watch for errors, and only remove the old path once everything downstream has been confirmed to expect the new one.

How small a step this needs to be tracks your deploy frequency: a team that can release on demand keeps each migration step tiny, while a team stuck releasing only once every 30 days or so ends up bundling bigger, riskier changes into each release by default, simply because that's the next chance they'll get1.

If your team can't yet deploy that often, the flag-based approach still helps: it lets you decouple turning on the new behavior from the less frequent code releases, so you're not stuck choosing between shipping the migration and waiting a month for the next window.

Watch the Signals That Actually Catch Problems

Error rate and write latency on the affected tables tell you more during a cutover than raw CPU or memory ever will. A migration can look perfectly healthy on infrastructure dashboards while quietly writing data in the wrong shape, because CPU doesn't know the difference between a correct write and a corrupt one.

A small reconciliation check, comparing old and new fields for a sample of rows, catches that kind of silent corruption before customers do, and it's worth writing before the migration starts, not after something looks wrong.

How do you decide the rollback trigger before you start?

Pick a specific, measurable trigger in advance: a defined error rate threshold, or a data integrity check failing on more than a small sample. Write it down before the migration begins.

A rollback decision made mid-incident, by a stressed team staring at ambiguous dashboards, is where good migrations turn into bad outages. Deciding the trigger in advance removes that judgment call from the moment when judgment is least reliable.

Make sure the rollback itself has actually been tested, not just written down. A rollback plan that has never been exercised outside of a real incident is a guess with good intentions, and the first time you run it shouldn't be the first time you find out it doesn't quite work.

A safe migration run follows this order:

  1. Add the new column, table, or field in a deploy that does not depend on it yet.
  2. Backfill the data on its own schedule, written so it can be re-run safely if it is interrupted.
  3. Ship the code that reads the new shape behind a flag while writes still populate both the old and new fields.
  4. Move a small slice of traffic first, watching error rate and write latency on the affected tables plus a reconciliation check of old and new fields.
  5. Remove the old path only after every reader has moved off it, and roll back if the trigger you wrote down in advance fires.
Executive Capability Standard

What Good Looks Like

Good migration practice means schema changes, data backfills, and code changes ship as separate, individually reversible steps behind a flag, with a rollback trigger defined before the migration starts rather than decided mid-incident.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read up on the expand-and-contract migration pattern and pick one recent migration your team ran to see which of its steps were actually reversible.
2. Do Manually:Write a short migration checklist, add before backfill before flag before cleanup, and have a second engineer confirm each step before the next one starts.
3. Delegate:Assign a migration reviewer role that rotates among senior engineers, so every schema change gets a second set of eyes on its rollback plan before it ships.
4. Automate:Build a lightweight reconciliation check that compares old and new fields on a sample of rows automatically during the flag-based rollout, rather than relying on someone remembering to spot-check manually.
5. Buy:Bring in a database migration framework or a fractional platform engineer if migrations keep causing incidents and nobody in-house has time to build the tooling around this properly.

How to Get Started

Frequently Asked Questions

Can we run a zero-downtime migration without feature flags?

It's possible for very simple, additive changes, but flags are what let you separate deploying new code from turning on new behavior. Without that separation, you're back to a big-bang release where the schema change and the code change succeed or fail together, which is exactly the risk this approach is meant to avoid.

What's the most common mistake teams make with migrations?

Renaming a column or table in place instead of adding the new shape alongside the old one. A rename can't be reversed cleanly once anything has written to the new name, while an addition can simply be left unused if something goes wrong, which is why expand-then-contract avoids renames almost entirely.

How long should we keep the old schema around after a migration?

Until every reader of the old field, including batch jobs, exports, and any downstream service, has been confirmed to use the new one instead. That's rarely as fast as engineers expect, since forgotten batch jobs and reports are usually the last consumers anyone checks.

What should actually trigger an automatic rollback during a migration?

A specific, pre-defined signal: a fixed error rate threshold being crossed, or a data integrity check failing on more than a small sample of rows. Deciding this before the migration starts matters more than the exact numbers, because it removes the decision from the middle of an incident.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides