Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

How to Upgrade a Major Dependency Without a Maintenance Window

The instinct with a big version migration, a database engine upgrade, a major framework bump, a breaking API change, is to schedule a maintenance window and do it all at once. That works right up until something in the new version behaves differently under real production load than it did in staging, and now you're debugging live with your users locked out and the whole team watching a clock.

Zero-downtime migrations aren't about being faster at the big-bang approach. They're about not doing a big-bang migration at all, and running old and new side by side long enough to be sure.

Why "We'll Just Do It Off-Hours" Isn't the Same as Zero Downtime

Off-hours deployment reduces how many users notice a problem, but it doesn't reduce the risk of the migration itself, and it adds a nasty side effect: your quietest traffic period is also the one where a subtle bug is least likely to surface before it's rolled out everywhere. A change that behaves fine under low load can fall over hours later once real daytime traffic hits it.

Treat off-hours timing as a courtesy to your users, not a safety mechanism. The actual safety comes from how the migration is sequenced, not when the clock says it happened.

The Expand-Contract Pattern for Schema and API Changes

Expand-contract breaks a breaking change into three steps that are each individually safe to deploy and roll back. First, expand: add the new column, endpoint, or field alongside the old one, so both exist at once. Second, migrate: move reads and writes over to the new path while the old one still works as a fallback. Third, contract: remove the old path only once nothing is using it anymore, verified by actual traffic data, not by assumption.

The unglamorous part is that step three often gets skipped, leaving two versions of everything live indefinitely. Put a date and an owner on the contract step when you plan the migration, not just on the expand step.

The expand-contract pattern runs in three safe steps:

  1. Expand: add the new column, endpoint or field alongside the old one, so both versions work at the same time.
  2. Migrate: move readers and writers to the new version gradually, keeping each deploy small and easy to roll back.
  3. Contract: once every consumer has moved and error rates stay flat, remove the old path instead of leaving it live.

Running Two Versions in Production at Once, Safely

Teams in DORA's top-performing cluster, the ones with the highest deployment frequency, tend to treat a version migration as a series of small, reversible deploys rather than one big cutover; teams whose deployment frequency is much lower1 default to big-bang migrations because that's the only rhythm they know, and building a gradual rollout muscle for the first time in the middle of a risky migration is a bad place to learn it.

Running both versions at once means routing a small, controlled slice of traffic to the new path first, watching error rates and latency against the old path's baseline, and only widening that slice once the new path has proven itself under real conditions, not synthetic tests. Decide in advance what "proven itself" means, specific error-rate and latency thresholds, so widening the rollout is a decision made from a plan and not a gut call under pressure.

The Rollback Test Most Teams Skip

Every migration plan has a rollback step written down. Far fewer teams have actually executed it before they need it for real. A rollback that only exists on paper often turns out to be one-way in practice, because the new version already wrote data in a shape the old version can't read.

Before you start migrating real traffic, run the rollback deliberately against a copy of production data and confirm the old version still functions correctly against anything the new version wrote during the migration window. If it can't, your rollback plan is fiction.

Sequencing a Multi-Service Migration Without a Big-Bang Cutover

When the change spans several services, order matters more than speed. Migrate the service with the fewest dependents first, so a mistake there has the smallest blast radius, and hold off on migrating anything that other unmigrated services still depend on directly. Trying to move everything in one release just recreates the big-bang problem you were avoiding, spread across more teams instead of contained to one.

Keep a simple, shared tracking document of which services are on the old version, which are on the new one, and which are mid-migration, so nobody accidentally builds a new feature against a path that's about to be removed. This sounds like overhead until the first time it prevents a team from shipping against a dependency that was scheduled for removal that same week.

Executive Capability Standard

What Good Looks Like

A good version migration runs old and new code side by side, moves traffic gradually based on real error and latency data, and has a rollback path that's actually been tested, not just written down.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read through the last major migration this team ran and note where the rollback plan was never actually exercised.
2. Do Manually:Write an expand-contract plan for the next migration with explicit owners and dates for each of the three steps.
3. Delegate:Assign a senior engineer to own the traffic-shifting schedule and the go or no-go call at each stage of the rollout.
4. Automate:Wire the traffic-shift percentages to your deployment pipeline so widening the rollout doesn't require a manual config change each time.
5. Buy:Bring in a contractor or fractional engineer experienced in large migrations when the change spans more services than your team has bandwidth to sequence safely.

How to Get Started

Frequently Asked Questions

How long should we run old and new versions side by side?

Long enough to see your real traffic patterns at least once, including your busiest regular day and any weekly or monthly peak, not just a quiet afternoon. For most teams that's somewhere between one and a few weeks, but the real signal is whether error rates and latency on the new path have stayed flat, not the calendar.

Do we need feature flags for every migration?

Not every one, but any migration touching a widely used code path benefits from a flag that lets you shift traffic gradually and roll back instantly without a redeploy. For small, isolated changes with few dependents, a flag can add more overhead than it saves; reserve them for changes where a bad rollout would be hard to reverse quickly.

What's the biggest sign a migration is being rushed?

Skipping the contract step, removing the old code path, because the team is eager to call the migration done. If the old path is still technically live weeks after the new one launched with no plan to remove it, that's usually a sign the migration was never actually finished, just started.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides