A Production Deployment Checklist That Actually Catches Problems
Most production incidents in a distributed system trace back to a deploy, not a design flaw. The code was fine; the rollout wasn't. Services deployed out of order, a rollback that turned out to be one-way, or a migration that ran before the code expecting it did.
This is a stage-by-stage checklist for the parts of a deployment that get skipped under deadline pressure, and the ones that cause the most expensive incidents when they do.
Before you deploy: confirm rollback actually works
A rollback plan that has never been tested is a guess. Before shipping anything that touches a shared dependency, database schema, or message format, confirm you can revert the code without reverting the data, or that reverting the data is also safe.
Database migrations are the usual trap: a migration that drops a column can't be rolled back by redeploying the old code, because the old code still expects that column to exist. Expand-then-contract migrations, where you add the new shape, deploy code that works with both, then remove the old shape later, avoid this.
Order matters when services depend on each other
In a distributed system, deploying the consumer of an API before the producer supports the new contract breaks requests immediately. Deploying the producer first and keeping it backward compatible until every consumer has updated is the safer default.
Write down the dependency order for any change that touches more than one service, and don't rely on remembering it under pressure during a release window.
Ship in stages, not all at once
A canary release to a small slice of traffic, watched against your key metrics before a full rollout, catches problems that pass every test but show up under real traffic patterns, uneven data, unusual request shapes, load that a staging environment never replicates.
Deploy frequency is a real signal of engineering health, and teams that release small, frequent changes tend to ship faster and with less risk per release than teams batching everything into a monthly window: research on deploy frequency puts the fastest-moving teams shipping on demand, essentially daily, while the slowest cluster manages roughly one release a month1. A smaller, more frequent release is also just easier to reason about: fewer changes bundled together means a regression points to a shorter list of suspects.
The checks that get skipped when a deploy is running late
- Verifying the feature flag actually defaults to off before the code ships
- Checking that the rollback path has been exercised in staging, not just written down
- Confirming downstream consumers know about a contract change before it goes out
- Watching error rates and latency for at least one full traffic cycle after the deploy, not just the first five minutes
Every one of these takes minutes to do right and hours to recover from if skipped.
What to do when a deploy goes wrong anyway
Decide in advance who has the authority to roll back without a meeting. The minutes spent debating whether a spike is real or noise are usually the most expensive minutes of an incident.
Write a short postmortem for every rollback, even the quiet ones that never became customer-facing incidents. The pattern that caused a near-miss this month is often the same one that causes a real incident next quarter if it goes unaddressed.
Coordinating a deploy across teams that don't talk daily
In a distributed system with several teams, the riskiest deploys aren't the big ones everyone plans around, they're the small ones that quietly change a contract another team depends on. A release calendar that other teams can actually see, not just your own, cuts down on the surprise factor.
For any change to a shared API or message format, give downstream teams a backward-compatibility window: ship the new shape alongside the old one, announce a removal date, and only take the old shape away once you've confirmed nobody's still calling it. Skipping that window is the single most common cause of a deploy that looks clean in isolation but breaks something two services away.
Say the payments team renames a field in an event it publishes. If three other services consume that event and only one of them is on the release calendar you checked, the other two start failing silently the moment the change ships, and nobody notices until a downstream report comes back wrong days later. A quick search across the codebase for who subscribes to an event, done before the change goes out, catches this in minutes instead of after the fact.
What Good Looks Like
A mature deployment process ships in small, reversible stages, has a rollback plan that's actually been tested, and gives someone clear authority to revert without waiting for a meeting.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How many people should be required to approve a production deploy?
One reviewer for routine changes is usually enough if your tests and staging environment are trustworthy. Reserve a second approval for changes touching shared infrastructure, database schemas, or anything with no clean rollback path.
Is a feature flag always safer than a full deploy?
It's safer for the rollback story, since flipping a flag is faster than redeploying, but it adds code complexity and a flag that never gets cleaned up becomes its own source of confusion. Use flags for anything risky, and set a reminder to remove them once the change is proven.
What's the minimum viable rollback plan for a small team?
A tested one-command revert to the last known-good deploy, and a rule that database migrations are always additive first. That covers most incidents without requiring a full blue-green or canary setup you don't have the headcount to maintain.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Canary, Blue-Green or Feature Flag: Matching the Rollout to the Risk
A decision guide for choosing between canary deployments, blue-green releases and feature flags, based on what kind of change you're actually shipping.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
Finding the Real Source of Latency in a Distributed System
A decision guide for narrowing down whether a slow request is a network problem, a database problem, a queue problem, or your own code.
Load Testing Numbers That Don't Match What Users Actually Feel
Why a clean throughput benchmark often fails to predict real-world scaling behavior, and how to build one around your real traffic mix and first bottleneck.
Cache Invalidation Is Still the Hard Part
A practical guide to choosing a caching layer and, more importantly, keeping it from serving stale or wrong data across a distributed system.
How to Run a Real Security Audit on a Distributed System
A working method for auditing service boundaries, credentials, and patch timelines across a distributed system instead of filling out a compliance checklist.