Blue-Green, Canary or Rolling: Picking a Deployment Strategy
Most small engineering teams pick a deployment strategy by copying whatever a blog post from a much larger company describes, then spend months fighting a setup that doesn't match their actual traffic pattern or database. The three common strategies, rolling, blue-green and canary, trade off differently on rollback speed, infrastructure cost and how much they can catch before all your users see a bad release.
The right choice depends on three things: how expensive a bad deploy is for you, whether your database schema changes can be made backward-compatible, and how much duplicate infrastructure you can justify running.
When is a rolling deployment the right default?
A rolling deployment replaces old instances with new ones a few at a time, so you never run double the infrastructure but you do run two versions of your code side by side for the duration of the rollout. This is the right default for most small teams: it's supported natively by most container orchestrators and needs no extra tooling.
The catch is that a bad release still reaches a slice of real users before you notice, and rollback means running another rolling deployment backward, which takes nearly as long as the original rollout. If your rollout takes ten minutes, a bad release can be live and serving traffic for most of that window before anyone catches it.
Running two code versions side by side also means your API and database schema have to tolerate both versions at once, even briefly. A rolling deploy that pairs old application code with a database column the new code expects to exist (or vice versa) is one of the more common self-inflicted outages, and it has nothing to do with which deployment strategy you picked.
Blue-green: instant rollback, at the cost of double capacity
Blue-green keeps two complete environments, one live and one idle, and switches traffic from one to the other all at once. The advantage is rollback speed: switching back is a routing change, not a redeploy, so a bad release is live for minutes instead of the length of a rollout. The cost is running two full copies of your production environment, which matters more once your infrastructure bill is a real number and less when you're small enough that doubling it is a rounding error.
Blue-green also complicates any deploy that includes a database migration, since both environments need to work against the same database during the switch. Migrations have to be written to be safe for the old and new code to run against simultaneously, which is a discipline worth having regardless of deployment strategy.
Say a ten-person team runs blue-green for a checkout service that processes payments. The doubled infrastructure cost for that one service, maybe a few hundred dollars a month, is small next to the cost of a bad release staying live even briefly on a payment path. The same team might reasonably run their internal admin dashboard as a plain rolling deployment, since the cost of a slow rollback there is measured in inconvenience, not revenue.
Canary: catch problems before they reach everyone
A canary release sends a small percentage of real traffic (for example, 5 percent) to the new version, watches error rates and latency against the old version, and only proceeds to full rollout if the numbers hold. This is the strategy that actually limits blast radius: a bad release affects a small slice of users for a short window instead of everyone.
The tradeoff is operational complexity. You need traffic splitting at the load balancer or service mesh layer, automated comparison of error rates between the canary and the baseline, and a clear rule for what triggers an automatic rollback. Teams that adopt canary releases without building the automated comparison step just end up watching a dashboard manually, which defeats most of the benefit.
How do you match a deployment strategy to your risk?
If a bad deploy means a visible outage for paying customers, canary is worth the setup cost even at ten engineers. If a bad deploy means an internal tool is broken for twenty minutes until someone rolls back manually, rolling deployments are enough and canary infrastructure would be solving a problem you don't have.
Teams that deploy frequently tend to need less elaborate deployment strategies, not more, because each individual change is smaller and easier to reason about. Organizations in DORA's highest-performing cluster practice on-demand deployment, often shipping several times a day, while lower-performing organizations may go weeks or months between releases1; that cadence alone shrinks the blast radius of any single release before you've added any deployment strategy on top.
A useful test: write down your last three bad deploys and how each one was actually caught, a customer complaint, an alert, an engineer noticing something off in a dashboard. If a customer complaint was the fastest detection path more than once, that's a stronger signal that you need canary-style traffic limiting than any theoretical argument about best practice.
Use these rules of thumb when choosing:
- Choose rolling when a bad deploy is tolerable and you want to avoid running duplicate infrastructure or adding new tooling.
- Choose blue-green when fast rollback matters most and you can justify running two full copies of production.
- Choose canary when a bad deploy means a visible outage for paying customers and you can take on the added operational setup.
- Before any of the three, confirm your API and database schema changes can tolerate old and new code running at the same time.
What Good Looks Like
A production deployment strategy match your actual risk: rollback time, database migration safety and blast radius are all deliberate choices, not accidents of whatever tooling you started with.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Can we mix strategies, like canary for the backend and rolling for a frontend?
Yes, and it's common. The backend usually carries more risk per deploy (data writes, downstream services) so it's the more common candidate for canary, while a static frontend can often ship with a simple rolling or even instant swap.
How small a canary percentage should we start with?
One to five percent of traffic is typical for a first canary stage, held for long enough to accumulate a statistically meaningful sample of your error rate, often 15 to 30 minutes depending on your traffic volume.
Does feature-flagging replace the need for a deployment strategy?
No. Feature flags control who sees a feature after code is already running in production; a deployment strategy controls how new code reaches production in the first place. Most mature setups use both together.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Canary Releases: How Much Traffic, How Fast
A canary that bakes for ten minutes at five percent traffic misses a memory leak that shows up an hour in. How to size and gate a canary release.
What to Track About Engineering Productivity Besides DORA
Why DORA's four metrics don't capture the whole picture of engineering health, and what to measure alongside them without turning metrics into a scoreboard.
Getting a New Engineer to Their First Production Deploy Faster
How to shrink the time between a new engineer's start date and their first production deploy, without cutting corners on access or review.
What 'Zero Trust' Actually Requires From Every Device on Your Network
What zero trust device verification actually requires in practice, beyond the buzzword, and where small teams should start first.
Improving Developer Experience Without Buying Another Tool
A practical way to measure and fix developer experience problems, from local setup time to documentation findability, before reaching for new software.
Building a Latency Budget Before You Chase Microsecond Fixes
Why teams that tune latency without a budget waste weeks on the wrong service, and how to build one that tells you exactly where to look first.