Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

A Canary Deployment Runbook That Catches Bad Releases Fast

A canary deployment sends a small slice of production traffic to a new version first, and it works only when the slice is large enough to reveal a regression and the metrics and thresholds pass or fail cleanly. Get the size, the metrics, or the wait time wrong and it either catches nothing or blocks every release.

This runbook works through the decisions in order: sizing the canary, picking metrics that predict a bad release rather than just describing one after the fact, and setting a rollback that fires automatically instead of waiting for someone to notice.

What Canary Actually Buys You Over Blue-Green

Blue-green deployment swaps all traffic to the new version at once, with the old version standing by for a fast rollback. That's simpler to reason about, but it means every user hits the new code at the same moment, so a regression that only shows up under real production load and concurrency hits everyone simultaneously.

A canary limits that blast radius on purpose. Instead of an instant full cutover, a small percentage of real traffic exercises the new version first, with the rest of your users still on the known-good release while you watch. The tradeoff is complexity: you need traffic splitting, per-version metrics, and a defined process for what happens next, not just a switch to flip.

Picking Your Canary Size and Traffic Split

Too small a canary means a real regression doesn't generate enough traffic to be statistically distinguishable from normal noise. Too large defeats the purpose of limiting blast radius. For most services, start small and ramp in stages rather than picking one fixed percentage: an initial slice large enough to generate a meaningful sample within a few minutes at your actual traffic volume, then a series of steps up if the metrics hold.

Low-traffic services need a longer canary window at each stage to accumulate enough requests to trust the comparison, not a smaller starting percentage. If your service handles a modest volume of requests per minute, a short canary window at any traffic split might not see enough real requests to say anything meaningful either way.

Choosing Metrics That Actually Predict a Bad Release

Error rate and latency are the obvious choices, and they matter, but they're lagging indicators: by the time error rate climbs, users have already hit the bug. Add at least one leading indicator specific to what the release actually changed, a business metric like completed checkouts if you touched the checkout flow, or a queue depth if you touched a background worker, compared directly between the canary and the baseline running at the same time.

Comparing the canary against last week's baseline instead of the current baseline running in parallel is a common mistake. Traffic patterns shift day to day and hour to hour for reasons that have nothing to do with your release, so the only comparison that isolates the release itself is canary versus baseline measured over the same window.

Setting an Automatic Rollback, Not Just a Manual One

A rollback that depends on someone watching a dashboard and noticing a problem is a rollback that happens late, especially outside business hours. Define the specific metric thresholds that trigger an automatic rollback before the release ships, not during an incident when everyone's under pressure to guess at a reasonable number.

Keep a manual rollback available too, for the cases automatic thresholds miss: a subtle correctness bug that doesn't move error rate or latency at all, just produces wrong results quietly. Automatic rollback handles the loud failures fast. A human watching the canary window for the first stretch after release still catches the quiet ones.

How Long to Wait Before Promoting to Full Traffic

There's no single right answer, but a few things should all be true before you promote:

  • The canary has run long enough to include at least one full cycle of your traffic pattern, a full business day for most B2B products, so you've seen the actual mix of request types it will face at scale.
  • Every leading and lagging metric you defined has stayed within threshold for the entire window, not just at the final check.
  • Nothing in your team's manual review of logs and traces from the canary looks off, even if the automated metrics look fine.
  • You have a plan for what happens if a regression shows up after full promotion anyway, since a canary reduces risk, it doesn't eliminate it.

Rushing the promotion because a release is time-sensitive defeats the purpose of running a canary at all. If the release truly can't wait for a full cycle, that's a signal to make the canary window shorter and more heavily instrumented, not to skip watching it.

Executive Capability Standard

What Good Looks Like

A good canary process means rollback thresholds and promotion criteria are defined before the release ships, not decided under pressure once something looks wrong.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Study your service's actual traffic pattern, request volume by hour and day, to figure out how long a canary window needs to run to see a representative sample.
2. Do Manually:Run your next release as a manually watched canary: a small traffic split, a fixed time window, and a person comparing canary and baseline metrics side by side before promoting.
3. Delegate:Have the team that owns each service define its own canary thresholds and window length, since traffic patterns and acceptable risk differ by service.
4. Automate:Wire up automatic rollback on defined metric thresholds so a bad release gets pulled back within minutes instead of waiting for someone to notice.
5. Buy:Use a deployment platform with canary support built in rather than building traffic splitting and automated rollback from scratch, unless canary releases are central enough to your product to justify owning that infrastructure.

How to Get Started

Frequently Asked Questions

How is a canary different from a feature flag rollout?

A canary splits traffic at the infrastructure or routing layer between two full deployments of your service. A feature flag rollout ships one deployment and toggles behavior inside it for a subset of users. They solve overlapping problems, but a canary also catches infrastructure and dependency issues a feature flag can't, since the flag doesn't change what's actually running.

What if our traffic is too low to run a meaningful canary?

Extend the canary window instead of shrinking the traffic split further. A low-traffic service needs more time at a given split to accumulate a comparable sample, not a smaller starting percentage, which would only make the sample size problem worse.

Should every deploy go through a canary?

For anything touching a core user flow or a shared dependency, yes. For a low-risk change like a copy update or an internal admin tool, a full canary process is often more overhead than the risk justifies, and a simpler rollout with good monitoring is enough.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides