AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

Build or Buy: Canary Deployments for a Small Team

A canary release sends a new version to a small slice of traffic before it reaches everyone, so a bad deploy hurts a handful of users instead of all of them. The idea is simple. Building it well enough to actually catch a bad deploy before it spreads is where teams either overbuild or underbuild.

The honest question isn't whether canaries are worth doing. It's whether you need a dedicated platform for it yet, or whether your existing load balancer and a bit of discipline already gets you most of the way there.

What a canary actually needs to catch

A canary is only useful if you're watching the right signal on the slice of traffic hitting it: error rate, latency, and any business metric that would flag a subtler problem, like a checkout completion rate that quietly drops without throwing a single error. Watching only server error codes misses the class of bug that degrades the product without technically failing.

Decide the promotion criteria before the first canary ships, not while staring at a dashboard mid-rollout. A vague sense that things look fine is how a bad deploy gets waved through under pressure.

Traffic shifting you can build with a load balancer alone

Most load balancers and ingress controllers already support weighted routing between two versions. Standing up a canary this way, ten percent of traffic to the new version, ninety to the old, for example, is a configuration change, not a new system. For a small team, this covers a surprising amount of what a canary needs to do.

What it doesn't give you is automatic analysis. Someone has to watch the metrics and decide, by hand, whether to promote or roll back. That's a real cost, but it's a cost you can absorb manually long before it's worth automating.

For example, a small team can run a weighted split of ten percent to the new version, then follow a written checklist: check error rate, latency, and one business metric at fixed intervals, promote only if all three match the old version, and roll back as soon as any of them diverges. The checklist turns a vague sense that things look fine into a decision anyone on call can make. When the manual steps start to eat real engineering time, that is the signal to consider automating them.

What a managed platform buys you beyond the basics

A dedicated progressive delivery platform automates the part a manual canary leaves to a human: it watches the metrics you define, shifts traffic upward on its own when they look healthy, and rolls back automatically when they don't. A canary rollout is what makes on-demand deployment survivable instead of risky, which is exactly the habit that separates DORA's highest-performing teams from everyone else1.

That automation is worth paying for once deploys are frequent enough that a person watching every one by hand becomes the bottleneck, not before.

The metrics that decide when a canary gets promoted

Error rate and latency are the obvious pair, but they're not sufficient on their own. Add at least one metric tied to what the feature is actually supposed to do, so a canary that's technically healthy but functionally broken still gets caught. And set a minimum observation window, because a canary that looks fine for ninety seconds can still be hiding a problem that only shows up under sustained load.

Write these thresholds into the deploy process itself where you can, even if enforcement is still manual at first. A checklist that says what green and red look like for this specific release is worth more under pressure than a person's memory of what normal usually looks like.

A written promotion rule should cover these points:

  • Error rate and latency on the slice of traffic that hits the canary.
  • At least one metric tied to what the feature is supposed to do, so a technically healthy but functionally broken canary is still caught.
  • A minimum observation window, since a canary that looks fine for ninety seconds can still hide a problem that appears under sustained load.
  • What green and red look like for this specific release, written down before the rollout starts.

When to build, and when the tradeoff flips

Start with manual, weighted traffic shifting on infrastructure you already run. Move to a managed platform once deploys are frequent enough that manual promotion decisions are eating real engineering time, or once a bad deploy has actually reached full traffic because nobody was watching closely enough during the canary window. That second trigger is the one that tends to make the decision for you.

There's a middle option worth knowing about too: some teams write a small script that polls their existing metrics platform and pages someone the moment a canary's error rate crosses a threshold, well short of building full automated promotion. That covers the worst failure mode, a bad deploy running unnoticed, without the cost of a full platform.

Executive Capability Standard

What Good Looks Like

Good canary practice means the promotion criteria are written down before the rollout starts, and at least one metric ties directly to what the feature does.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pick your two riskiest recent deploys and work out what a canary would have needed to show to catch each one.
2. Do Manually:Start weighted traffic shifting through your existing load balancer, watching metrics by hand before promoting each release.
3. Delegate:Give a specific engineer ownership of promotion criteria for each service, written down rather than decided ad hoc during each rollout.
4. Automate:Automate the promotion and rollback decision against defined metric thresholds once manual promotion becomes a bottleneck.
5. Buy:Adopt a managed progressive delivery platform once deploy frequency makes hand-watched canaries too slow to keep up with.

How to Get Started

Frequently Asked Questions

Do we need a dedicated canary platform to get started?

No. Most load balancers already support weighted traffic splitting, which covers the mechanics of a canary release. What you'll do manually at first is watching the metrics and deciding whether to promote, which a platform later automates.

What metrics should trigger an automatic rollback?

Error rate and latency at minimum, plus at least one metric tied to what the feature actually does. A canary that's technically error-free but silently breaks a business flow, like checkout completion, won't get caught by error rate alone.

How long should a canary run before promotion?

Long enough to see the traffic pattern you actually care about. A brief window can miss a problem that only appears under sustained load or a specific time of day, so set a minimum observation period rather than promoting the moment early numbers look clean.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides