Feature flagsPlaybook3 min readUpdated September 2026

Canary Releases Using Feature Flags: A Staged Rollout Playbook

A flag-based canary release ships new code to production turned off, then enables it for a small group of users, watches guardrail metrics and widens exposure in steps. If a metric crosses a limit, you switch the flag off instead of redeploying.

The technique works because deploying code and exposing it are separate decisions. Below is a playbook: what to decide before the first user sees the change, how to stage the rollout, when to roll back and how to finish.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What do you decide before turning the flag on?

A canary goes better when you prepare before flipping the flag. Write these down first:

  1. The change and its blast radius. What behavior differs, which users and services does it touch, and what data does it write?
  2. Guardrail metrics. The signals that say 'stop': error rate, latency at the slow end, and one business metric such as checkout completion or login success.
  3. Thresholds. The exact numbers that trigger rollback, agreed in advance. Deciding mid-incident is how teams talk themselves out of rolling back.
  4. Who watches, and how long. A named person, a dashboard and a minimum observation time at each stage.
  5. The off switch. Confirm that turning the flag off really restores the old path, including for users already mid-session.

If the new code changes a database schema, split that work first, as described in the last section.

How do you stage the rollout?

Widen exposure in steps, and hold at each one long enough to see real traffic:

  • Stage 0, internal users. Enable for employees or a test tenant. Catches obvious breakage with no customer impact.
  • Stage 1, a small slice of production. For example, a few percent of users chosen by a stable hash of their ID, so each person keeps seeing the same version.
  • Stage 2, a larger slice. Move up only when guardrails stay flat across a full cycle of your usual traffic pattern, including peak hours.
  • Stage 3, most users, then everyone.

Percentages are examples, not rules. Pick stages that give you enough traffic to see a problem in your metrics, so a low-volume service needs bigger early slices than a busy one. Prefer consistent bucketing by user or account, not by request, or one person will bounce between old and new behavior.

When should you roll back, and how?

Roll back when a guardrail crosses its threshold, when support reports a pattern, or when you can't explain a change in the metrics. Don't wait for certainty.

Make the rollback a single action:

  • Flip the flag to off for everyone, or for the affected segment if you can isolate it.
  • Confirm in the dashboard that traffic returned to the old path and that errors fell.
  • Record the time, the metric that fired and the suspected cause.
  • Leave the flag off until someone owns the fix.

Sometimes the flag service itself is the problem. Decide the fallback value when the service can't be reached, and make it the safe path, which is usually the old behavior. Test that path deliberately in staging by blocking the flag service.

Which mistakes make flag-based canaries fail?

These come up repeatedly:

  • Metrics can't separate the groups. If dashboards don't split by flag variant, a problem affecting a small slice of users disappears in the average. Tag events with the variant.
  • Sticky data changes. New code writes data in a format old code can't read, so turning the flag off breaks people. Keep writes compatible in both directions until the end.
  • Sessions and caches. Cached pages or long-lived sessions keep the old or new version longer than the flag says.
  • Ramping on a schedule instead of evidence. Moving up because it's Friday isn't a rollout plan.
  • Too many flags at once. Two canaries running together make failures hard to attribute.
  • Nobody owns the end. The flag stays fully on for a year.

For tool selection, the comparison of LaunchDarkly, Split and Flagsmith explains what to check.

How do you handle database changes and finish cleanly?

Schema changes are the hardest part of any canary. Use an expand, migrate, contract sequence:

  1. Expand. Add the new column or table in a change both old and new code can work with, and deploy it before the flag.
  2. Dual write or backfill. Have the new path write to both, or backfill, while the flag ramps up.
  3. Switch reads once the data is complete and verified.
  4. Contract. After the flag has been at full exposure and stable, remove the old columns and old code path.

When the rollout is complete, delete the flag. A long-lived release flag is debt, as covered in the flag cleanup guide. Use a clear name that shows the flag is temporary, following the naming convention, and add the removal task to the ticket when you create the flag.

Executive Capability Standard

What Good Looks Like

Every risky change ships behind a flag, exposure widens in stages against pre-agreed guardrails, and one action restores the old behavior.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read how percentage rollouts and user targeting work in your flag tool, and identify your three most important guardrail metrics.
2. Do Manually:Run one low-risk change through internal, small and wider stages by hand, with a named watcher and written rollback thresholds.
3. Delegate:Give each release an owner who approves every stage change and records decisions in the release ticket.
4. Automate:Tag telemetry by flag variant and trigger alerts or automatic flag-off when a guardrail crosses its threshold.
5. Buy:Adopt a dedicated flag platform with targeting, audit history and progressive rollout support once several teams ship this way.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

What is a canary release?

It's a rollout where a new version or feature reaches a small group of users first, while you watch for problems. If metrics stay healthy, you expand exposure gradually. If they don't, you stop or roll back, so few users are affected.

How is a flag-based canary different from a deployment canary?

A deployment canary sends a share of traffic to new servers or containers. A flag-based canary deploys the code everywhere and controls who runs it with a flag. Flags give finer control by user or account, and rollback takes seconds without a redeploy.

How long should each canary stage last?

Long enough to see representative traffic and enough events to notice a problem, typically covering at least one peak period. Low-traffic products need longer holds or larger early slices. Set the minimum time before you start so pressure doesn't shorten it.

What metrics should I monitor during a canary?

Error rate, latency for the slowest requests, and one or two business metrics tied to the feature, such as signup or checkout completion. Break each down by flag variant so you compare the canary group against everyone else.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides