Data Engineering & Real-Time Event StreamsPlaybook3 min readUpdated September 2026

Blue-Green, Canary, or Rolling: Deploying Stream Processors

Choose a rolling deploy for backward-compatible stream processor changes, blue-green when output schema or side effects change, and a canary when the change is genuinely uncertain. Stateful processors carry checkpoints, consumer offsets, and local state that a bad deploy can corrupt silently, so rollback needs a plan you've actually tested.

This is a guide to picking a deploy strategy for that kind of system, and to building a rollback plan you've actually tested, not just written down in a runbook nobody has opened since it was created.

Why a stateful stream processor breaks the usual playbook

A web service deploy is forgiving: if the new version misbehaves, you roll back and the next request just works. A stream processor's new version might have already committed offsets past events it processed incorrectly, or written bad state into a store that the old version will read from after rollback. The mistake doesn't stay contained to the deploy window; it travels forward in the data.

Before picking a strategy, decide how you'll detect a bad deploy fast: a spike in dead-letter volume, a drop in output record counts, or a specific business metric a new version is expected to move. Without that signal, any deploy strategy just buys you a slower way to notice the same failure.

Rolling deploys: the default, and where it falls short

Rolling deploys replace consumer instances one at a time, which is simple and works fine for stateless processing logic. The risk is that old and new versions run side by side during the rollout, consuming from the same topic, and if the new version's output schema or business logic has changed, downstream consumers see an inconsistent mix of old and new behavior mid-deploy.

Rolling deploys are the right default when a change is backward compatible: a performance fix, a log line, a dependency bump. They're the wrong choice for anything that changes what a message means, because there's no clean line between old and new behavior during the rollout.

Blue-green deploys: a clean cutover, at double the compute cost

A blue-green deploy runs the new version as a full, separate consumer group against the same topics, lets it catch up to the current offset, then switches traffic over in one move. It gives you a clean, instant rollback (switch back to the old group) and no mixed-version window, at the cost of running two full copies of your processing fleet during the cutover.

This is worth it for changes that touch output schema, side effects like writing to a downstream database, or anything where a mixed old and new behavior window would be genuinely harmful rather than just untidy. It's overkill for routine fixes where a rolling deploy's mixed window doesn't actually matter.

Canary deploys: testing on a slice of real traffic first

A canary deploy routes a small percentage of partitions, or a single low-traffic topic, to the new version while the rest of the fleet stays on the old one. It catches a bad deploy on a fraction of your data instead of all of it, which matters most when the change is genuinely uncertain, not just routine.

The tradeoff is that a canary needs partition or topic level traffic splitting, which not every streaming setup supports cleanly, and it takes longer to fully roll out than the other two approaches since you're deliberately going slow on the riskiest changes.

How deploy frequency should change as the pipeline matures

The top DORA performance cluster ships on-demand deployments, often more than once a day, while the lowest cluster can go as long as 180 days between releases1. A mature real-time pipeline should be closer to the first end of that range for backward-compatible changes, and deliberately slower and more careful for anything that touches state or schema.

The goal isn't deploying as fast as possible everywhere. It's matching deploy frequency to risk: routine, backward-compatible changes should ship often and automatically, while schema or state-changing changes should go through the blue-green or canary path above even if that means shipping them less frequently.

Build a rollback plan you test, not one you just write

A rollback plan that's only ever been read, never executed, usually fails the first time it's needed for real, typically because a consumer offset or a piece of state assumed to still be compatible with the old version isn't. Run a rollback drill on a non-production topic at least once before you rely on it in an incident.

Document exactly what "rollback" means for your specific setup: reverting code only, resetting consumer offsets to a known point, or restoring state from a snapshot. Those are three different operations with three different blast radiuses, and conflating them during an actual incident is how a bad deploy turns into a bad outage.

Before you ship a stream processor change, check the following:

  • Decide how you'll detect a bad deploy quickly, such as a spike in dead-letter volume or a drop in output record counts.
  • Use a rolling deploy for backward-compatible changes, where old and new versions running side by side won't confuse downstream consumers.
  • Choose blue-green when output schema or side effects change and you can afford a second full consumer group during cutover.
  • Choose a canary when the change is genuinely uncertain and you can split traffic by partition or topic.
  • Run a rollback drill on a non-production topic and confirm offsets and state stay compatible with the old version.
Executive Capability Standard

What Good Looks Like

Deployment is production-grade when every change is classified as backward compatible or not before it ships, the deploy strategy matches that classification, and rollback has been tested, not just documented.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map your current deploy process against the three strategies above and identify which changes are shipping through the wrong one.
2. Do Manually:Run a rollback drill on a non-production topic so the team knows exactly what reverting actually involves for your setup.
3. Delegate:Give a senior engineer ownership of the backward compatibility check that decides which deploy strategy a given change uses.
4. Automate:Wire schema compatibility checks and dead-letter volume alerts into the deploy pipeline so a bad deploy is caught within minutes.
5. Buy:Bring in a streaming infrastructure specialist to set up blue-green or canary tooling once if your platform doesn't support it natively.

How to Get Started

Frequently Asked Questions

Which deploy strategy should a small team default to?

Rolling deploys for anything backward compatible, which covers most day-to-day changes. Reserve blue-green for changes to output schema or side effects, and canary for genuinely uncertain changes where you want to see real production behavior on a slice of traffic before committing the whole fleet to it.

How do we know if a schema change is actually backward compatible?

Test it against your schema registry's compatibility check before deploying, not just by inspection. A field addition is usually safe; a field removal, rename, or type change almost never is for any consumer still running the old code. When in doubt, treat it as breaking and use blue-green.

Is a canary deploy worth the added complexity for a small pipeline?

Only once you have enough topics or partitions to meaningfully split traffic, and enough at stake in a bad deploy to justify the setup work. Below that threshold, a well-tested rollback plan on a standard rolling or blue-green deploy usually gets you most of the same protection for less operational overhead.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides