Blue-Green, Canary, or Rolling: Deploying Stream Processors
Choose a rolling deploy for backward-compatible stream processor changes, blue-green when output schema or side effects change, and a canary when the change is genuinely uncertain. Stateful processors carry checkpoints, consumer offsets, and local state that a bad deploy can corrupt silently, so rollback needs a plan you've actually tested.
This is a guide to picking a deploy strategy for that kind of system, and to building a rollback plan you've actually tested, not just written down in a runbook nobody has opened since it was created.
Why a stateful stream processor breaks the usual playbook
A web service deploy is forgiving: if the new version misbehaves, you roll back and the next request just works. A stream processor's new version might have already committed offsets past events it processed incorrectly, or written bad state into a store that the old version will read from after rollback. The mistake doesn't stay contained to the deploy window; it travels forward in the data.
Before picking a strategy, decide how you'll detect a bad deploy fast: a spike in dead-letter volume, a drop in output record counts, or a specific business metric a new version is expected to move. Without that signal, any deploy strategy just buys you a slower way to notice the same failure.
Rolling deploys: the default, and where it falls short
Rolling deploys replace consumer instances one at a time, which is simple and works fine for stateless processing logic. The risk is that old and new versions run side by side during the rollout, consuming from the same topic, and if the new version's output schema or business logic has changed, downstream consumers see an inconsistent mix of old and new behavior mid-deploy.
Rolling deploys are the right default when a change is backward compatible: a performance fix, a log line, a dependency bump. They're the wrong choice for anything that changes what a message means, because there's no clean line between old and new behavior during the rollout.
Blue-green deploys: a clean cutover, at double the compute cost
A blue-green deploy runs the new version as a full, separate consumer group against the same topics, lets it catch up to the current offset, then switches traffic over in one move. It gives you a clean, instant rollback (switch back to the old group) and no mixed-version window, at the cost of running two full copies of your processing fleet during the cutover.
This is worth it for changes that touch output schema, side effects like writing to a downstream database, or anything where a mixed old and new behavior window would be genuinely harmful rather than just untidy. It's overkill for routine fixes where a rolling deploy's mixed window doesn't actually matter.
Canary deploys: testing on a slice of real traffic first
A canary deploy routes a small percentage of partitions, or a single low-traffic topic, to the new version while the rest of the fleet stays on the old one. It catches a bad deploy on a fraction of your data instead of all of it, which matters most when the change is genuinely uncertain, not just routine.
The tradeoff is that a canary needs partition or topic level traffic splitting, which not every streaming setup supports cleanly, and it takes longer to fully roll out than the other two approaches since you're deliberately going slow on the riskiest changes.
How deploy frequency should change as the pipeline matures
The top DORA performance cluster ships on-demand deployments, often more than once a day, while the lowest cluster can go as long as 180 days between releases1. A mature real-time pipeline should be closer to the first end of that range for backward-compatible changes, and deliberately slower and more careful for anything that touches state or schema.
The goal isn't deploying as fast as possible everywhere. It's matching deploy frequency to risk: routine, backward-compatible changes should ship often and automatically, while schema or state-changing changes should go through the blue-green or canary path above even if that means shipping them less frequently.
Build a rollback plan you test, not one you just write
A rollback plan that's only ever been read, never executed, usually fails the first time it's needed for real, typically because a consumer offset or a piece of state assumed to still be compatible with the old version isn't. Run a rollback drill on a non-production topic at least once before you rely on it in an incident.
Document exactly what "rollback" means for your specific setup: reverting code only, resetting consumer offsets to a known point, or restoring state from a snapshot. Those are three different operations with three different blast radiuses, and conflating them during an actual incident is how a bad deploy turns into a bad outage.
Before you ship a stream processor change, check the following:
- Decide how you'll detect a bad deploy quickly, such as a spike in dead-letter volume or a drop in output record counts.
- Use a rolling deploy for backward-compatible changes, where old and new versions running side by side won't confuse downstream consumers.
- Choose blue-green when output schema or side effects change and you can afford a second full consumer group during cutover.
- Choose a canary when the change is genuinely uncertain and you can split traffic by partition or topic.
- Run a rollback drill on a non-production topic and confirm offsets and state stay compatible with the old version.
What Good Looks Like
Deployment is production-grade when every change is classified as backward compatible or not before it ships, the deploy strategy matches that classification, and rollback has been tested, not just documented.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Which deploy strategy should a small team default to?
Rolling deploys for anything backward compatible, which covers most day-to-day changes. Reserve blue-green for changes to output schema or side effects, and canary for genuinely uncertain changes where you want to see real production behavior on a slice of traffic before committing the whole fleet to it.
How do we know if a schema change is actually backward compatible?
Test it against your schema registry's compatibility check before deploying, not just by inspection. A field addition is usually safe; a field removal, rename, or type change almost never is for any consumer still running the old code. When in doubt, treat it as breaking and use blue-green.
Is a canary deploy worth the added complexity for a small pipeline?
Only once you have enough topics or partitions to meaningfully split traffic, and enough at stake in a bad deploy to justify the setup work. Below that threshold, a well-tested rollback plan on a standard rolling or blue-green deploy usually gets you most of the same protection for less operational overhead.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Sizing a Canary Deployment So It Actually Catches Bad Releases
How to size a canary deployment, pick the metrics that actually catch a bad release, and decide when to build this in house versus buy a platform.
Verifying Every Service That Talks to Your Pipeline
Which parts of zero-trust verification to build and which to buy, so every producer and consumer on a streaming pipeline proves its identity.
Where Latency Actually Hides in a Growing Data Pipeline
A walkthrough of where latency hides as a real-time pipeline grows, from producer batching to consumer lag, so you can find your own bottleneck fast.
Decoupling Services With Events Without Losing Traceability
A worked example of decoupling two services with an event queue, and the specific traceability and ordering problems that show up once you do.
Mapping SOC 2 Controls to a Real-Time Streaming Pipeline
How SOC 2 trust service criteria actually map onto a streaming pipeline's controls, and where a governance policy has to go beyond what a tool tracks.
How to Run a Security Audit on a Real-Time Data Pipeline
A step by step way to check access, encryption, and patch timelines on your event streams before an incident or an auditor finds the gap first.