Building a CI/CD Pipeline That Understands Streaming Code
A CI/CD pipeline for streaming code needs checks a web service pipeline lacks: schema compatibility gates, topic configuration as code, ordered producer and consumer deploys, and a tested rollback path. Those are what actually break streaming systems, and automating them means they run no matter who is deploying.
Here's how to build that pipeline out step by step, so those checks run automatically instead of depending on whoever's deploying that day remembering to do them by hand.
Why should topic configuration live in code?
If topics, retention settings, and partition counts are created through a web console rather than a version-controlled configuration file, there's no history of what changed, no review step, and no way to reproduce the setup in a new environment. Move topic configuration into your infrastructure-as-code tooling alongside everything else, so a change to retention or partition count goes through the same pull request review as a code change.
This also gives you a single source of truth to diff against when something drifts, which happens more often than teams expect once more than one person has console access.
Put a schema compatibility gate before the merge, not after
Wire your schema registry's compatibility check into the pull request pipeline itself, so a breaking schema change fails CI before it can be merged, rather than being caught in a post-merge check that still lets the change reach a deploy queue. The earlier a breaking change fails, the cheaper it is to fix, since the author still has full context on what they were changing and why.
This is a small addition to most CI setups, usually a single step calling the registry's existing compatibility API, and it closes one of the most common gaps in streaming CI/CD.
In what order should producer and consumer deploys ship?
A producer that starts emitting a new field before any consumer expects it is usually harmless; a producer that removes a field before every consumer has been updated to stop relying on it is not. Build your deploy pipeline to enforce the right order explicitly for changes that need it: consumer changes that add tolerance for a new shape ship first, then the producer change, then any cleanup that removes old fallback logic.
Document this ordering requirement directly in the pull request or deploy ticket for any change that needs it, rather than relying on the team remembering the right sequence from a previous, similar change.
Automate the rollback path, not just the deploy path
Most CI/CD pipelines automate getting a change out and leave rollback as a manual, one-off action taken during an incident, which is exactly the wrong time to be improvising. Build an automated rollback path (reverting to the previous consumer group version, or restoring a previous topic configuration) that's been tested outside of an incident, so triggering it during one doesn't introduce a second problem on top of the first.
Match the pipeline's speed to your actual DORA cluster
Teams in the top DORA performance cluster ship on-demand deployments, often more than once a day, while teams in the lowest cluster can go as long as 180 days between releases1. If your CI/CD pipeline still requires a manual sign-off on every deploy, that ceiling is usually the actual limiter on how often you can safely ship, not anything about the code itself.
Work toward removing manual gates one at a time, starting with the lowest-risk deploys (routine, backward-compatible changes), and keep manual review specifically for the schema and state-changing changes that still genuinely need a human looking at them.
Give the pipeline its own test environment, not a shared one
Running CI against a shared staging broker means one team's test run can pollute another's: a topic filled with test messages, a schema registered under a name someone else was about to use, a consumer group that collides with a real one. Spin up an ephemeral broker per test run, seeded fresh each time, so CI runs are isolated from each other and from whatever manual testing is happening in shared staging at the same time.
This is more infrastructure to manage than a single shared environment, but it removes an entire class of flaky, hard-to-reproduce CI failures that come from test runs interfering with each other rather than from an actual bug in the code being tested.
Build the pipeline in this order:
- Move topic configuration, retention settings, and partition counts into version-controlled infrastructure code that goes through pull request review.
- Add the schema registry's compatibility check to the pull request pipeline so breaking changes fail before merge.
- Sequence deploys so consumers that tolerate a new shape ship before the producer change.
- Automate a rollback path and test it outside of an incident.
- Run CI against an ephemeral broker per test run instead of a shared staging broker.
What Good Looks Like
CI/CD is streaming-ready when topic configuration lives in version control, schema compatibility is checked before merge, and rollback has been automated and tested outside of an incident.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should topic creation always go through a pull request?
For anything beyond a quick local experiment, yes. A topic created through a console click with no review has no record of who created it, why, or what its retention and access settings should be, which is exactly the kind of gap that turns into an audit finding or an access review headache months later.
How do we know if our producer and consumer deploys are out of order?
The clearest sign is a spike in consumer errors or dead-letter volume immediately after a producer deploy, before any consumer code changed. That pattern almost always means the producer shipped a change no consumer was ready for, which is exactly what an explicit deploy sequencing step is meant to prevent.
Is it worth automating rollback if incidents are rare?
Yes, because rare incidents are exactly when an untested manual rollback is most likely to go wrong. Building and testing the automated path once, outside of an incident, costs far less than improvising a rollback under pressure with an on-call engineer who's never actually run it before.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Building a CI/CD Pipeline That Actually Catches Bugs
How to build a pipeline that blocks real regressions instead of just style errors, from test selection to what actually belongs as a merge gate.
Blue-Green, Canary, or Rolling: Deploying Stream Processors
A decision guide to rolling, blue-green, and canary deploys for stateful stream processors, plus the rollback plan most teams never actually test.
Building a Test Suite That Actually Catches a Bad Pipeline Change
A worked example of setting up schema, data quality, and contract tests for a streaming pipeline, so a bad change fails in CI instead of in production.
Verifying Every Service That Talks to Your Pipeline
Which parts of zero-trust verification to build and which to buy, so every producer and consumer on a streaming pipeline proves its identity.
Do You Actually Need Contract Tests for Your Event Streams?
Answers to the questions teams actually have about contract testing for event streams: what it catches that schema checks miss, and when to skip it.
What Your CI/CD Pipeline Actually Costs You
A way to think about CI/CD pipeline cost beyond the compute bill, including engineer waiting time, flaky test triage, and what to fix first.