Data Engineering & Real-Time Event StreamsPlaybook3 min readUpdated September 2026

Running Your First Chaos Drill on a Streaming Pipeline

Run your first chaos drill by choosing one failure your architecture should already survive, rehearsing it in a lower environment, and setting a blast radius and stop condition before you start. Most teams learn how their pipeline behaves during a broker failure only when one actually fails, and a drill lets you find out on your own schedule.

The hardest part of a first drill is rarely the tooling. It is convincing the team that deliberately breaking something in a controlled way is safer than continuing to not know what happens when it breaks on its own.

How do you pick the first failure for a chaos drill?

Start with a single, well-understood failure, such as killing one broker in a cluster that should tolerate the loss, rather than trying to simulate a full regional outage on your first attempt. The goal of the first drill is building the muscle of running one safely, not maximizing how much you learn from it.

A good first candidate is a failure you already believe your architecture handles gracefully. If it does not, you have found a real gap cheaply. If it does, you have confirmed an assumption instead of just hoping it holds.

Run it in a lower environment before production

Rehearse the exact drill in staging or a dedicated test cluster first, with production-like traffic if you can generate it, so you find tooling problems and unexpected side effects before you are running the same drill against real customer data. This step is easy to skip under time pressure and is exactly the step that prevents a chaos drill from becoming a real incident.

Only move to production once the staging run went cleanly and the team is confident in the rollback procedure if something goes differently than expected.

How do you set a blast radius and stop condition before a drill?

Decide in advance exactly which systems the drill is allowed to touch, and define a specific, measurable condition that means abort immediately, such as consumer lag crossing a threshold that would affect real customers. Write both down before the drill starts, not during it, since deciding limits in the moment tends to bias toward continuing rather than stopping.

Have one person whose only job during the drill is watching for the stop condition and calling it, separate from whoever is running the injection itself, so the decision to abort is not competing for attention with the mechanics of the drill.

Watch what people do, not only what the systems do

A chaos drill tests your team's response as much as it tests your infrastructure. Note how long it took someone to notice the failure, whether the right person got paged, and whether the runbook they reached for was actually current. These human-side gaps are often more valuable findings than anything about the infrastructure itself.

Run the drill without announcing the exact timing to the on-call engineer in some cases, once the team has enough experience with planned drills to make an unannounced one useful rather than just alarming.

Turn every finding into a tracked fix, not a story

A drill that surfaces three gaps and results in zero follow-up tickets has taught the team something without actually making the pipeline more resilient. Convert every finding into a specific, owned task with a deadline, the same way you would track a bug found in a real incident.

Schedule the next drill before the current one's findings are even closed out, so chaos engineering becomes a recurring practice with a cadence, rather than a one-time event the team did once and can point to.

Expand scope only after the basics are boring

Once single-broker failures reliably surface nothing new, that is the signal to widen scope, not a reason to stop running drills. Move toward failures that combine two problems at once, such as a broker loss during a deploy, or toward larger blast radii like a full availability zone outage, since those combinations are where the more subtle coordination gaps between teams tend to show up.

Resist the urge to jump straight to the most dramatic scenario, like a full regional failover, before the team has built confidence on smaller failures. A drill that is too ambitious too early tends to produce so many findings at once that none of them get properly tracked or fixed, which defeats the purpose covered in the previous step.

Put together, a first drill runs in this order:

  1. Choose a single well-understood failure, such as killing one broker in a cluster built to tolerate that loss.
  2. Rehearse the exact drill in staging or a test cluster, using production-like traffic where you can generate it.
  3. Write down the systems the drill may touch and the measurable condition that means abort immediately.
  4. Run the drill with a separate person watching the stop condition, and record how long people took to notice the failure.
  5. Turn every finding into an owned task with a deadline, and schedule the next drill before the current findings close.
Executive Capability Standard

What Good Looks Like

A useful chaos drill targets one well-defined failure at a time, gets rehearsed in a lower environment first, runs against a predefined blast radius and stop condition, and produces tracked follow-up fixes rather than just a story about what happened.

Building The Capability (5-Stage Skill Ladder)

1. Learn:List the failure modes your architecture claims to tolerate and pick the one you are least confident actually works.
2. Do Manually:Run a manual, single-broker failure drill in staging and document exactly what happened against what you expected.
3. Delegate:Assign an engineer to own the chaos engineering program, including scheduling drills and tracking findings to closure.
4. Automate:Use a chaos engineering tool to schedule and inject failures on a recurring cadence instead of running each drill manually.
5. Buy:Bring in outside reliability engineering expertise to design your first production drills if your team has never run one before.

How to Get Started

Frequently Asked Questions

Is it safe to run chaos engineering drills directly in production?

Only after you have rehearsed the exact drill in a lower environment and are confident in the rollback path. Start in staging, and move to production once the team has run the drill enough times to trust the tooling and the response.

How do we decide which failure to simulate first?

Pick a failure you believe your architecture already handles gracefully, such as losing one broker in a cluster built to tolerate it. That gives you either a cheap confirmation that your assumption holds or a real gap found safely, both of which are useful first outcomes.

Who should be involved in a chaos drill besides the engineer running it?

At minimum, someone whose only job is watching for the predefined stop condition, separate from whoever is running the injection. Including the on-call engineer who would normally respond to that failure is also valuable, since their real response is part of what the drill is testing.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides