How to Run a Chaos Engineering Drill Without Causing a Real Outage
Chaos engineering has a marketing problem: the name makes it sound like the point is to break things randomly and see what happens. Done well, it's closer to a fire drill than an act of chaos, a deliberate, contained test of a specific failure hypothesis, run with a rollback plan and someone watching the exits the entire time.
Done badly, it's an unannounced outage with extra steps. The difference between the two isn't the failure you inject, it's the discipline around containing it and following up on what it finds.
Why should a chaos drill start with a hypothesis?
A chaos drill without a specific hypothesis is just breaking something to see what happens, and it produces interesting stories more often than useful fixes. Start instead with a testable claim: if the primary database connection pool is exhausted, the service should degrade gracefully and shed non-critical requests rather than failing entirely.
The drill then either confirms that claim or, more usefully, disproves it in a contained setting instead of during a real incident where the same discovery costs a lot more than an afternoon of planned testing.
How do you pick a blast radius you can contain?
The blast radius, how much of production the failure can actually touch, should be the smallest slice that still tests the hypothesis honestly: a single instance behind a load balancer, a single non-critical service, a small percentage of traffic routed through a feature flag.
Expanding the blast radius should be earned by drills that went cleanly at a smaller scale, not assumed from the start because a bigger test feels more thorough. A contained drill that actually finishes teaches you more than an ambitious one that has to be aborted halfway through.
Document the blast radius in writing before the drill starts, not just in the head of the engineer running the injection. A written boundary, this load balancer, this single instance, this feature flag at a small traffic percentage, gives the person watching dashboards something concrete to check the failure against, and it gives the team something to compare afterward if the drill drifted wider than planned.
The Steady-State Metric That Tells You If It's Working
Before injecting anything, define the one or two metrics that represent normal operation, error rate, latency at a specific percentile, and confirm they're stable before you start the injection. Without a clear steady state to compare against, you can't tell whether a spike during the drill is the failure you intended or an unrelated problem happening to occur at the same time.
Write that baseline down before the drill begins, not from memory afterward. A remembered baseline has a way of shifting to match whatever the drill produced.
Pick a metric sensitive enough to move when the injected failure actually happens, but stable enough on a normal day that a false alarm doesn't waste the drill. Error rate at the specific endpoint you're testing usually works better than an aggregate, system-wide error rate, which can stay flat even while the one path you're testing degrades badly.
Running the Drill: Roles, Rollback, and a Kill Switch
Assign roles before the drill starts: someone injecting the failure, someone watching dashboards, and someone with the explicit authority to call it off immediately if the blast radius creeps wider than planned. The kill switch, whatever mechanism stops the injected failure instantly, should be tested on its own before the drill.
A kill switch you've never actually exercised is not something to trust mid-incident, and finding out it doesn't work cleanly is a much cheaper lesson to learn during a planned drill than during the drill itself going wrong.
A contained drill runs in this order:
- State a specific, testable hypothesis about how the system should behave when one component fails.
- Choose the smallest blast radius that still tests that hypothesis honestly.
- Confirm the steady-state metrics, such as error rate and latency, are stable before you inject anything.
- Assign three roles: one person injecting the failure, one watching dashboards and one with authority to call it off.
- Inject the failure, watch for the radius widening, and roll back immediately if it does.
Turning a Drill Into a Fix, Not Just a War Story
If you're targeting 99.99 percent uptime, you're defending about 52.6 minutes of downtime for the entire year1, which is exactly why a drill that surfaces a real gap and then doesn't get followed up on is worse than not running the drill at all: the gap is now documented, known, and still there.
Every drill should end with a short write-up and, for anything the hypothesis disproved, a ticket with an owner and a rough timeline, not just a retelling of what happened at the next team meeting.
Deciding When a Drill Is Actually Worth Repeating
Not every system needs a recurring drill on the same fixed cadence. A service you're actively rewriting is a poor candidate this month, since the failure mode you'd test is about to change anyway, while a service nobody has touched in two years can be more valuable to test precisely because nobody remembers exactly how it fails.
Revisit the list of hypotheses each quarter rather than running the same rotation indefinitely. A hypothesis the team has confirmed twice in a row is a candidate to retire from the regular schedule and replace with one you haven't tested yet, since repeating an already-confirmed result teaches the team less each time it runs.
What Good Looks Like
Good chaos engineering means a specific failure hypothesis, tested in a contained blast radius against a known steady-state baseline, with a tested kill switch and a follow-up ticket for anything it disproves.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How often should we run chaos drills?
Regularly enough that the muscle stays sharp, monthly is a reasonable cadence for most teams, and always after a significant architecture change that could have altered how a previously tested failure behaves. A drill run once and never repeated only tells you about the system as it existed on that one day.
Is it safe to run chaos drills in production?
It can be, once you've proven the pattern works in staging first and you've deliberately contained the blast radius. Production is where the real value comes from, since staging traffic rarely matches production load and dependency behavior closely enough to be fully trustworthy, but earn your way there with smaller, successful drills first.
What's a good first failure to test?
Pick something with a clear, specific hypothesis and a contained blast radius: killing a single non-critical service instance, or exhausting a connection pool for one low-traffic dependency. Save cross-region failover and full dependency outages for later drills, once the team has practiced the roles and rollback process on something smaller.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Running Your First Chaos Engineering Drill Without Breaking Production
A practical way to run your team's first chaos engineering drill: small blast radius, a clear hypothesis, and a plan to stop it fast.
Running Your First Chaos Drill on a Streaming Pipeline
A runbook for a first chaos engineering drill on a real-time streaming pipeline, from picking a safe failure to injecting it without causing a real one.
A First Chaos Drill: What to Break, and How to Do It Safely
A step-by-step first chaos drill for small engineering teams, including how to pick a safe failure to inject and what to measure while it runs.
Where Production Deployment Budgets Quietly Leak
The recurring places engineering teams overspend on production deployment architecture, and a practical order for fixing them without a full rebuild.
Running Your First Chaos Drill Without Breaking Prod
How to scope, run, and learn from a controlled failure drill without turning a resilience test into the real outage you were trying to prevent.
Running Chaos Drills Without Breaking Production for Real
A practical way to start chaos engineering drills, from picking a safe first failure to inject to deciding when a drill is ready to run in production.