Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

How to Run a Chaos Engineering Drill Without Causing a Real Outage

Chaos engineering has a marketing problem: the name makes it sound like the point is to break things randomly and see what happens. Done well, it's closer to a fire drill than an act of chaos, a deliberate, contained test of a specific failure hypothesis, run with a rollback plan and someone watching the exits the entire time.

Done badly, it's an unannounced outage with extra steps. The difference between the two isn't the failure you inject, it's the discipline around containing it and following up on what it finds.

Why should a chaos drill start with a hypothesis?

A chaos drill without a specific hypothesis is just breaking something to see what happens, and it produces interesting stories more often than useful fixes. Start instead with a testable claim: if the primary database connection pool is exhausted, the service should degrade gracefully and shed non-critical requests rather than failing entirely.

The drill then either confirms that claim or, more usefully, disproves it in a contained setting instead of during a real incident where the same discovery costs a lot more than an afternoon of planned testing.

How do you pick a blast radius you can contain?

The blast radius, how much of production the failure can actually touch, should be the smallest slice that still tests the hypothesis honestly: a single instance behind a load balancer, a single non-critical service, a small percentage of traffic routed through a feature flag.

Expanding the blast radius should be earned by drills that went cleanly at a smaller scale, not assumed from the start because a bigger test feels more thorough. A contained drill that actually finishes teaches you more than an ambitious one that has to be aborted halfway through.

Document the blast radius in writing before the drill starts, not just in the head of the engineer running the injection. A written boundary, this load balancer, this single instance, this feature flag at a small traffic percentage, gives the person watching dashboards something concrete to check the failure against, and it gives the team something to compare afterward if the drill drifted wider than planned.

The Steady-State Metric That Tells You If It's Working

Before injecting anything, define the one or two metrics that represent normal operation, error rate, latency at a specific percentile, and confirm they're stable before you start the injection. Without a clear steady state to compare against, you can't tell whether a spike during the drill is the failure you intended or an unrelated problem happening to occur at the same time.

Write that baseline down before the drill begins, not from memory afterward. A remembered baseline has a way of shifting to match whatever the drill produced.

Pick a metric sensitive enough to move when the injected failure actually happens, but stable enough on a normal day that a false alarm doesn't waste the drill. Error rate at the specific endpoint you're testing usually works better than an aggregate, system-wide error rate, which can stay flat even while the one path you're testing degrades badly.

Running the Drill: Roles, Rollback, and a Kill Switch

Assign roles before the drill starts: someone injecting the failure, someone watching dashboards, and someone with the explicit authority to call it off immediately if the blast radius creeps wider than planned. The kill switch, whatever mechanism stops the injected failure instantly, should be tested on its own before the drill.

A kill switch you've never actually exercised is not something to trust mid-incident, and finding out it doesn't work cleanly is a much cheaper lesson to learn during a planned drill than during the drill itself going wrong.

A contained drill runs in this order:

  1. State a specific, testable hypothesis about how the system should behave when one component fails.
  2. Choose the smallest blast radius that still tests that hypothesis honestly.
  3. Confirm the steady-state metrics, such as error rate and latency, are stable before you inject anything.
  4. Assign three roles: one person injecting the failure, one watching dashboards and one with authority to call it off.
  5. Inject the failure, watch for the radius widening, and roll back immediately if it does.

Turning a Drill Into a Fix, Not Just a War Story

If you're targeting 99.99 percent uptime, you're defending about 52.6 minutes of downtime for the entire year1, which is exactly why a drill that surfaces a real gap and then doesn't get followed up on is worse than not running the drill at all: the gap is now documented, known, and still there.

Every drill should end with a short write-up and, for anything the hypothesis disproved, a ticket with an owner and a rough timeline, not just a retelling of what happened at the next team meeting.

Deciding When a Drill Is Actually Worth Repeating

Not every system needs a recurring drill on the same fixed cadence. A service you're actively rewriting is a poor candidate this month, since the failure mode you'd test is about to change anyway, while a service nobody has touched in two years can be more valuable to test precisely because nobody remembers exactly how it fails.

Revisit the list of hypotheses each quarter rather than running the same rotation indefinitely. A hypothesis the team has confirmed twice in a row is a candidate to retire from the regular schedule and replace with one you haven't tested yet, since repeating an already-confirmed result teaches the team less each time it runs.

Executive Capability Standard

What Good Looks Like

Good chaos engineering means a specific failure hypothesis, tested in a contained blast radius against a known steady-state baseline, with a tested kill switch and a follow-up ticket for anything it disproves.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Write down one specific, testable failure hypothesis for a system your team already suspects is fragile.
2. Do Manually:Run a small, manually triggered failure injection in staging against that hypothesis and record the steady-state metrics before and after.
3. Delegate:Assign an engineer ownership of the drill calendar and the write-up process for each run.
4. Automate:Build a repeatable, scheduled drill for the failure modes you've already validated are safe to test regularly.
5. Buy:Bring in a consultant experienced in chaos engineering if your team has never run a drill and wants a safer first attempt guided by someone who has.

How to Get Started

Frequently Asked Questions

How often should we run chaos drills?

Regularly enough that the muscle stays sharp, monthly is a reasonable cadence for most teams, and always after a significant architecture change that could have altered how a previously tested failure behaves. A drill run once and never repeated only tells you about the system as it existed on that one day.

Is it safe to run chaos drills in production?

It can be, once you've proven the pattern works in staging first and you've deliberately contained the blast radius. Production is where the real value comes from, since staging traffic rarely matches production load and dependency behavior closely enough to be fully trustworthy, but earn your way there with smaller, successful drills first.

What's a good first failure to test?

Pick something with a clear, specific hypothesis and a contained blast radius: killing a single non-critical service instance, or exhausting a connection pool for one low-traffic dependency. Save cross-region failover and full dependency outages for later drills, once the team has practiced the roles and rollback process on something smaller.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides