Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Running Chaos Drills Without Breaking Production for Real

The fear that stops most teams from trying chaos engineering is reasonable: deliberately breaking something sounds like a bad way to spend a Tuesday. The fix isn't to skip it, it's to start small enough that a failed experiment teaches you something instead of paging the whole company.

This is a practical path from a first, low-risk drill to something closer to real production chaos testing, without needing a dedicated chaos engineering team to get started.

Start With a Failure You Already Understand

Pick a failure mode your team has already recovered from before, a single instance restart, a dependency timing out, a cache going cold, rather than inventing a novel scenario for your first drill. The point of an early drill is confirming your known recovery path actually works when triggered on purpose, not discovering an entirely new failure class.

Run it in staging first, even if your eventual goal is production chaos testing. A drill that reveals a scripting mistake in staging is a non-event; the same mistake triggered directly in production is its own incident.

Define the Blast Radius Before You Press Go

Decide, in writing, exactly what the experiment is allowed to touch: one service, one instance, one percentage of traffic, and nothing beyond that boundary. Have a specific person watching dashboards in real time with authority to abort immediately if the blast radius starts to spread past what was agreed.

An experiment without a defined stop condition isn't a controlled drill, it's just breaking something and hoping for the best.

Write the abort procedure down alongside the blast radius, not just the intention to abort if needed. Knowing exactly which command reverts the failure, and confirming it works before the drill starts, is what actually keeps a contained experiment contained if something doesn't go the way you expected.

Write these items down before starting a drill:

  • The exact scope the experiment may touch: one service, one instance, or one slice of traffic, and nothing beyond it.
  • The named person watching dashboards in real time, with authority to abort immediately.
  • The specific command that reverts the failure, confirmed to work before the drill starts.
  • The stop condition that tells the watcher the blast radius has spread past what was agreed.

What a Drill Is Actually Protecting

Chaos drills exist to confirm your system degrades gracefully within the downtime budget you've committed to, not to test whether it can survive anything. A team promising 99.9% availability has roughly 0.365 days, about nine hours, of downtime to spend across an entire year, and a drill's job is proving that a single dependency failure doesn't quietly eat a large slice of that budget in one afternoon1.

Framing the goal this way also tells you when to stop running bigger drills: once you've confirmed your known failure modes are contained within budget, further escalation should target new, specific risks, not general bravado.

Escalating to Real Production Testing

Move to production only after staging drills have run clean multiple times, and only against a small, clearly bounded percentage of real traffic at first. Schedule it for a low-traffic window, with the same real-time monitoring and abort authority as your staging drills, just with higher stakes if something goes wrong.

Tell the on-call team the drill is happening before it starts. A chaos experiment that pages an unsuspecting on-call engineer at 2 a.m. teaches the wrong lesson about what chaos engineering is for.

The Mistake: Treating a Drill Like a One-Time Event

A chaos drill run once, written up, and never repeated tells you your system handled a specific failure on a specific day. It says nothing about whether a code change three months later reintroduced the same fragility. Recurring drills, even simple ones repeated on a set cadence, catch regressions that a single celebrated success never will.

The goal isn't a growing list of dramatic experiments. It's a short list of realistic failure modes, tested regularly enough that a regression gets caught before a real incident finds it first.

Keep that list short on purpose. Five or six failure modes that map to your system's real dependencies, tested on a repeating schedule, teach you more over a year than a long backlog of exotic scenarios that only ever get run once for a demo.

For example, a team could pick three failure modes that match its real dependencies: a single instance restart, a dependency timing out, and a cold cache. It schedules each on a repeating cadence, records how long recovery took, and compares the result with the previous run. If a code change makes recovery slower, the drill catches it before a real incident does. A useful decision rule is to add a new failure mode only when a real incident, a new dependency, or an architecture change creates a risk the existing list does not cover.

Executive Capability Standard

What Good Looks Like

Good chaos engineering practice means every drill has a written blast radius, a defined stop condition, and a specific owner watching in real time, escalating from staging to a small, bounded slice of production only after known failure modes run clean repeatedly.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read through your last two or three real incidents and identify which ones involved a failure mode simple enough to serve as a safe first drill.
2. Do Manually:Run a single, scripted staging drill for a known failure mode, with one engineer watching dashboards and a written stop condition agreed on beforehand.
3. Delegate:Give a rotating owner responsibility for scheduling recurring drills on a set cadence, so testing known failure modes doesn't stop once the first success is written up.
4. Automate:Script the drill itself, the failure injection, the monitoring checks, and the abort condition, so it can run repeatably without someone manually recreating the steps each time.
5. Buy:Bring in a chaos engineering platform once you're running drills often enough that scripting and coordinating each one manually is taking more time than the drills themselves.

How to Get Started

Frequently Asked Questions

Is chaos engineering too risky for a small engineering team to try?

Not if you start small. Running a known, already-understood failure mode against staging, with a defined blast radius and someone watching in real time, is low risk and still teaches you whether your recovery path actually works. Save production testing for after those early drills run clean repeatedly.

What failure should we test first in a chaos engineering drill?

Something your team has already recovered from before, like a single instance restart or a dependency timing out. The first drill should confirm a known recovery path works on purpose, not discover a brand new failure class, which is a much higher-risk exercise better suited for later.

How do we know when we're ready to run chaos experiments in production?

Move to production once staging drills for the same failure mode have run cleanly more than once. Start with a small, clearly bounded slice of real traffic during a low-traffic window, with real-time monitoring and someone able to abort immediately if the blast radius starts spreading.

Should the on-call team know a chaos drill is happening in advance?

Yes, always. Surprising an unsuspecting on-call engineer defeats the purpose of a controlled experiment and just creates a real incident with extra steps. Tell the team the drill's scope and timing beforehand, so any real response can be distinguished cleanly from the deliberate test.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides