Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

Running Your First Chaos Drill Without Breaking Prod

Chaos engineering has a branding problem: it sounds like breaking things on purpose for the thrill of it. Done well, it's the opposite, a controlled, small, reversible test of a specific failure you already suspect your system might not handle gracefully, run at a time and scope you chose rather than waiting to find out during a real incident at 3 a.m.

The teams that get value from this run small drills often. The teams that get burned run one big, unscoped drill, cause a real outage, and rarely try again.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What should your first chaos drill test?

A useful first drill tests one concrete hypothesis: what happens if this specific database replica disappears, what happens if this specific downstream API starts timing out instead of failing fast. Vague goals produce vague results and make it hard to know afterward whether the drill actually proved anything. Write the hypothesis down before you start, something like: this system should fail over within thirty seconds if this dependency disappears, and design the drill to test exactly that claim.

How do you limit the blast radius of a chaos drill?

The single most important decision in any drill is how much of the system, and how many real users, it can possibly affect if something goes wrong. Start with the smallest scope that can still test the hypothesis: a single instance, a single low traffic region, a single internal service with no customer facing dependency. Only expand scope once smaller drills have run clean and the team trusts the abort mechanism. A drill that could plausibly take down the whole product is not a first drill, no matter how confident the team is going in.

Have a kill switch that works faster than the failure you're testing

Every drill needs a way to instantly stop the injected failure and return the system to normal, tested separately from the drill itself before it ever runs against anything real. If the kill switch takes longer to work than the failure takes to cause damage, the drill isn't actually safe, it's a real incident with extra steps. Practice pulling the kill switch on a system that isn't currently broken first, so the team knows exactly what it does and how fast it works before relying on it under pressure.

Run drills often enough that the muscle stays sharp

A single chaos drill run once a year tests the system as it existed a year ago, not as it exists today after a dozen deploys changed its behavior. Smaller, more frequent drills, monthly or even weekly for critical paths, catch regressions closer to when they were introduced and keep the team's on call response current. On-demand deploy teams in the fastest DORA cluster, the ones shipping every day, tend to treat drills as a routine part of that cadence, while teams stuck deploying once a month or less often let drills lapse for the same reason they let everything else slow down1. The drill schedule and the deploy schedule end up reinforcing each other, for better or worse.

Write down what actually happened, especially when it wasn't expected

The value of a drill isn't the drill itself, it's the gap between what you expected to happen and what actually happened. A drill that confirms the system behaved exactly as designed is a fine outcome, but a drill that surfaces a missing timeout, an alert that never fired, or a fallback path that was quietly broken is the one worth running again to confirm the fix. Document every drill's hypothesis, actual result, and any follow up work, in the same place every time, so patterns across drills are visible instead of scattered across individual engineers' memories.

A first drill runs in this order:

  1. Choose one concrete failure to test, such as a specific replica disappearing or one downstream API timing out.
  2. Set the smallest scope that can still test the idea: a single instance, a low traffic region or an internal service.
  3. Confirm the kill switch stops the injected failure quickly, testing it separately before the drill touches anything real.
  4. Run the drill and compare what actually happened against what you expected, noting every surprise.
  5. Document the findings, fix gaps such as a missing timeout or a silent alert, and rerun to confirm the fix.

The most common mistake: running a drill as a demo instead of a test

A drill scheduled for a leadership review, with a known outcome and a rehearsed narrative, teaches the team almost nothing, because everyone already knows what's supposed to happen and quietly makes sure it does. The whole value of a drill comes from genuine uncertainty about the outcome. If your team can predict exactly how a drill will go before running it, either the hypothesis is too safe to be useful or the drill has drifted into theater. Keep drills small enough that they're routine, not performances, and resist the urge to polish one into a showcase the first time someone senior wants to watch.

Executive Capability Standard

What Good Looks Like

Good chaos engineering means one specific, written hypothesis per drill, a scoped blast radius with a tested kill switch, and documented results, run often enough that regressions get caught close to when they're introduced.

Building The Capability (5-Stage Skill Ladder)

1. Learn:List three specific failure scenarios your team suspects the system might not handle gracefully, and pick the smallest, safest one to test first.
2. Do Manually:Manually kill one low traffic instance during business hours and time how long it takes the system to recover on its own.
3. Delegate:Assign an owner for the drill calendar and results log so drills happen on a schedule and findings don't get lost between them.
4. Automate:Once manual drills run clean, automate the failure injection and kill switch so drills can run more frequently with less manual setup.
5. Buy:Bring in outside expertise to design the first drill program if nobody on the team has run a scoped chaos test before.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

ClickUp

logging each drill's hypothesis, scope, and result in a tool like ClickUp turns scattered one off tests into a pattern the team can actually learn from

Visit ClickUp→

Frequently Asked Questions

Is chaos engineering only useful once we're at a large scale?

No, small companies benefit even more in some ways, since a single failure is more likely to be catastrophic when there's no redundancy to absorb it. A useful first drill for a small team can be as simple as killing one instance during low traffic hours and confirming the system recovers the way you expect. Scope and frequency should match your team's size, the underlying discipline doesn't require scale to be worth doing.

How do we convince leadership this is worth the risk?

Frame the first drill by its scope and its kill switch, not by the word chaos. A drill scoped to one low traffic instance with a tested rollback is a controlled experiment, not a risk to the business, and showing the specific blast radius and abort plan usually resolves the concern faster than a general explanation of the practice.

What's the most common mistake teams make with their first drill?

Scoping it too broadly because the team wants to prove something impressive on the first try. A first drill that could plausibly cause a real customer facing outage is testing the team's nerve, not the system's resilience. Start smaller than feels necessary, and only widen scope once the smallest version has run clean more than once.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides