Enterprise DevSecOps & Automated CompliancePlaybook3 min readUpdated September 2026

A First Chaos Drill: What to Break, and How to Do It Safely

The hardest part of chaos engineering isn't the tooling. It's picking a first failure small enough that a bad outcome is a minor incident, not a real outage, while still being realistic enough to teach the team something true about how the system actually behaves under stress. Here's a runbook for that first drill.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you pick a first failure with a known blast radius?

Start with something contained: killing a single non-primary instance behind a load balancer, or introducing latency on calls to a non-critical downstream service. Avoid anything touching your primary database or payment path for a first drill; the goal is building confidence in the process, not testing your worst-case scenario on day one.

Write down, before you start, exactly what you expect to happen. If you can't state a hypothesis, you're not ready to run the drill yet; you're just breaking something to see what happens, which teaches less and risks more. A useful test for whether a scenario is small enough for a first attempt: could you explain the worst plausible outcome to your CEO in one calm sentence, after the fact, without it sounding like an incident report.

Step 2: run it in a low-traffic window with a kill switch ready

Even a contained experiment deserves a low-traffic window and a fast way to stop it. Confirm, before injecting the failure, that whoever's running the drill can revert it within seconds, not minutes. A drill that can't be stopped quickly isn't a controlled experiment; it's an unplanned outage with extra steps.

Have a second person watching dashboards in real time, separate from whoever triggers the failure, so the person running the experiment isn't also the one who has to notice it's gone wrong.

Before injecting the failure, confirm each of these:

  • You have written down a hypothesis stating exactly what you expect to happen, so the drill tests something specific.
  • The drill is scheduled for a low-traffic window rather than during peak usage.
  • Whoever runs the drill can revert the failure within seconds, not minutes.
  • A second person is watching dashboards in real time, separate from the person injecting the failure.

Step 3: measure what you actually expected to break

Watch the specific metric your hypothesis predicted, not just the general dashboard. If the hypothesis was "the retry logic should absorb this without customer impact," the metric that matters is customer-facing error rate, not just the health of the instance you killed. A drill that confirms the instance died as expected but doesn't check downstream impact hasn't actually tested anything useful.

Capture the result in writing immediately afterward, while details are fresh: what you expected, what happened, and the gap between them if there was one.

Why fix the gap before running a bigger drill?

If the drill revealed a gap, close it before scaling up to a riskier scenario. Teams that skip this step and move straight to bigger, more ambitious chaos experiments are usually just discovering the same class of problem twice, at a higher cost the second time.

Only once a class of failure is handled cleanly, with the system recovering the way your hypothesis predicted, does it make sense to graduate to a more aggressive version: killing multiple instances at once, or injecting failure into a service that's actually on the critical path.

Building a regular cadence, not a one-time event

A single chaos drill proves the system handled one specific failure on one specific day. A regular cadence, even a modest one, catches the drift that happens as the system changes: a retry policy that got quietly removed during a refactor, a timeout that got misconfigured in a recent deploy. Run a drill on a fixed schedule, not just when someone remembers to, and treat a drill that suddenly fails after passing for months as a real regression worth investigating immediately.

Getting the team comfortable with deliberately breaking things

The cultural hurdle is often bigger than the technical one. Engineers who spend most of their time preventing failures can find it genuinely uncomfortable to cause one on purpose, even a small, contained one. Frame the first few drills explicitly as low-stakes and reversible, and share the results, including the gaps found, openly with the team rather than treating a discovered gap as something to quietly patch and move past.

A team that sees a chaos drill surface a real gap and watches that gap get fixed, credited to the drill, builds confidence in the practice faster than any amount of explaining why chaos engineering matters in the abstract.

Executive Capability Standard

What Good Looks Like

Good chaos engineering practice means every drill starts from a written hypothesis, runs with a fast kill switch, and gets measured against the specific outcome it predicted.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read through your system's existing retry and failover logic before your first drill, so your hypothesis about what should happen is grounded in what the code actually does.
2. Do Manually:Run a single-instance kill drill by hand during a low-traffic window, with a second engineer watching dashboards and a documented rollback step.
3. Delegate:Assign a rotating drill owner each month so the responsibility, and the resulting learning, doesn't sit with one person indefinitely.
4. Automate:Script the failure injection and the rollback so a drill can be triggered consistently without manual setup each time.
5. Buy:Bring in a dedicated chaos engineering platform once you're running drills often enough across enough services that manual scripting stops scaling.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Do we need a dedicated chaos engineering tool to start?

No. A first drill can be as simple as manually stopping an instance or adding artificial latency with a proxy configuration change. Dedicated tooling earns its cost once you're running drills regularly enough that manual setup becomes the bottleneck.

How do we get buy-in to run a drill against production instead of staging?

Start in staging if production buy-in isn't there yet, but be honest that staging traffic patterns rarely match production closely enough to catch every real failure mode. Use a successful staging drill as the evidence that earns trust for a small, contained production drill later.

What's a reasonable first cadence for chaos drills?

Monthly is a reasonable starting cadence for a small team: frequent enough to catch drift, infrequent enough that each drill gets proper attention rather than becoming a box-checking exercise nobody prepares for.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides