A First Chaos Drill: What to Break, and How to Do It Safely
The hardest part of chaos engineering isn't the tooling. It's picking a first failure small enough that a bad outcome is a minor incident, not a real outage, while still being realistic enough to teach the team something true about how the system actually behaves under stress. Here's a runbook for that first drill.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How do you pick a first failure with a known blast radius?
Start with something contained: killing a single non-primary instance behind a load balancer, or introducing latency on calls to a non-critical downstream service. Avoid anything touching your primary database or payment path for a first drill; the goal is building confidence in the process, not testing your worst-case scenario on day one.
Write down, before you start, exactly what you expect to happen. If you can't state a hypothesis, you're not ready to run the drill yet; you're just breaking something to see what happens, which teaches less and risks more. A useful test for whether a scenario is small enough for a first attempt: could you explain the worst plausible outcome to your CEO in one calm sentence, after the fact, without it sounding like an incident report.
Step 2: run it in a low-traffic window with a kill switch ready
Even a contained experiment deserves a low-traffic window and a fast way to stop it. Confirm, before injecting the failure, that whoever's running the drill can revert it within seconds, not minutes. A drill that can't be stopped quickly isn't a controlled experiment; it's an unplanned outage with extra steps.
Have a second person watching dashboards in real time, separate from whoever triggers the failure, so the person running the experiment isn't also the one who has to notice it's gone wrong.
Before injecting the failure, confirm each of these:
- You have written down a hypothesis stating exactly what you expect to happen, so the drill tests something specific.
- The drill is scheduled for a low-traffic window rather than during peak usage.
- Whoever runs the drill can revert the failure within seconds, not minutes.
- A second person is watching dashboards in real time, separate from the person injecting the failure.
Step 3: measure what you actually expected to break
Watch the specific metric your hypothesis predicted, not just the general dashboard. If the hypothesis was "the retry logic should absorb this without customer impact," the metric that matters is customer-facing error rate, not just the health of the instance you killed. A drill that confirms the instance died as expected but doesn't check downstream impact hasn't actually tested anything useful.
Capture the result in writing immediately afterward, while details are fresh: what you expected, what happened, and the gap between them if there was one.
Why fix the gap before running a bigger drill?
If the drill revealed a gap, close it before scaling up to a riskier scenario. Teams that skip this step and move straight to bigger, more ambitious chaos experiments are usually just discovering the same class of problem twice, at a higher cost the second time.
Only once a class of failure is handled cleanly, with the system recovering the way your hypothesis predicted, does it make sense to graduate to a more aggressive version: killing multiple instances at once, or injecting failure into a service that's actually on the critical path.
Building a regular cadence, not a one-time event
A single chaos drill proves the system handled one specific failure on one specific day. A regular cadence, even a modest one, catches the drift that happens as the system changes: a retry policy that got quietly removed during a refactor, a timeout that got misconfigured in a recent deploy. Run a drill on a fixed schedule, not just when someone remembers to, and treat a drill that suddenly fails after passing for months as a real regression worth investigating immediately.
Getting the team comfortable with deliberately breaking things
The cultural hurdle is often bigger than the technical one. Engineers who spend most of their time preventing failures can find it genuinely uncomfortable to cause one on purpose, even a small, contained one. Frame the first few drills explicitly as low-stakes and reversible, and share the results, including the gaps found, openly with the team rather than treating a discovered gap as something to quietly patch and move past.
A team that sees a chaos drill surface a real gap and watches that gap get fixed, credited to the drill, builds confidence in the practice faster than any amount of explaining why chaos engineering matters in the abstract.
What Good Looks Like
Good chaos engineering practice means every drill starts from a written hypothesis, runs with a fast kill switch, and gets measured against the specific outcome it predicted.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
A chaos drill that reveals a service recovers slower than expected is also worth checking against Tenable's exposure data, since a slow-recovering service is often also one that's gone longer than it should between patch cycles.
Alternative enterprise solution for scaling Enterprise DevSecOps: Chaos Engineering and Resilience.
Frequently Asked Questions
Do we need a dedicated chaos engineering tool to start?
No. A first drill can be as simple as manually stopping an instance or adding artificial latency with a proxy configuration change. Dedicated tooling earns its cost once you're running drills regularly enough that manual setup becomes the bottleneck.
How do we get buy-in to run a drill against production instead of staging?
Start in staging if production buy-in isn't there yet, but be honest that staging traffic patterns rarely match production closely enough to catch every real failure mode. Use a successful staging drill as the evidence that earns trust for a small, contained production drill later.
What's a reasonable first cadence for chaos drills?
Monthly is a reasonable starting cadence for a small team: frequent enough to catch drift, infrequent enough that each drill gets proper attention rather than becoming a box-checking exercise nobody prepares for.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Running Your First Chaos Engineering Drill Without Breaking Production
A practical way to run your team's first chaos engineering drill: small blast radius, a clear hypothesis, and a plan to stop it fast.
How to Run a Chaos Engineering Drill Without Causing a Real Outage
Chaos engineering works when it tests one hypothesis in a contained blast radius. Here is how to run a drill that produces a fix instead of a war story.
Running Your First Chaos Drill Without Breaking Prod
How to scope, run, and learn from a controlled failure drill without turning a resilience test into the real outage you were trying to prevent.
Running Your First Chaos Drill on a Streaming Pipeline
A runbook for a first chaos engineering drill on a real-time streaming pipeline, from picking a safe failure to injecting it without causing a real one.
Running Chaos Drills Without Breaking Production for Real
A practical way to start chaos engineering drills, from picking a safe first failure to inject to deciding when a drill is ready to run in production.
Running a Chaos Drill Against Your RAG Pipeline Without Breaking Production
A step-by-step guide to running chaos engineering drills against a production RAG and vector search pipeline, from picking a failure to reviewing results.