Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Running Your First Chaos Engineering Drill Without Breaking Production

A chaos engineering drill is a small, controlled experiment that tests whether your system handles a failure the way you assume it does, before that failure happens for real. It's often mistaken for deliberately breaking production for fun, but done well it is the opposite: limited in scope, closely watched, and designed to teach.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What hypothesis should a chaos engineering drill start with?

Every drill should start with a specific, falsifiable claim: "if the payment provider times out, checkout falls back to the retry queue within five seconds." Not "let's see what happens if we kill this service." A hypothesis gives you a clear pass or fail, and a random failure just gives you a story.

Write the hypothesis down before you run anything. If you can't state what you expect to happen, you're not ready to run the experiment yet, you're ready to go read the code for that path first.

How small should the blast radius of a first drill be?

Your first drill should be able to fail safely: a single instance, a single non-critical service, during low-traffic hours, with a person watching and a clear stop button. Save the ambitious, wide-blast-radius experiments for after your team has run several small ones successfully.

This isn't caution for its own sake. A drill that goes wrong and causes a real customer-facing outage teaches your organization that chaos engineering is dangerous, and it will be much harder to get buy-in for the next one.

Size the Drill Against What You Can Actually Recover From

Your allowed downtime for the year is a real budget, and a drill that goes wrong spends from that same budget just like a real outage does1. Schedule drills with that in mind: don't run an aggressive experiment the same week you're already tight against your quarterly downtime allowance.

Teams that deploy often have an advantage here, because a small, contained failure is closer to their normal operating conditions and easier to recover from quickly than it is for a team that ships rarely and hasn't exercised its rollback path recently.

Watch the Right Signal, Not Just the One You're Testing

It's easy to focus entirely on the metric your hypothesis is about and miss that the experiment caused a different, unrelated problem elsewhere. Have someone watching overall system health, not just the specific path under test, and agree on the stop condition before you start, not while you're already mid-experiment and deciding under pressure.

For example, a team tests whether killing one worker instance triggers a clean restart. The hypothesis holds, but a second observer notices that latency for an unrelated job type climbed during the restart. Watching only the restart path would have shown a clean pass. Because someone was watching overall health, the team files a second finding about a shared connection pool. Give that finding an owner like any other, and add the pool behavior to the next drill's hypothesis. A pass on the tested path doesn't prove the experiment was harmless elsewhere.

Turn Every Drill Into a Fix, Not Just a Finding

A drill that reveals a gap and doesn't lead to a fix within a reasonable window teaches your team that the gap is acceptable. Track findings the same way you'd track a bug from a real incident, with an owner and a deadline, or the same failure mode will show up again the next time it's not a drill.

Getting Buy-In From a Team That's Never Done This

The biggest obstacle to a first drill usually isn't technical, it's convincing a skeptical team or a nervous founder that deliberately triggering a failure is safe. Bring evidence, not enthusiasm: walk through the small blast radius, the stop plan, and the specific hypothesis before asking for the go-ahead.

Running the very first drill on a low-stakes internal tool, something with no customer impact even in the worst case, is an easy way to build that trust before asking to test anything closer to revenue-critical systems.

Document the Drill Like You'd Document an Incident

A drill that produces a finding but no written record is a finding your team will rediscover the hard way in six months, during a real outage that feels oddly familiar. Write up what you tested, what you expected, what actually happened, and what changed as a result, the same format you'd use for a real incident review.

Over time, this record becomes genuinely useful: a growing list of verified assumptions about how your system behaves under specific failures, which is a far more concrete asset than a vague sense that the team has tried a few things.

A safe first drill follows this order:

  1. Write a specific, falsifiable hypothesis before running anything, such as how a fallback path should behave when a dependency times out.
  2. Pick a single instance or non-critical service during low traffic, with a person watching and a stop button ready.
  3. Agree on the stop condition in advance and assign someone to watch overall system health, not only the path under test.
  4. Check your remaining downtime budget so an experiment that goes wrong doesn't overspend it.
  5. Record what you tested, expected, observed, and changed, then give each finding an owner and a deadline.
Executive Capability Standard

What Good Looks Like

A real chaos drill tests a specific written hypothesis with a small blast radius, a clear stop condition agreed in advance, and a tracked fix for whatever it finds.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read through one recent real incident and write the hypothesis a chaos drill would have tested to catch it earlier.
2. Do Manually:Manually trigger a small, contained failure during a low-traffic window with someone watching and a clear stop plan.
3. Delegate:Assign a specific engineer to own the quarterly drill calendar and hypothesis list so it doesn't depend on one enthusiast.
4. Automate:Automate the safe, repeatable drills (a known instance failure, a known dependency timeout) once you've run them manually and trust the result.
5. Buy:Bring in outside expertise for your first few drills if no one on the team has run controlled failure testing before.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Is chaos engineering worth it for a small engineering team?

Yes, at a small scale. You don't need a dedicated chaos engineering platform to start, a single manually triggered failure during a low-traffic window with someone watching is a real drill. The value comes from finding gaps before a real outage does, not from the sophistication of the tooling.

How often should a small team run chaos drills?

Quarterly is a reasonable starting cadence for most small teams, timed to avoid your busiest traffic periods. The goal is steady practice and a growing list of verified assumptions, not a high frequency that turns into its own operational risk.

What's the most common mistake teams make on their first drill?

Starting too big: testing a critical, customer-facing path with a wide blast radius before the team has practiced on something smaller and lower-stakes. Start narrow, build confidence in your stop procedure, then expand scope over several drills.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides