Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

A Chaos Drill Walkthrough: From Hypothesis to Fixed Bug

A chaos drill that just confirms your system behaves the way you already expected teaches you nothing worth the effort. The useful ones are built around a specific, falsifiable guess about how something will fail, run somewhere you can safely be wrong, and followed by an actual fix. Here's what that looks like from start to finish, using one real class of drill as the example.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What makes a good chaos engineering hypothesis?

"Let's kill a node and see what happens" is not a hypothesis. "If the primary database connection pool is exhausted, we expect the application tier to shed load gracefully and recover within a set time once the pool frees up" is one, because it can be wrong in a specific, discoverable way. Write the hypothesis down before the drill, including what you expect to happen and what would count as a failure of that expectation, so the outcome isn't reinterpreted after the fact to sound like it went fine.

How large should a chaos drill's blast radius be?

Run the first version of any new drill in staging, or in production against a small, deliberately chosen slice of traffic with a fast rollback available, not against your full production fleet on the first attempt. The point of starting small is that your hypothesis about the blast radius is itself part of what you're testing; if the failure spreads further than expected even at small scale, you want to discover that with a rollback ready, not while you're mid-incident.

Inject the failure and watch the parts you didn't expect to watch

During the drill, watch not just the component you're directly testing but its neighbors: the services that depend on it, the ones it depends on, and anything sharing its infrastructure. A connection pool exhaustion test often reveals that a completely unrelated service sharing the same database instance also degrades, which is exactly the kind of hidden coupling a drill is supposed to surface. If nobody is watching beyond the target component, that discovery gets missed.

Write down what actually happened, including where you were wrong

Compare the outcome against the hypothesis explicitly: did the application tier shed load gracefully, did it recover within the expected time, and if not, exactly where did the behavior diverge from what was predicted. Resist the pull to write the finding as a success because nothing caught fire; a drill where recovery took three times longer than expected is a valuable finding even if the system technically survived, and writing it up honestly is what makes the next drill worth running.

Turn the finding into a tracked fix, not a Slack thread

The most common way chaos engineering stops delivering value is that findings live in a debrief document nobody revisits. File the specific gap the drill uncovered as a ticket with an owner and a deadline, the same as any other bug, and schedule a follow-up drill against the same hypothesis once the fix ships, to confirm the behavior actually changed rather than assuming the fix worked because it looked right in review.

Run each drill in this order:

  1. Write a falsifiable hypothesis, including what would count as a failure of your expectation.
  2. Run the first version in staging, or against a small slice of production with a fast rollback ready.
  3. Inject the failure and watch neighboring and dependent services, not just the component under test.
  4. Record the outcome against the hypothesis, including exactly where the system behaved differently than predicted.
  5. File each gap as a ticket with an owner and deadline, then rerun the drill once the fix ships.

Build a recurring calendar, not a one-time event

A single chaos drill tells you about the state of the system on one day. Recurring drills, run on a set schedule against a rotating set of hypotheses, catch regressions that a one-time exercise never will, since a fix that worked when it shipped can quietly stop working after an unrelated change six months later. Treat the drill calendar as part of the system's ongoing maintenance, not a special project that wraps up once the first round is done.

Keep the whole team in the loop, not just whoever's running it

A drill run silently by one engineer without telling the rest of the on-call rotation risks being mistaken for a real incident, or worse, having its findings dismissed because nobody else trusts what happened during an exercise they didn't know was happening. Announce the drill window ahead of time to whoever might reasonably be paged during it, and share the write-up afterward with the full team, not just whoever proposed the hypothesis, so the lesson actually spreads instead of staying with one person.

Executive Capability Standard

What Good Looks Like

A useful chaos drill starts from a specific, falsifiable hypothesis, runs at a blast radius the team can tolerate being wrong in, is documented honestly including where the outcome diverged from the prediction, and produces a tracked fix that gets re-tested once it ships.

Building The Capability (5-Stage Skill Ladder)

1. Learn:read through your last major incident and write the chaos-drill hypothesis that, if tested beforehand, would have caught it
2. Do Manually:run one small, manually orchestrated drill against a single dependency in staging and write up the outcome against your hypothesis
3. Delegate:give a specific engineer or rotation ownership of the drill calendar, separate from day-to-day feature work, so it doesn't quietly stop happening
4. Automate:schedule recurring drills against a rotating set of hypotheses and automatically file a ticket when an outcome diverges from the prediction
5. Buy:if you need to show an auditor that resilience testing actually happens on a schedule, a compliance platform like Vanta can track that evidence alongside your other controls

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Vanta

If a compliance framework you're pursuing expects documented, recurring resilience testing, a platform like Vanta is a reasonable place to keep that evidence alongside your other controls rather than in a folder of debrief documents.

Visit Vanta→

Frequently Asked Questions

How do we pick which failure to test first?

Start with the dependency whose failure you're least confident about and whose blast radius you can most easily contain, typically something with a clear owner and an existing rollback mechanism. Save wider, cross-team failure scenarios for once the team has practice running smaller drills safely.

Should chaos drills run in production?

Eventually, for the drills that matter most, since staging often can't reproduce real traffic patterns and scale. Start in staging or against a small, deliberately scoped slice of production traffic with a fast rollback ready, and only expand scope once you've built confidence in the drill and the rollback path.

What counts as evidence that a drill was worth running?

A written hypothesis, a documented outcome that says explicitly where the system matched or diverged from that hypothesis, and at least one concrete fix that came out of it. A drill that produces none of those three is closer to a fire drill than an engineering exercise.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides