Running Chaos Drills Without Breaking Production for Real
The fear that stops most teams from trying chaos engineering is reasonable: deliberately breaking something sounds like a bad way to spend a Tuesday. The fix isn't to skip it, it's to start small enough that a failed experiment teaches you something instead of paging the whole company.
This is a practical path from a first, low-risk drill to something closer to real production chaos testing, without needing a dedicated chaos engineering team to get started.
Start With a Failure You Already Understand
Pick a failure mode your team has already recovered from before, a single instance restart, a dependency timing out, a cache going cold, rather than inventing a novel scenario for your first drill. The point of an early drill is confirming your known recovery path actually works when triggered on purpose, not discovering an entirely new failure class.
Run it in staging first, even if your eventual goal is production chaos testing. A drill that reveals a scripting mistake in staging is a non-event; the same mistake triggered directly in production is its own incident.
Define the Blast Radius Before You Press Go
Decide, in writing, exactly what the experiment is allowed to touch: one service, one instance, one percentage of traffic, and nothing beyond that boundary. Have a specific person watching dashboards in real time with authority to abort immediately if the blast radius starts to spread past what was agreed.
An experiment without a defined stop condition isn't a controlled drill, it's just breaking something and hoping for the best.
Write the abort procedure down alongside the blast radius, not just the intention to abort if needed. Knowing exactly which command reverts the failure, and confirming it works before the drill starts, is what actually keeps a contained experiment contained if something doesn't go the way you expected.
Write these items down before starting a drill:
- The exact scope the experiment may touch: one service, one instance, or one slice of traffic, and nothing beyond it.
- The named person watching dashboards in real time, with authority to abort immediately.
- The specific command that reverts the failure, confirmed to work before the drill starts.
- The stop condition that tells the watcher the blast radius has spread past what was agreed.
What a Drill Is Actually Protecting
Chaos drills exist to confirm your system degrades gracefully within the downtime budget you've committed to, not to test whether it can survive anything. A team promising 99.9% availability has roughly 0.365 days, about nine hours, of downtime to spend across an entire year, and a drill's job is proving that a single dependency failure doesn't quietly eat a large slice of that budget in one afternoon1.
Framing the goal this way also tells you when to stop running bigger drills: once you've confirmed your known failure modes are contained within budget, further escalation should target new, specific risks, not general bravado.
Escalating to Real Production Testing
Move to production only after staging drills have run clean multiple times, and only against a small, clearly bounded percentage of real traffic at first. Schedule it for a low-traffic window, with the same real-time monitoring and abort authority as your staging drills, just with higher stakes if something goes wrong.
Tell the on-call team the drill is happening before it starts. A chaos experiment that pages an unsuspecting on-call engineer at 2 a.m. teaches the wrong lesson about what chaos engineering is for.
The Mistake: Treating a Drill Like a One-Time Event
A chaos drill run once, written up, and never repeated tells you your system handled a specific failure on a specific day. It says nothing about whether a code change three months later reintroduced the same fragility. Recurring drills, even simple ones repeated on a set cadence, catch regressions that a single celebrated success never will.
The goal isn't a growing list of dramatic experiments. It's a short list of realistic failure modes, tested regularly enough that a regression gets caught before a real incident finds it first.
Keep that list short on purpose. Five or six failure modes that map to your system's real dependencies, tested on a repeating schedule, teach you more over a year than a long backlog of exotic scenarios that only ever get run once for a demo.
For example, a team could pick three failure modes that match its real dependencies: a single instance restart, a dependency timing out, and a cold cache. It schedules each on a repeating cadence, records how long recovery took, and compares the result with the previous run. If a code change makes recovery slower, the drill catches it before a real incident does. A useful decision rule is to add a new failure mode only when a real incident, a new dependency, or an architecture change creates a risk the existing list does not cover.
What Good Looks Like
Good chaos engineering practice means every drill has a written blast radius, a defined stop condition, and a specific owner watching in real time, escalating from staging to a small, bounded slice of production only after known failure modes run clean repeatedly.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is chaos engineering too risky for a small engineering team to try?
Not if you start small. Running a known, already-understood failure mode against staging, with a defined blast radius and someone watching in real time, is low risk and still teaches you whether your recovery path actually works. Save production testing for after those early drills run clean repeatedly.
What failure should we test first in a chaos engineering drill?
Something your team has already recovered from before, like a single instance restart or a dependency timing out. The first drill should confirm a known recovery path works on purpose, not discover a brand new failure class, which is a much higher-risk exercise better suited for later.
How do we know when we're ready to run chaos experiments in production?
Move to production once staging drills for the same failure mode have run cleanly more than once. Start with a small, clearly bounded slice of real traffic during a low-traffic window, with real-time monitoring and someone able to abort immediately if the blast radius starts spreading.
Should the on-call team know a chaos drill is happening in advance?
Yes, always. Surprising an unsuspecting on-call engineer defeats the purpose of a controlled experiment and just creates a real incident with extra steps. Tell the team the drill's scope and timing beforehand, so any real response can be distinguished cleanly from the deliberate test.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Running Your First Chaos Engineering Drill Without Breaking Production
A practical way to run your team's first chaos engineering drill: small blast radius, a clear hypothesis, and a plan to stop it fast.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
Finding the Real Source of Latency in a Distributed System
A decision guide for narrowing down whether a slow request is a network problem, a database problem, a queue problem, or your own code.
A First Chaos Drill: What to Break, and How to Do It Safely
A step-by-step first chaos drill for small engineering teams, including how to pick a safe failure to inject and what to measure while it runs.
How to Run a Chaos Engineering Drill Without Causing a Real Outage
Chaos engineering works when it tests one hypothesis in a contained blast radius. Here is how to run a drill that produces a fix instead of a war story.