Setting Headroom Targets So Traffic Spikes Don't Take You Down
Most capacity conversations start with an average: average requests per second, average database load, average queue depth. Averages are the wrong number to plan around, because outages rarely happen on an average day. They happen the day a promotional email lands in tens of thousands of inboxes at once, or the day a batch job and a traffic spike collide.
Taj, MeetMyCTO's AI CTO, sees the same gap across founder-led engineering teams again and again: everyone agrees headroom matters, but almost nobody has written down how much headroom is enough, for which part of the system, or who checks it on a schedule. This guide walks through setting that number layer by layer and keeping it current.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How much capacity headroom do you need? Size to your worst day
Pull the last two or three traffic spikes your system actually survived, or barely survived, and use the ratio between that peak and your normal daily average as your burst multiplier. Say your typical afternoon runs at a certain steady load and your worst spike hit four or five times that, plan your headroom around that ratio, not around a plan built on averages that assumes traffic barely moves.
Do this separately for each layer, because they don't scale together. A stateless API layer might ride out a large spike fine while its database connection pool exhausts far sooner, because connections are a hard ceiling that autoscaling web servers alone won't fix.
Give Compute, Database, Queue, and Egress Their Own Targets
Compute headroom is the easiest to reason about because it's usually elastic: add instances, add pods. Database headroom is harder, because connection limits, replica lag, and disk IOPS all cap out at different points, and none of them respond to autoscaling the way a web server does. Queue depth needs its own target too: a queue that's merely long isn't the same problem as a queue that's growing faster than consumers can drain it.
Egress and third-party API limits are the layer teams forget entirely. A payment processor's rate limit or a partner API's request cap doesn't care how much headroom you built into your own infrastructure.
Say your steady-state database connection pool usually runs comfortably below its limit and jumps close to it during your worst recorded spike, while your compute layer barely moves under the same event. That's two very different headroom pictures living inside one dashboard that looks green from the top. Treat each layer's number on its own terms instead of assuming one healthy-looking chart means the whole stack has room to spare.
When should capacity alerts fire? Before the spike, not during it
How aggressively you alert should track your availability target. A team promising 99.9% uptime is working with a downtime budget of about 0.365 days a year, roughly nine hours, so a capacity alert that only fires once you're already degraded is spending directly out of that budget1.
Pick a sustained-load threshold, a resource pool holding above a set level for several consecutive minutes, rather than alerting on any single spiky data point, or you'll train your on-call rotation to ignore the pages.
Put Capacity Reviews on a Fixed Calendar
A headroom target set once and never revisited goes stale the moment your traffic pattern changes, whether that's a new integration, a new customer segment, or a feature that suddenly gets heavy use. Put a short capacity review on the calendar, monthly for a fast-growing product, quarterly for a stable one, and treat a missed review the same way you'd treat a missed security patch.
Tie the review to real release and growth planning, not just to whoever remembers to look at a dashboard. If engineering leadership already meets to plan the roadmap, a two-minute capacity check belongs on that same agenda.
A short capacity review can cover these checks:
- Compare your latest traffic spikes against your normal daily average, and update the burst multiplier if that ratio has shifted.
- Check database limits separately, including connection limits, replica lag, and disk IOPS, since none of them respond to autoscaling the way web servers do.
- Review queue depth against its own target instead of folding it into the compute numbers.
- Confirm autoscaling ceilings, such as maximum instance counts and downstream vendor rate limits, still match the peak you expect.
- Note any new integration, customer segment, or heavily used feature that has changed the traffic pattern since the last review.
Headroom Is Not the Same Thing as Redundancy
Redundancy protects you when one instance or node fails, by having another one ready to take over. Headroom protects you when total demand grows, by having spare capacity across the whole system. Those are different problems, and it's easy to confuse them.
A fully redundant setup where every replica is already running close to its own limit can still fall over the moment a spike hits, because redundancy alone doesn't add capacity, it just adds a spare copy of whatever capacity you already had.
The pattern shows up most often right after a team adds a second availability zone or a standby replica and quietly stops worrying about capacity, treating the redundancy work as if it also solved the headroom question. It didn't: a second copy of a system that's already near its ceiling just means you now have two systems near their ceiling instead of one.
What Good Looks Like
Good capacity planning means every layer of your stack, compute, database, queue, and network, has a documented headroom target based on your worst realistic traffic day, checked on a fixed schedule instead of after an outage.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How much extra infrastructure capacity should we keep in reserve?
There's no universal safe percentage, because it depends on your own burst ratio, not an industry rule of thumb. Compare your worst survived traffic spike with your normal average, use that ratio as your headroom target for each layer, and revisit it whenever your traffic pattern shifts, like a new integration or a large marketing push.
Does autoscaling mean we don't need to plan capacity manually?
Autoscaling extends how far your headroom stretches, but it doesn't remove the need to plan it. Autoscaling has ceilings too, a maximum instance count, a database connection limit, or a rate limit from a downstream vendor, and a person still has to set and periodically raise those ceilings.
How do we plan capacity for a launch or a big marketing push?
Treat it like your worst traffic day, but planned in advance instead of discovered live. Estimate the likely spike from the campaign's expected reach, load-test that estimate against staging if you can, and raise autoscaling ceilings and connection limits ahead of time rather than letting autoscaling react during the event itself.
What's the actual difference between headroom and redundancy?
Redundancy is about surviving a failure: another instance is ready if one goes down. Headroom is about surviving growth: there's spare capacity across the system when total demand rises. A system can have full redundancy and still run out of headroom if every replica is already near its ceiling.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Sizing Platform Capacity Around How Often Your Team Ships
A way to size infrastructure headroom against your traffic pattern, deploy cadence, and uptime target, instead of picking a round percentage and hoping.
How to Build an Infrastructure Headroom Worksheet Before You Need One
A worksheet-based way for CTOs to track infrastructure headroom by service, so capacity decisions happen before an outage forces them.
How Much Infrastructure Headroom Is Actually Enough?
Capacity planning usually means reacting to a page instead of a forecast. Here is how to pick a headroom target and spot your next constraint before it hits.
How Much Cloud Headroom Should You Actually Keep?
A practical way to size compute headroom against real traffic spikes, so engineering isn't paying for capacity it never uses or scrambling when demand jumps.
A Capacity Planning Runbook for Teams Tired of Fire Drills
A concrete way to set headroom targets, watch the right leading indicators, and decide what to pre-provision before the next launch catches you flat.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.