A Capacity Planning Runbook for Teams Tired of Fire Drills
Capacity planning means setting an explicit headroom target for each service, watching the indicators that lead toward it, and deciding in advance what you pre-provision. Most incidents come from nobody owning a number: a service creeps toward its ceiling for weeks while the dashboard shows it, until a launch or seasonal spike pushes it over.
This is a runbook for setting that rule, not a survey of metrics you could theoretically watch.
Pick a headroom target before you need one
Every service that can fail under load needs an explicit ceiling: a level of CPU, memory, connection pool usage, or queue depth past which you consider it in trouble. Say you decide no stateful service should sustain more than 70% memory utilization for longer than 15 minutes. That number becomes the thing people check, not "does it feel slow." Pick the number with the team that owns the service, write it down next to the service's runbook, and revisit it after any architecture change. A target picked in a meeting six months ago and never revisited is worse than no target, because it gives false confidence.
Which capacity indicators lead and which ones lag?
CPU and memory are lagging indicators: by the time they are pinned, you are already in the incident. Queue depth, connection pool wait time, and the growth rate of your busiest table or index are leading indicators, because they tell you where you are heading before you arrive. For a queue-backed service, a queue that is growing faster than it drains for more than a few minutes is a better signal than any percentage. Build your alerting around the trend, not just the threshold. Say a service is climbing toward its ceiling a little more each day; that trend deserves a ticket even while the service is still comfortably under the line, because by the time it crosses the line the fix usually takes longer than the warning gave you time for. A dashboard that only turns red at the ceiling is a dashboard that tells you about the incident after it has already started.
What scales in minutes and what scales in weeks?
Stateless compute behind an autoscaler can absorb a spike in minutes. A database's storage volume, a message broker's partition count, a third-party API's rate limit, and a cloud provider's quota for a given instance family cannot. Make a short list, per service, of which dependencies fall into the slow category, and treat a request to raise one of those limits as something you file weeks ahead of a known event, not the week before. If you cannot tell how long a given increase takes, ask the vendor or your cloud account team directly instead of assuming it is instant. Quota increases in particular are an easy thing to forget, because they sit outside your own codebase and rarely show up in a normal deployment checklist; the team that remembers them is usually the team that got burned by one once and wrote it down afterward so nobody had to learn it twice.
Run the numbers before a launch, not during it
Before any event you expect to change traffic, work through the arithmetic on paper first. Say your current peak is a known number of requests per second and marketing is planning a push that historically multiplies traffic for a week. Multiply your peak by that factor, check it against every layer the request touches (load balancer, application tier, cache, database, any rate-limited third party), and find the layer that breaks first. That is your bottleneck, and it is usually not the layer everyone assumed; teams tend to assume the database will be the constraint and are often surprised to find a third-party API's rate limit or a cache eviction policy gets there first. Fix that one layer, then redo the math, rather than adding capacity everywhere evenly, since evenly-distributed extra capacity is usually a more expensive way to solve a problem that lives in one specific place.
Decide, per layer, what you pre-provision and what you autoscale
Autoscaling is right for capacity you can add and remove in minutes with no state to carry: web and application tiers, most background workers. Pre-provisioning is right for anything with a cold-start cost, a contractual minimum, or state that cannot simply appear on demand: database read replicas, cache cluster size, reserved capacity you have committed to for cost reasons. Write this decision down per layer once, so the next person planning a launch is not re-litigating whether the database read replica should "just autoscale." It cannot, and pretending otherwise is how launches go wrong late at night when the on-call engineer discovers, mid-incident, that the thing they assumed would scale itself needed a change request filed days earlier.
Sort each layer into one of these two groups:
- Autoscale web and application tiers and most background workers, since capacity can be added and removed in minutes with no state to carry.
- Pre-provision database read replicas and cache cluster size, where state or cold-start cost means capacity cannot simply appear on demand.
- Pre-provision anything under a contractual minimum, such as reserved capacity you have committed to for cost reasons.
- Start requests to raise slow-scaling limits, like storage volume, partition count, API rate limits and instance quotas, well ahead of the launch.
Write the plan down somewhere the next launch will actually find it
A capacity plan that lives only in one engineer's memory disappears the day that engineer is on vacation or has moved to a different team. Keep a short, living document per service: its headroom target, its slow-to-expand dependencies, and the layer that broke first the last time someone ran the arithmetic. Update it after every launch, whether or not anything went wrong, since a launch that went smoothly still tells you something useful about where the next bottleneck is likely to be.
What Good Looks Like
Good capacity planning means every service that can fail under load has an explicit headroom target, an owner who checks it against real usage on a set schedule, and a written answer for which of its dependencies can autoscale versus which need lead time to expand.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How much headroom should a service actually carry?
There is no single right number, which is exactly why picking one explicitly matters more than picking the "correct" one. Start conservative for anything customer-facing, tighten it once you have a few months of real usage data, and treat the target as a living decision the owning team revisits after load tests or architecture changes, not a constant.
Is autoscaling the same thing as capacity planning?
No. Autoscaling handles short-term variation within a range you have already provisioned for. Capacity planning decides what that range is, which dependencies cannot autoscale at all, and how far ahead you need to request more of them. Teams that treat autoscaling as a substitute for planning usually discover the gap at a database or a vendor rate limit.
How do we plan for a launch whose traffic we can't predict well?
Plan for the highest multiple of current traffic you can defend, then find the first layer that breaks at that level. Give that layer a specific fallback: a queue to absorb overflow, a feature flag to shed non-essential work, or a rate limit that degrades gracefully instead of failing outright. Repeat the exercise after each launch, since the layer that breaks first can change.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Sizing Platform Capacity Around How Often Your Team Ships
A way to size infrastructure headroom against your traffic pattern, deploy cadence, and uptime target, instead of picking a round percentage and hoping.
How to Build an Infrastructure Headroom Worksheet Before You Need One
A worksheet-based way for CTOs to track infrastructure headroom by service, so capacity decisions happen before an outage forces them.
Sizing GPU Headroom So Your Inference Cluster Doesn't Choke
How to size spare GPU capacity for a model serving cluster: set a headroom floor, know when autoscaling helps, and weigh what extra capacity costs.
How Much Infrastructure Headroom Is Actually Enough?
Capacity planning usually means reacting to a page instead of a forecast. Here is how to pick a headroom target and spot your next constraint before it hits.
How Much Cloud Headroom Should You Actually Keep?
A practical way to size compute headroom against real traffic spikes, so engineering isn't paying for capacity it never uses or scrambling when demand jumps.
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.