Enterprise DevSecOps & Automated CompliancePlaybook3 min readUpdated September 2026

How to Build an Infrastructure Headroom Worksheet Before You Need One

Most capacity problems don't show up as a slow, visible climb. They show up as a page that used to load in 200 milliseconds now taking four seconds during a traffic spike nobody modeled for. The fix isn't a bigger budget line for infrastructure. It's a worksheet you actually keep current: one row per service, its current utilization, its growth rate, and the date it runs out of room.

This is a plan you can build in an afternoon and revisit every sprint. It doesn't require a platform purchase or a consultant. It requires picking the right five columns and refusing to let the sheet go stale.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

The five columns that matter

Track, per service: current peak utilization (CPU, memory, connection pool, or whatever the actual bottleneck is), the trailing 90-day growth rate, the hard ceiling (instance type limit, database connection cap, queue depth where latency falls apart), the lead time to add capacity, and the date you cross most of the ceiling at the current growth rate, say four-fifths of it. That last column is the one that turns a spreadsheet into an early-warning system.

Don't track every service. Track the ones with a real ceiling: your primary database, your job queue, your cache layer, anything with a hard connection or IOPS limit. A stateless API server behind autoscaling doesn't need a row; it needs a max instance count that's high enough to be a non-issue.

Each service row on the worksheet needs these five fields:

  • Current peak utilization, measured on whatever resource is the real bottleneck, such as CPU, memory, or a connection pool.
  • The trailing three-month growth rate, so the sheet reflects how fast the service is actually filling up.
  • The hard ceiling: an instance type limit, a database connection cap, or the queue depth where latency falls apart.
  • The lead time to add capacity, using the slowest realistic path rather than the fastest one.
  • The date you cross about four-fifths of the ceiling at the current growth rate, which is the early-warning column.

Set the lead time honestly

The lead time column is where most worksheets lie. If provisioning a bigger database instance takes fifteen minutes but re-sharding a table that's outgrown its instance class takes three weeks of migration work, the lead time for that service is three weeks, not fifteen minutes. Write down the slowest realistic path, including approvals, vendor lead time on reserved capacity, and the time to test the change in staging.

Once lead time is honest, the warning date writes itself: it's the date you hit the ceiling, minus the lead time, minus a buffer you're comfortable defending to your CEO if you're wrong. A team that ships a warning date without that buffer is really just moving the outage a few weeks later and calling it planning.

Where the budget actually leaks

The common leak isn't underprovisioning, it's overprovisioning that nobody revisits. A service gets sized for a launch spike, the spike passes, and the instance count never comes back down because nobody owns that decision. Put an owner's name next to every row. Say a service has been running under a third of its capacity for two full growth cycles: that's a resize ticket, not a monitoring dashboard entry.

The other leak is reserved capacity bought for a growth curve that didn't happen. Reserved instances and committed-use discounts are worth it only when your 90-day growth rate is stable enough to trust a one-year commitment. If a service's growth rate swings by more than half between quarters, keep it on-demand until the curve settles. Locking in a discount on a number you don't trust turns a pricing decision into a bet you didn't mean to make.

Reviewing the worksheet without it becoming busywork

A fifteen-minute review at the start of each sprint, covering only rows whose warning date moved closer, keeps this from turning into a chore nobody does. Assign the review to whoever's on call that week; they already have the freshest picture of what's tight. Anything that crosses the threshold in the next two lead-time windows becomes a ticket, not a conversation.

When a row's warning date keeps slipping closer every sprint, that's a signal the growth-rate estimate is wrong, not that you need a bigger buffer. Recompute the trailing rate on real numbers instead of adding padding on top of a stale one.

Treat a near-miss as a process bug, not bad luck

If a service ever crosses its hard ceiling before the worksheet flagged it, the fix isn't a bigger safety margin bolted onto the existing process. Find the specific step that failed: was the growth rate stale, was the lead time underestimated, or did nobody own the row. Fixing that specific step keeps the worksheet trustworthy; padding every number with extra margin just delays the next miss and makes the sheet harder to read.

Keep a short log of near-misses next to the worksheet itself. A pattern of misses in the same column, across different services, usually points at one bad assumption baked into how that column gets filled in, and that's worth fixing once rather than patching service by service.

Executive Capability Standard

What Good Looks Like

Good capacity planning means every service with a hard ceiling has a named owner, a current growth rate, and a warning date that triggers action before the lead time to fix it runs out.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your cloud provider's documentation on the specific limits your primary database and queue actually hit (connection caps, IOPS, partition counts) so the ceiling column reflects real limits, not guesses.
2. Do Manually:Build the five-column worksheet in a spreadsheet and update it by hand each sprint from your monitoring dashboards.
3. Delegate:Hand the sprint review to your on-call engineer with a standing fifteen-minute slot and a rule for what crosses into a ticket.
4. Automate:Pull utilization and growth-rate numbers into the worksheet with a scheduled query against your monitoring API instead of copying numbers by hand.
5. Buy:Bring in a cloud cost and capacity platform once you have more than a handful of services with hard ceilings and the manual worksheet stops keeping up.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Tenable

When a capacity ceiling review turns up an old, forgotten instance still running, that's also the moment to check whether it's still patched; Tenable's exposure scans catch the two problems with the same inventory pass.

Visit Tenable→

Frequently Asked Questions

How far out should the warning date look?

Set it at your longest lead time plus one full review cycle. If provisioning takes three weeks and you review every two weeks, your warning threshold should trigger at least five weeks before the ceiling, so a missed review cycle still leaves room to act.

Should this worksheet include cost, not just capacity?

Add a monthly cost column once the capacity columns are stable and trusted. Mixing cost review into the first version usually turns a capacity tool into a budget argument, and the ceiling dates stop getting the attention they need.

What if a service doesn't have a clean utilization metric?

Use the closest proxy that predicts failure: queue depth for a job processor, connection count for a database, open file descriptors for anything that leaks them. A rough proxy tracked consistently beats a precise metric nobody agrees on.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides