Building a Latency Budget Before You Chase Microsecond Fixes
Latency tuning goes wrong in a predictable order: someone notices the app feels slow, profiles the first service they can access, shaves 20 milliseconds off a function that wasn't the bottleneck, and the user-facing number doesn't move. The fix isn't a better profiler. It's building a latency budget first, so you know which hop in the request path is actually allowed to be slow.
A latency budget is just an allocation: if your product needs a page to respond in 400 milliseconds, you decide up front how much of that 400 milliseconds each service, database call and network hop is allowed to spend.
How do you map the request path before profiling?
Pick the single slowest user-facing flow you've had a complaint about: a search, a checkout step, a dashboard load. Draw every hop it takes: browser to load balancer, load balancer to API gateway, API gateway to each backend service it calls, each service to its database or cache, and back. Most teams have never drawn this diagram for their own product, which is why the first optimization attempt usually targets the wrong hop.
For each hop, write down what you'd guess its current latency is. You'll be wrong about several of them, and that's the point: the map tells you where to actually measure.
This exercise usually takes less than an hour and surfaces hops nobody remembered existed, like an internal auth check that calls out to a third service, or a feature flag lookup that hits a separate cache on every request. Those forgotten hops are disproportionately likely to be where your budget is actually being spent, precisely because nobody has looked at them since they were added.
Instrument the hops you guessed wrong about
Add tracing spans (OpenTelemetry is the common choice now) around each hop from your map, not around every function in your codebase. You're looking for the one or two hops that eat most of the budget, not a complete performance profile of everything. For example, a team running a typical three-tier web app often finds that a single unindexed database query or an uncached call to a third-party API accounts for more of the total latency than every application-code path combined.
Run the instrumented flow under realistic load, not a single request from your laptop. A query that takes 5 milliseconds when the database is idle can take 200 milliseconds when ten other requests are competing for the same connection pool.
How do you set per-hop latency budgets?
Once you know each hop's real latency, allocate your total budget across them with headroom: for example, keep the sum of your hop budgets to about 70 percent of the total, so a slow day in one hop doesn't blow the whole page. A hop that's already over its budget is your actual bottleneck; everything else is noise until that one is fixed.
This is also where deployment cadence matters: teams that deploy small changes often can isolate which specific change moved a hop's latency, while teams that batch weeks of changes into one release have to guess. Teams in DORA's highest-performing cluster practice on-demand deployment, often shipping several times in a single day, which is partly why they catch latency regressions within hours instead of discovering them a sprint later1.
Fix the bottleneck, then re-measure the whole path
Resist fixing more than one hop at a time. If you add a cache layer, switch a query index and upsize a database instance in the same week, you won't know which change actually mattered when the next regression shows up. Ship one change, re-run the same instrumented flow under the same load, and compare against your budget.
A common mistake here is declaring victory when the average latency improves but the tail doesn't. If your p50 drops from 300ms to 150ms but your p99 is still 2 seconds, a meaningful share of real users are still having the slow experience your budget was supposed to prevent.
Keep the budget alive after the fire drill ends
A latency budget that only gets checked during an incident stops paying for itself. Add a lightweight alert on the one or two hops you identified as tightest, set, for example, at roughly 80 percent of their budget, so you get warned before a hop breaches rather than after users complain. Revisit the whole map when you add a new service to the request path, not on a fixed calendar, since that's when budgets actually drift.
Write the budget down somewhere the whole team can see it, next to the architecture diagram if you have one, rather than in a document only the engineer who ran the original exercise remembers exists. A budget nobody else knows about doesn't change how new code gets written, and new code is where the next regression will come from.
Run the whole process in this order:
- Pick the slowest user-facing flow and draw every hop it takes, from the browser through each service and database and back.
- Add tracing spans around each hop on that map, not around every function in the codebase.
- Allocate the total budget across the hops with headroom, and treat any hop already over its share as the bottleneck.
- Change one hop at a time, then re-run the same instrumented flow under the same load and compare against the budget.
- Alert on the tightest hops before they breach, and redraw the map whenever a new service joins the request path.
What Good Looks Like
Good latency management means a documented budget per hop in your critical request paths, instrumentation that shows which hop is closest to breaching it, and a habit of re-measuring the whole path after any single fix.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How do we pick the total latency budget in the first place?
Start from a number your product already has an opinion about, like a checkout flow you've promised feels instant, and work backward. If you have no existing target, 200 to 400 milliseconds for an interactive API response is a reasonable starting point to refine once you have real measurements.
Should we budget for p50, p95 or p99 latency?
Budget your worst tolerable p95 or p99, not your average. Averages hide the slow requests that actually drive complaints and churn, since half your users by definition experience latency at or above the median.
What's the fastest way to find the worst hop without full tracing infrastructure?
Add timestamped log lines at the entry and exit of each service in the request path and compute the deltas from your existing logs. It's cruder than distributed tracing, but it will surface an obviously dominant hop within a day, which is often enough to start.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
How to Benchmark an API Gateway Without Fooling Yourself
How to run an API gateway latency benchmark that actually reflects your real traffic, instead of a number that looks good and means little.
What to Track About Engineering Productivity Besides DORA
Why DORA's four metrics don't capture the whole picture of engineering health, and what to measure alongside them without turning metrics into a scoreboard.
Finding Your Real Latency Bottleneck Before Customers Do
A practical approach to latency benchmarking: how to define what slow means, set a budget, and find where the time actually goes before users complain.
Blue-Green, Canary or Rolling: Picking a Deployment Strategy
A decision guide for choosing between blue-green, canary and rolling deployments based on your traffic, database and rollback needs, not what's trendy.
Getting a New Engineer to Their First Production Deploy Faster
How to shrink the time between a new engineer's start date and their first production deploy, without cutting corners on access or review.
Budgeting Latency for Security Scanning Without Slowing Releases
How to set latency budgets that account for security scanning and endpoint agents, so compliance checks don't quietly become your slowest code path.