Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

Finding Your Real Latency Bottleneck Before Customers Do

A latency benchmark run against a quiet staging environment tells you almost nothing about how your API behaves at nine in the morning when three customers kick off large reports at once. Most latency problems that reach a support queue were visible in the data weeks earlier, just not measured the way that would have caught them.

This is a walkthrough for defining what slow actually means for your product, setting a budget you can hold people to, and finding where time is really going before you start rewriting code.

What should you measure when benchmarking latency?

Averages hide the problem. A p50 of 120 milliseconds can sit next to a p99 of four seconds on the same endpoint, and the average will look fine while a slice of your users are furious. Pick p95 and p99 per endpoint, not a single blended number across your whole API, since a slow admin report endpoint will otherwise drown out a fast, high-traffic one in the average.

Measure under realistic concurrency, not a single request at a time. A benchmark run with one client hitting an idle server will miss lock contention and connection pool exhaustion entirely, and those are usually where the real regressions live.

Set a latency budget the way you would set an uptime budget

The same logic teams already apply to uptime works for latency. A 99.9% availability target allows a little under 8.76 hours of downtime a year1, and a latency budget works the same way: you decide upfront how much slowness is acceptable before it counts as a regression, instead of arguing about it after a customer complains.

Write the budget per endpoint, in milliseconds, and treat crossing it the same way you would treat an uptime breach: an alert, an owner, and a fix window, not a Slack message that scrolls away.

Where is the time in a slow request probably going?

Before blaming the database, check the boring stuff first:

  • N+1 queries hiding behind an ORM that looks fine in code review
  • Cold starts on serverless functions that only show up on the first request after a quiet period
  • Third-party API calls made synchronously inside the request path instead of queued
  • Payload serialization on large responses that never gets profiled because it works fine locally
  • Connection pool limits that quietly queue requests once traffic passes a threshold nobody tested

Each of these shows up cleanly in a distributed trace and almost never shows up in a load test that only measures total request time.

Make regressions reproducible, not anecdotal

Run load tests against a staging environment seeded with production-sized data and a traffic shape that matches your real usage pattern, not a flat rate of identical requests. A synthetic test that hits the same cached endpoint a thousand times a second tells you about your cache, not your application.

When a regression shows up in production, correlate the latency graph against your deploy timeline first. A comparison of observability platforms is worth reading if your current tooling cannot answer that question in under five minutes, because that five minutes is usually the difference between a quick rollback and a multi-hour investigation.

The fix that usually pays off first

Before reaching for a bigger database instance, look at what you are sending over the wire and how often you are asking for it twice. Caching hot reads that do not change every request, trimming response payloads down to what the client actually renders, and reusing connections instead of opening new ones per request tend to close most of the gap between a slow endpoint and an acceptable one.

Only after those are done does it make sense to look at horizontal scaling or a faster data store, since scaling a query that should not have run twice just makes the waste more expensive.

For example, suppose a report endpoint shows a p99 of several seconds while its database queries finish quickly. A trace might show the time going into serializing a very large response that the client mostly ignores. Trimming the payload to the fields the page actually renders fixes that endpoint without touching the database at all. The common mistake is to skip this check and upgrade the database instance first, which costs more and leaves the slow endpoint slow. Before any infrastructure change, compare the time spent in the query, the serialization and the network transfer for one slow request. The largest piece is your first target.

Executive Capability Standard

What Good Looks Like

Good here means you can name your current p95 and p99 for your three busiest endpoints without opening a dashboard, because someone checks and reports them on a fixed schedule.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your own latency dashboards for a week and note which endpoints spike under real traffic, not synthetic pings against an idle system.
2. Do Manually:Run a manual load test against a staging environment seeded with production-sized data, and record p50, p95, and p99 for the endpoints that matter most.
3. Delegate:Assign an engineer to own a latency budget per critical endpoint and report against it on a fixed monthly cadence.
4. Automate:Add distributed tracing and automated regression checks to your deploy pipeline so a latency regression blocks a release instead of surfacing in a support ticket.
5. Buy:Bring in an observability platform with built-in tracing if your current stack cannot show you where time goes inside a single request.

How to Get Started

Frequently Asked Questions

What latency counts as too slow for a typical API?

There is no universal cutoff that applies across products. A reporting endpoint and a login endpoint have different tolerances. Set your budget against your own historical p95 and your competitors' observed behavior for a similar action, then treat any regression against that baseline as the signal, not an arbitrary millisecond target.

Should we optimize for average latency or the tail?

The tail matters more than it looks like it should. A p99 spike means your heaviest users, often your best customers running the largest workloads, are the ones hitting it. Fixing the tail usually improves the average too, but optimizing only the average can leave the tail untouched.

How often should we re-run benchmarks after a fix?

Re-benchmark after any release that touches the hot path for a given endpoint, and run a full load test on a fixed schedule, quarterly at minimum, so a slow creep does not go unnoticed between major releases. A regression caught the week it shipped is a small fix; caught three months later, it is an investigation.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides