Finding the Real Source of Latency in a Distributed System
A slow request in a distributed system almost never has one cause. The request crosses a network hop, waits on a database, maybe sits in a queue, and gets serialized twice before the user sees anything, and "latency is bad" tells you nothing about which of those to fix first.
This guide is a decision process for narrowing that down: where to look first, what each pattern of slowness usually means, and when the fix is architectural instead of a tuning knob.
Why start with the shape of the slowness, not the average?
A rising average latency and a rising p99 with a flat average are different problems. If the average moves, something is systemically slower, often a database query that changed plans or a dependency that got heavier. If only the tail moves, you're usually looking at contention: a lock, a connection pool, or a noisy neighbor on shared infrastructure.
Pull both numbers before you touch anything. Tuning for the average when the real problem is tail latency wastes a sprint on the wrong layer, and the two usually need different fixes: one is about raw capacity, the other is about contention that only shows up under specific conditions you have to go find.
Separate network time from processing time
Distributed tracing exists for exactly this question. Break a slow request into spans and look at how much of the total time is actual work versus waiting on a network hop or a downstream service that hasn't responded yet.
Say a request spends most of its time waiting on one downstream call. That has a very different fix than one where every span is individually fast but there are forty of them chained in sequence. The first needs that one service fixed; the second needs fewer round trips, often by batching or parallelizing calls that don't actually depend on each other.
Why check the database before touching the code?
A surprising share of latency complaints turn out to be a missing index, a query that used to hit cache and stopped, or a connection pool sized for last year's traffic. These are cheap to check and expensive to miss, because engineers often jump straight to rewriting application code when the actual bottleneck is one slow query running on every request.
Look at query plans for anything in the hot path before assuming the problem is architectural. A query that used an index last month and stopped usually means the table's data distribution shifted enough that the planner picked a different strategy, which is a one-line fix once you spot it and a multi-day investigation if you don't.
Decide whether this is a tuning problem or a design problem
Some slowness responds to tuning: bigger connection pools, warmer caches, better query plans. Other slowness is a sign the design itself doesn't fit the traffic pattern, and no amount of tuning fixes a synchronous call chain that should have been asynchronous from the start.
A useful test: if fixing the obvious bottleneck only moves the slowness to the next hop in the chain, you're dealing with a design problem, not a tuning problem, and the fix is usually to remove a hop rather than speed one up.
Common mistakes that make latency work harder than it needs to be
A few habits slow down the diagnosis itself:
- Optimizing the service that's easiest to change instead of the one the trace actually implicates
- Adding a cache in front of a slow query instead of fixing the query, which just moves the problem
- Comparing latency across environments with different data volumes, which makes the comparison meaningless
- Chasing a single slow trace instead of looking at the distribution across a full traffic window
Fixing the process, not just the query, is usually what keeps latency from creeping back.
A worked example: chasing a slow checkout endpoint
Say checkout requests are averaging 400 milliseconds but the p99 has crept up to 3 seconds over the past month. The trace shows most of the time sitting in a call to an inventory service, and that service's own trace shows its database query time climbing in lockstep.
The first instinct is often to add a cache in front of checkout. The trace already tells you that's the wrong layer: the problem lives in the inventory service's query, likely a missing index after a recent schema change. Fixing the query there resolves it for every caller of that service, not just checkout, and avoids adding a cache that would have masked the real regression until it got worse.
What Good Looks Like
Good latency management means every service in a critical path has a defined budget, tracing that can show where time actually goes, and a documented answer for whether a given slowdown is tuning or design.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should we set a latency budget for every service?
Yes, once you have more than a handful of services in a call chain. A per-service budget makes it obvious which service is over its allowance instead of everyone assuming the problem is somewhere else. Set budgets based on what the end-to-end experience actually needs, not round numbers.
Is caching always the right first fix for a slow endpoint?
No. Caching hides a slow underlying query or call, it doesn't fix it, and it adds a new failure mode: stale data. Fix the actual bottleneck first, then add caching where the data genuinely doesn't need to be fresh on every request.
How do we tell if latency is a capacity problem instead of a code problem?
Check whether latency correlates with load. If p99 stays flat regardless of traffic, it's probably code or a specific query. If it climbs as concurrent requests climb, you're running out of some resource, threads, connections, or CPU, before you run out of correctness.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Benchmarking API Gateway Latency the Right Way
A methodology for benchmarking API gateway latency that reflects real traffic, the mistakes that produce misleading numbers, and what to test beyond raw speed.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
Cache Invalidation Is Still the Hard Part
A practical guide to choosing a caching layer and, more importantly, keeping it from serving stale or wrong data across a distributed system.
Load Testing Numbers That Don't Match What Users Actually Feel
Why a clean throughput benchmark often fails to predict real-world scaling behavior, and how to build one around your real traffic mix and first bottleneck.
Why Your Redis Lock Let Two Jobs Run at Once (and How to Fix It)
A walkthrough of a real double-charge bug caused by a Redis lock's TTL expiring mid-job, and the fencing-token pattern that actually fixes it.