Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Finding the Real Source of Latency in a Distributed System

A slow request in a distributed system almost never has one cause. The request crosses a network hop, waits on a database, maybe sits in a queue, and gets serialized twice before the user sees anything, and "latency is bad" tells you nothing about which of those to fix first.

This guide is a decision process for narrowing that down: where to look first, what each pattern of slowness usually means, and when the fix is architectural instead of a tuning knob.

Why start with the shape of the slowness, not the average?

A rising average latency and a rising p99 with a flat average are different problems. If the average moves, something is systemically slower, often a database query that changed plans or a dependency that got heavier. If only the tail moves, you're usually looking at contention: a lock, a connection pool, or a noisy neighbor on shared infrastructure.

Pull both numbers before you touch anything. Tuning for the average when the real problem is tail latency wastes a sprint on the wrong layer, and the two usually need different fixes: one is about raw capacity, the other is about contention that only shows up under specific conditions you have to go find.

Separate network time from processing time

Distributed tracing exists for exactly this question. Break a slow request into spans and look at how much of the total time is actual work versus waiting on a network hop or a downstream service that hasn't responded yet.

Say a request spends most of its time waiting on one downstream call. That has a very different fix than one where every span is individually fast but there are forty of them chained in sequence. The first needs that one service fixed; the second needs fewer round trips, often by batching or parallelizing calls that don't actually depend on each other.

Why check the database before touching the code?

A surprising share of latency complaints turn out to be a missing index, a query that used to hit cache and stopped, or a connection pool sized for last year's traffic. These are cheap to check and expensive to miss, because engineers often jump straight to rewriting application code when the actual bottleneck is one slow query running on every request.

Look at query plans for anything in the hot path before assuming the problem is architectural. A query that used an index last month and stopped usually means the table's data distribution shifted enough that the planner picked a different strategy, which is a one-line fix once you spot it and a multi-day investigation if you don't.

Decide whether this is a tuning problem or a design problem

Some slowness responds to tuning: bigger connection pools, warmer caches, better query plans. Other slowness is a sign the design itself doesn't fit the traffic pattern, and no amount of tuning fixes a synchronous call chain that should have been asynchronous from the start.

A useful test: if fixing the obvious bottleneck only moves the slowness to the next hop in the chain, you're dealing with a design problem, not a tuning problem, and the fix is usually to remove a hop rather than speed one up.

Common mistakes that make latency work harder than it needs to be

A few habits slow down the diagnosis itself:

  • Optimizing the service that's easiest to change instead of the one the trace actually implicates
  • Adding a cache in front of a slow query instead of fixing the query, which just moves the problem
  • Comparing latency across environments with different data volumes, which makes the comparison meaningless
  • Chasing a single slow trace instead of looking at the distribution across a full traffic window

Fixing the process, not just the query, is usually what keeps latency from creeping back.

A worked example: chasing a slow checkout endpoint

Say checkout requests are averaging 400 milliseconds but the p99 has crept up to 3 seconds over the past month. The trace shows most of the time sitting in a call to an inventory service, and that service's own trace shows its database query time climbing in lockstep.

The first instinct is often to add a cache in front of checkout. The trace already tells you that's the wrong layer: the problem lives in the inventory service's query, likely a missing index after a recent schema change. Fixing the query there resolves it for every caller of that service, not just checkout, and avoids adding a cache that would have masked the real regression until it got worse.

Executive Capability Standard

What Good Looks Like

Good latency management means every service in a critical path has a defined budget, tracing that can show where time actually goes, and a documented answer for whether a given slowdown is tuning or design.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read the full trace of your three slowest endpoints end to end and note which spans are work versus waiting.
2. Do Manually:Add timing logs around each hop in one critical call chain and watch how the split between network and processing changes under load.
3. Delegate:Assign one engineer to own latency budgets per service and to flag any deploy that pushes a service over its allowance.
4. Automate:Instrument distributed tracing across every service so slow-request investigation doesn't start from scratch each time.
5. Buy:Bring in a fractional CTO or performance specialist to review your architecture once tuning stops moving the numbers.

How to Get Started

Frequently Asked Questions

Should we set a latency budget for every service?

Yes, once you have more than a handful of services in a call chain. A per-service budget makes it obvious which service is over its allowance instead of everyone assuming the problem is somewhere else. Set budgets based on what the end-to-end experience actually needs, not round numbers.

Is caching always the right first fix for a slow endpoint?

No. Caching hides a slow underlying query or call, it doesn't fix it, and it adds a new failure mode: stale data. Fix the actual bottleneck first, then add caching where the data genuinely doesn't need to be fresh on every request.

How do we tell if latency is a capacity problem instead of a code problem?

Check whether latency correlates with load. If p99 stays flat regardless of traffic, it's probably code or a specific query. If it climbs as concurrent requests climb, you're running out of some resource, threads, connections, or CPU, before you run out of correctness.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides