Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

Diagnosing Slow Requests Before You Blame the Database

To diagnose slow requests, first confirm where the time goes: the network, your application code, or the database. Most latency fixes fail because teams add a cache, upgrade an instance, or rewrite a query before checking, and the slowdown returns because the real bottleneck was never touched.

Diagnosing latency properly takes less time than a bad fix does. Here's the order to work through it.

Find out whether the slowdown is the network, the app, or the database

Every request has three places to lose time: getting to your server, your application code doing its work, and any database or downstream call it makes along the way. Trace a slow request end to end and log a timestamp at each handoff, rather than guessing which layer is responsible.

Say a checkout request takes four seconds. If three of those seconds are spent waiting on a single database query, no amount of front-end optimization or CDN configuration will touch that. Find the layer first, then work on it.

Read your own percentiles, not just the average

An average response time hides the requests that actually matter. If your median response is fast but your 95th or 99th percentile is far slower, you have a specific subset of requests, often a particular query, a particular customer's data size, or a specific endpoint, causing real pain for a real slice of users while your dashboard looks fine.

Pull your percentile breakdown by endpoint before you decide anything is or isn't a problem. A slow p99 on your busiest endpoint usually matters more than a slow average across everything.

Rule out the noisy-neighbor problem before you touch code

If you're on shared or burstable infrastructure, a slowdown can come from something else entirely sharing the same underlying hardware, not your own code at all. Check your provider's own resource metrics (CPU steal time, throttling, or credit balance, depending on the platform) before you spend a day rewriting a function that was never the problem.

This step gets skipped constantly because it feels like giving up on finding a real answer. In practice it's often the fastest one to check and the one most likely to save you from a wasted week.

Where teams over-fix: caching layers that hide the real bottleneck

Caching is the most common overcorrection. It makes a slow query feel fast without fixing why the query was slow, which is fine until the cache misses, the data changes, or you add a new endpoint that needs the same data in a slightly different shape and the whole problem reappears.

Use caching to buy time or handle genuinely expensive, rarely-changing lookups, not as a substitute for finding out why a query or a call is slow in the first place.

A four-step process for the next time latency spikes

When it happens again, work through this in order:

  • Trace one slow request end to end and time each handoff
  • Pull the p95 and p99 for the affected endpoint, not just the average
  • Check host-level resource metrics for contention outside your own code
  • Only then decide whether the fix is a query, an architecture change, or a cache

Skipping straight to step four is how the same slowdown comes back in a different form a few weeks later.

When the fix is architectural, not a quick patch

Sometimes tracing points at something that isn't a quick fix at all: a service that has to call three other services sequentially to answer one request, a database that's being asked to do both fast lookups and heavy reporting queries at once, or a queue that backs up under normal load rather than only during spikes. Adding an index or a cache doesn't touch a bottleneck built into the shape of the system itself.

When you find one of these, resist patching around it with a bigger instance, since that buys time but not much of it. Put the redesign on the roadmap as a real project with an owner and a timeline, and be honest with the rest of the team about the fact that a quick fix isn't coming. A clearly scoped architectural project, even a slow one, beats a string of temporary patches that each buy a few weeks before the same complaint comes back.

Executive Capability Standard

What Good Looks Like

Good here means you can point to exactly which layer, and often which query or endpoint, is responsible for a slowdown before you change anything, using your own traced data instead of a guess.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Learn to read a request trace and a percentile breakdown well enough to tell which layer of a slow request is actually responsible.
2. Do Manually:Trace a handful of your slowest real requests by hand and time each handoff, rather than relying only on an aggregate dashboard.
3. Delegate:Give one engineer ownership of your latency dashboards and the authority to block a release that regresses a key percentile.
4. Automate:Set alerts on your own p95 and p99 baselines per endpoint, not generic thresholds, so a real regression pages someone quickly.
5. Buy:Bring in outside performance engineering help for a one-time deep dive when a core system is slow and no one on the team has traced it before.

How to Get Started

Frequently Asked Questions

Is a CDN going to fix my latency problem?

Only if the slowdown is in delivering static assets or content to users far from your servers. A CDN does nothing for a slow database query, a slow API call, or application code that's doing too much work per request. Trace the request first to see where the time actually goes.

How much latency is acceptable for a typical web app?

It depends entirely on what the request does and what your users expect, so there's no single acceptable number. What matters more is consistency: a page that's usually fast but occasionally very slow for a subset of users is a worse experience than one that's evenly a bit slower.

Should we upgrade our database instance size to fix slow queries?

Only after confirming the database is actually the bottleneck and that the queries themselves are reasonably efficient. A bigger instance can mask a poorly indexed query for a while, but the query will eventually outgrow the extra headroom, and you'll pay for the upgrade either way.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides