Diagnosing Slow Requests Before You Blame the Database
To diagnose slow requests, first confirm where the time goes: the network, your application code, or the database. Most latency fixes fail because teams add a cache, upgrade an instance, or rewrite a query before checking, and the slowdown returns because the real bottleneck was never touched.
Diagnosing latency properly takes less time than a bad fix does. Here's the order to work through it.
Find out whether the slowdown is the network, the app, or the database
Every request has three places to lose time: getting to your server, your application code doing its work, and any database or downstream call it makes along the way. Trace a slow request end to end and log a timestamp at each handoff, rather than guessing which layer is responsible.
Say a checkout request takes four seconds. If three of those seconds are spent waiting on a single database query, no amount of front-end optimization or CDN configuration will touch that. Find the layer first, then work on it.
Read your own percentiles, not just the average
An average response time hides the requests that actually matter. If your median response is fast but your 95th or 99th percentile is far slower, you have a specific subset of requests, often a particular query, a particular customer's data size, or a specific endpoint, causing real pain for a real slice of users while your dashboard looks fine.
Pull your percentile breakdown by endpoint before you decide anything is or isn't a problem. A slow p99 on your busiest endpoint usually matters more than a slow average across everything.
Rule out the noisy-neighbor problem before you touch code
If you're on shared or burstable infrastructure, a slowdown can come from something else entirely sharing the same underlying hardware, not your own code at all. Check your provider's own resource metrics (CPU steal time, throttling, or credit balance, depending on the platform) before you spend a day rewriting a function that was never the problem.
This step gets skipped constantly because it feels like giving up on finding a real answer. In practice it's often the fastest one to check and the one most likely to save you from a wasted week.
Where teams over-fix: caching layers that hide the real bottleneck
Caching is the most common overcorrection. It makes a slow query feel fast without fixing why the query was slow, which is fine until the cache misses, the data changes, or you add a new endpoint that needs the same data in a slightly different shape and the whole problem reappears.
Use caching to buy time or handle genuinely expensive, rarely-changing lookups, not as a substitute for finding out why a query or a call is slow in the first place.
A four-step process for the next time latency spikes
When it happens again, work through this in order:
- Trace one slow request end to end and time each handoff
- Pull the p95 and p99 for the affected endpoint, not just the average
- Check host-level resource metrics for contention outside your own code
- Only then decide whether the fix is a query, an architecture change, or a cache
Skipping straight to step four is how the same slowdown comes back in a different form a few weeks later.
When the fix is architectural, not a quick patch
Sometimes tracing points at something that isn't a quick fix at all: a service that has to call three other services sequentially to answer one request, a database that's being asked to do both fast lookups and heavy reporting queries at once, or a queue that backs up under normal load rather than only during spikes. Adding an index or a cache doesn't touch a bottleneck built into the shape of the system itself.
When you find one of these, resist patching around it with a bigger instance, since that buys time but not much of it. Put the redesign on the roadmap as a real project with an owner and a timeline, and be honest with the rest of the team about the fact that a quick fix isn't coming. A clearly scoped architectural project, even a slow one, beats a string of temporary patches that each buy a few weeks before the same complaint comes back.
What Good Looks Like
Good here means you can point to exactly which layer, and often which query or endpoint, is responsible for a slowdown before you change anything, using your own traced data instead of a guess.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is a CDN going to fix my latency problem?
Only if the slowdown is in delivering static assets or content to users far from your servers. A CDN does nothing for a slow database query, a slow API call, or application code that's doing too much work per request. Trace the request first to see where the time actually goes.
How much latency is acceptable for a typical web app?
It depends entirely on what the request does and what your users expect, so there's no single acceptable number. What matters more is consistency: a page that's usually fast but occasionally very slow for a subset of users is a worse experience than one that's evenly a bit slower.
Should we upgrade our database instance size to fix slow queries?
Only after confirming the database is actually the bottleneck and that the queries themselves are reasonably efficient. A bigger instance can mask a poorly indexed query for a while, but the query will eventually outgrow the extra headroom, and you'll pay for the upgrade either way.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Benchmark Your Own Gateway Before You Trust Anyone Else's Numbers
Vendor latency numbers are measured on their best day with synthetic traffic. How to build a benchmark against your own traffic shape instead.
How to Ship a Risky Change Without a 2am Rollback
A concrete walkthrough of how to plan a risky production deployment: how to split it, what to watch, and when to decide the rollback trigger.
Build or Buy for Verifying Every Device That Connects?
How to split device identity from device posture checking, what building either one in house actually costs, and where a platform earns its keep instead.
Building a Throughput Benchmark You Can Actually Trust
A worksheet approach to benchmarking throughput: what load pattern to test, what to record, and how synthetic benchmarks lie about real capacity.
Catching a Breaking API Change Before It Ships
How contract testing catches a breaking change between services before it reaches production, and how to set one up without slowing every deploy down.
Three Ways to Cut Cloud Spend, and When Each One Works
Rightsizing, committed-use discounts, and architecture changes all cut cloud spend differently. Here's how to pick the right one for your situation.