API Security, Identity & Zero-TrustPlaybook3 min readUpdated September 2026

Keeping Auth Checks Fast as Your API Traffic Grows

Zero trust means checking identity and permission on every request, and that check has to stay fast even as traffic grows tenfold. The usual failure isn't the authorization logic itself, it's a database round trip on the hot path that was fine at low volume and becomes the whole system's bottleneck at scale.

Walk through where that latency actually comes from, what to fix first, and how the fixes interact with each other so you don't solve one bottleneck only to immediately hit the next one.

Where does your authorization check actually spend its time?

Say a request takes 40 milliseconds end to end, and profiling shows 25 of those milliseconds are a database lookup to check the caller's permissions, that's the bottleneck to fix, not the business logic that runs afterward in 15 milliseconds. Most teams guess at where scaling problems will come from instead of measuring, and guess wrong more often than not.

Profile a representative sample of real production requests, not a synthetic benchmark, since the actual permission-check pattern under real traffic often looks different from what you'd predict from reading the code. A synthetic load test that hammers one endpoint with one identity will never surface the specific query pattern that shows up when a thousand distinct callers, each with different roles, hit different endpoints at once, which is exactly the pattern real growth produces.

Cache permission decisions, not just the underlying data

Caching a user's profile record is common; caching the result of "can this identity perform this action on this resource" is less common and often more valuable, since that's the check running on every single request. Set a short time-to-live, seconds rather than minutes, and invalidate immediately on any permission change rather than waiting for the cache to expire naturally.

The tradeoff is real: a cached permission decision is technically stale for its TTL window, so keep that window short enough that the risk is acceptable for your specific system, and never cache a decision for an action with serious consequences, like a permission revocation that should take effect immediately.

For example, an endpoint checks whether a caller may edit a project on every request, and each check queries a permissions table. Caching the answer to that specific question for a few seconds, keyed by identity, action and resource, removes most of those queries. Then wire the role-change flow so that granting or removing a role deletes the affected cache entries immediately. The short lifetime bounds the risk if an invalidation is missed, and endpoints with severe consequences, such as revoking another user's access, should skip the cache entirely.

Should you use opaque tokens or self-contained tokens?

An opaque token requires a lookup on every request to resolve into an identity and permission set, which is simple and revokes instantly but adds a round trip. A self-contained signed token, carrying claims directly, skips that round trip but makes instant revocation harder, since the token itself is still technically valid until it expires even if you've revoked the underlying grant.

Many teams land on a hybrid: short-lived self-contained tokens, on the order of minutes, with revocation enforced by the short expiry itself rather than an active blocklist check on every request. This keeps the fast path fast while bounding how long a revoked token stays usable. A common mistake is picking self-contained tokens for the latency win and then setting their expiry to a full day for convenience, which quietly reintroduces the exact revocation delay the hybrid approach was meant to avoid.

Match your target latency to your actual availability commitments

A team promising 99.9% availability has roughly nine hours of downtime budget a year to work with, and a slow authorization layer that causes timeouts under load eats directly into that budget1. Set an explicit latency target for your auth check, not just for the request as a whole, and alert on that specifically, since a slow auth layer often gets masked by acceptable overall response times until traffic spikes.

Test this target under load before you need it, not during your first real traffic spike, since that's the worst possible time to discover the authorization layer is where everything falls over. Revisit the target itself periodically too, not just whether you're currently hitting it, since a latency budget that was generous at last year's traffic volume can quietly become tight as request volume grows even when nothing about the code has changed.

Keep auth checks fast as traffic grows by working through this order:

  1. Profile a sample of real production requests to find where the authorization check spends its time, since guesses are often wrong.
  2. Cache permission decisions with a short time-to-live and invalidate them the moment a permission changes.
  3. Choose between opaque and self-contained tokens deliberately, and keep self-contained token expiry short so revocation still bites.
  4. Set an explicit latency target for the auth check itself, not just the whole request, and alert on it.
  5. Load test against that target before a real traffic spike, and revisit it as request volume grows.
Executive Capability Standard

What Good Looks Like

Good practice means you know exactly where your authorization check spends its time under real load, cache permission decisions with a short and deliberate TTL, and have an explicit latency target for the auth layer that you test before a traffic spike forces the issue.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Profile a sample of real production requests to find out how much of total latency comes from the authorization check specifically.
2. Do Manually:Add a short-TTL cache for your highest-volume permission check by hand and measure the latency improvement before rolling it out more broadly.
3. Delegate:Assign an engineer to own auth-path performance specifically, with a defined latency target and alerting distinct from general API performance monitoring.
4. Automate:Build load tests that specifically exercise the authorization layer at your target scale, run on a schedule rather than only before a known traffic event.
5. Buy:Bring in infrastructure advisory help if your authorization layer is architecturally tied to a single database in a way that caching alone can't fix.

How to Get Started

Frequently Asked Questions

How short should a cached permission decision's TTL be?

Seconds, typically five to thirty, for anything where a stale decision has real consequences if permissions change. Longer TTLs are reasonable only for low-stakes, frequently-checked permissions where a brief delay in reflecting a change is genuinely acceptable.

Is a self-contained token less secure than an opaque one?

Not inherently, but it trades instant revocation for lower latency. Keeping the token's own expiry short, on the order of minutes, bounds how long a revoked grant can still be used, which addresses most of the practical risk.

What's the first thing to profile when auth checks get slow under load?

The actual time breakdown of a representative request, not a synthetic benchmark. Most teams assume the bottleneck before measuring it, and the real answer is often a specific database query rather than the authorization logic itself.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides