API Security, Identity & Zero-TrustPlaybook3 min readUpdated September 2026

The Real Latency Cost of Zero Trust, and How to Measure It

Every security control you add to an API call, mTLS handshake, token validation, a policy engine lookup, has a latency cost. Most teams never measure it, so when a request feels slow, "it's probably the security stack" becomes an unfalsifiable excuse instead of a number.

This is about isolating that cost so you can decide, per check, whether it belongs on the request path at all.

How do you measure the latency of each security check?

A single "auth middleware took 40ms" span tells you almost nothing, because it usually bundles token parsing, a signature check, a policy decision, and sometimes a network call to a permissions service. Break these into separate spans in your tracing so you can see, for a given request, exactly which step cost what.

Once you do this, the common finding is that the cryptographic work (verifying a JWT signature, doing a TLS handshake) is fast, and the slow part is a network round trip to an external policy engine or identity provider that could be cached or run in-process instead.

Keep the spans even after you fix the obvious problems. Without them, a slow dependency added six months later hides inside the same "auth middleware" number you already stopped questioning, and nobody notices until a customer does.

Separate connection-time cost from per-request cost

mTLS and connection-level authentication happen once per connection, not once per request, if you're reusing connections properly. If you're seeing that cost on every request, the problem usually isn't zero trust, it's that your service is opening a new connection every time instead of pooling them. Fix the connection reuse before you consider weakening the security check.

Say your P99 latency budget for an internal call is 50 milliseconds and a fresh mTLS handshake alone costs 15 to 20 milliseconds; that's a connection-pooling problem wearing a security costume.

Which identity and policy decisions can you safely cache?

Not every check needs to hit a live service. A policy decision that depends on a role that changes rarely can be cached for seconds to a couple of minutes without meaningfully increasing risk, while a decision that depends on something that changes constantly, like a session's revocation status, generally shouldn't be cached at all. The right cache duration is a judgment call about how stale is tolerable for that specific decision, not a single number you apply everywhere.

Document which checks are cached, for how long, and why, so the next engineer who's debugging a stale-permission bug knows where to look first.

Imagine a support engineer whose access was just revoked after an offboarding: if their permission check is cached for two minutes, that's a two-minute window where a revoked account can still act, which is a reasonable tradeoff for most low-sensitivity actions but not for one that deletes data. Match the cache duration to what the action can do, not to whatever default your library ships with.

Know which checks you can move off the synchronous path

Some checks have to block the request: you can't return data before confirming the caller is allowed to see it. Others, like writing an audit log entry or updating a rarely-read usage counter, can happen after the response goes out. Moving the second category off the critical path is often the single biggest latency win available, and it has nothing to do with weakening security.

Be honest about which category each check is in. A team under latency pressure sometimes reclassifies a real authorization check as "logging" to move it off the path; that's a shortcut that shows up later as an incident.

Set a latency budget per security layer before you optimize

Without a budget, every optimization conversation turns into "this feels slow" versus "this feels necessary." Set an explicit millisecond budget for authentication, for authorization, and for any policy lookups, based on your actual P99 targets, and measure against it. If a layer blows its budget, that's a concrete engineering problem to fix, not an argument for cutting the layer entirely.

Review the budget when your traffic pattern changes meaningfully, a new region, a big new customer with different request shapes, rather than only when someone complains.

Say your authentication layer is budgeted for five milliseconds and your policy engine for ten; if a new feature adds a second policy lookup per request, that's a concrete, visible breach of the ten-millisecond line, not a vague sense that things got slower, and it points a specific engineer at a specific fix.

Trim the security cost of a slow request in this order:

  1. Split the auth middleware into separate spans for token parsing, signature checks, policy decisions and any network calls.
  2. Check whether connection setup is repeating on every request, and fix connection pooling before touching the security check itself.
  3. Cache decisions that depend on slow-changing roles for a short time, and never cache fast-changing ones like session revocation.
  4. Move work that doesn't need to block the response, such as audit log writes, off the synchronous path.
  5. Set a millisecond budget for authentication, authorization and policy lookups, then treat any layer that blows its budget as an engineering problem.
Executive Capability Standard

What Good Looks Like

Good practice here means you can name, in milliseconds, exactly what each security check costs on the request path, and you've made a deliberate decision about which checks are synchronous and which are cached or deferred.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Add per-check tracing spans to your existing authentication and authorization middleware so the cost of each step is visible, not bundled into one number.
2. Do Manually:Profile your ten highest-traffic endpoints by hand and note where time goes, then fix the obvious offenders like uncached policy lookups.
3. Delegate:Give one engineer ownership of a per-layer latency budget and have them review it whenever traffic patterns change.
4. Automate:Add latency assertions for your security middleware to your test suite so a regression gets caught in CI instead of by a customer.
5. Buy:Bring in outside help to redesign the auth path if the fix requires architectural change, like moving from a remote policy engine to a local decision cache, and your team hasn't done that migration before.

How to Get Started

Frequently Asked Questions

How much latency should zero trust controls realistically add?

There's no single honest number, because it depends entirely on your implementation: a well-cached, connection-pooled setup can add low single-digit milliseconds, while a naive one calling out to a remote policy engine on every request can add tens of milliseconds. Measure your own stack rather than trusting a benchmark from a different architecture.

Should we skip zero trust checks on internal, low-latency-sensitive paths?

Skipping the check is rarely the right fix; the check is almost always cheaper than it looks once it's properly cached and running over a reused connection. Reach for that before deciding a path is too latency-sensitive to authenticate.

What's the first thing to check if auth suddenly got slower?

Check whether connections are being reused. A surprisingly common cause of a latency regression is a deploy that accidentally disabled connection pooling or certificate caching, turning a once-per-connection cost into a once-per-request one.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides