Where Caching Helps a Zero Trust API and Where It Creates Risk
Caching and zero trust pull in opposite directions: caching wants to avoid repeated checks, and zero trust wants every request checked fresh. Both goals are legitimate, and the resolution isn't choosing one over the other, it's being precise about which layer of your system actually benefits from caching and which one can't afford it.
Here's how the tradeoffs actually play out across the layers where teams typically cache something, and what tends to go wrong at each one when the caching decision is made without thinking through the permission implications first.
Is it safe to cache API response data under zero trust?
Caching the actual response payload for a read-heavy, rarely-changing resource is generally fine and delivers a real performance win, as long as the cache key includes the caller's identity and permission scope, not just the resource being requested. A cache keyed only on the resource ID will happily serve one user's cached response to a different user who shouldn't see it.
This is the most common caching mistake in practice: the cache works correctly for the first caller, passes every test written from that caller's perspective, and silently leaks data to a second caller with different permissions the first time someone tests it that way. A shared cache in front of a multi-tenant API is the highest-risk place for this specific bug, since a key collision there means one tenant's cached response leaking to a different tenant entirely, not just a different user within the same account.
How long can you safely cache a permission decision?
Caching the yes-or-no result of "can this identity do this action" saves a real database round trip on high-volume endpoints, but it means a revoked permission stays effectively active until the cache entry expires. Keep this TTL short, seconds rather than minutes, and treat any permission change as an event that should actively invalidate the relevant cache entries rather than waiting them out.
For high-consequence actions, deleting a resource, changing another user's role, skip the cache entirely and check fresh every time. The latency cost is worth it for anything where a stale "yes" is a real problem, and the volume of high-consequence actions is usually low enough that skipping the cache there doesn't meaningfully affect overall system load.
Negative caching: the case teams forget to handle at all
Most teams cache the "yes, allowed" result and never think about caching the "no, denied" one, which means a caller repeatedly probing an endpoint they don't have access to generates a fresh database check on every single attempt. A short-lived negative cache, denying the same request without a full round trip, both improves performance under that kind of load and makes a probing pattern easier to spot, since the requests all hit the cache instead of your primary database.
Keep the negative cache TTL even shorter than the positive one, since a caller who was just granted access shouldn't be stuck seeing a stale denial. A mismatch here, where a newly granted permission takes longer to take effect than a newly revoked one, is a confusing experience for the exact users you most want to get right.
Decide caching policy per endpoint, not once for the whole API
A single blanket caching policy applied across every endpoint almost always ends up wrong somewhere, too aggressive for a sensitive action, too conservative for a genuinely low-risk, high-volume read. Walk through your endpoints individually and classify each one by consequence and volume, then apply the caching approach that fits that specific classification rather than a single default.
Say you have a hundred endpoints: the handful handling account deletion or permission changes deserve the fresh-check-always treatment, a larger set of high-volume reads deserve short-TTL response caching, and everything in between gets whatever combination fits its actual risk. This takes longer to set up than one global rule, but it's the difference between a caching strategy that's actually been thought through and one that's just a default nobody revisited.
Apply caching layer by layer using these rules:
- Include the caller's identity and permission scope in every response cache key, so one user's cached response never reaches another.
- Keep permission decision caches on a short lifetime and invalidate entries when a permission changes.
- Add a brief negative cache for denied requests so repeated probing doesn't trigger a full check every time.
- Classify each endpoint by consequence and volume, then choose its caching approach instead of applying one policy everywhere.
What Good Looks Like
Good practice means response caches are keyed on identity and permission scope, permission decision caches use a short TTL with active invalidation on change, and high-consequence actions skip caching entirely in favor of a fresh check.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Can we cache an authorization decision at the edge or CDN layer?
Only for low-consequence, high-volume checks with a very short TTL, and never for anything involving per-request context like the specific resource being accessed. Edge caching an authorization decision for a sensitive action is a significant risk for a modest latency gain.
Does caching conflict with zero trust as a principle?
Not inherently. Zero trust means never assuming trust without verification, not that every verification has to be a fresh database call. A short-lived, correctly-scoped cache is still a verified decision, just one reused for a bounded window.
How do we invalidate a cache immediately when someone's access is revoked?
Publish an event on permission change that actively clears the relevant cache entries, rather than relying on TTL expiry alone. This closes the gap between when access is revoked and when the cache actually reflects it.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
The Real Latency Cost of Zero Trust, and How to Measure It
How to find out how much latency your zero trust controls actually add, which checks are worth the cost, and which ones you can move off the hot path.
How to Audit Whether Your APIs Actually Enforce Zero Trust
A step-by-step method for testing whether your APIs enforce zero trust in practice, not just on paper, and what to do with what you find.
Distributed Locking With Redis: Where Redlock Actually Falls Short
A practical guide to distributed locks with Redis, including where the Redlock algorithm's guarantees break down and when to use a database lock instead.
Keeping Auth Checks Fast as Your API Traffic Grows
A worked example for keeping zero trust authorization checks fast as request volume grows, and where teams usually add latency without noticing.