Caching Context So Your Agents Don't Pay for It Twice
Caching in an agentic system isn't one decision, it's several, and each one trades a different kind of speed and cost against a different kind of staleness risk. Treating "add caching" as a single project tends to produce either an under-cached system that's still slow and expensive, or an over-cached one quietly serving stale data.
Prompt and context caching: the easiest win, and its limits
Many model providers support caching the static parts of a prompt, system instructions and tool definitions that don't change between calls, so you're not paying to reprocess them every single turn. This is close to a free win where it's supported, since it doesn't touch data freshness at all; it only helps the parts of context that are genuinely static, not the tool results that change from call to call.
Tool result caching: where the real tradeoff lives
Caching the output of a tool call, a lookup that returns a product's specifications, for example, can cut both latency and cost meaningfully if that data doesn't change often. The tradeoff is straightforward: cache too long and the agent works from stale data; cache too short and you've added complexity for little benefit. Set expiry per tool based on how often the underlying data actually changes, not a single blanket rule across every tool.
A product catalog entry might be safe to cache for an hour. A customer's current account balance is not safe to cache at all.
Also decide what identifies a cached result. A tool result that depends on who is asking, such as an order lookup for a specific customer, must be keyed by that user as well as by the query, or one customer's data can be served to another. Test this deliberately by requesting the same lookup as two different users and confirming each gets their own answer. Shared, user-independent data such as product specifications can safely use a simpler key.
When caching is the wrong tool for the job
For anything financial, anything involving current account state, or anything where a stale answer could lead to a wrong action, skip caching and accept the latency cost of a live lookup. The failure mode of a stale cache in these cases isn't just a slightly outdated answer, it can be an agent confidently taking an action based on data that was already wrong when it read it.
Invalidate deliberately, not just on a timer
Where possible, invalidate a cached tool result immediately when the underlying data changes through a normal write path, rather than relying solely on a time-based expiry to eventually catch up. This matters most for data an agent might read and act on shortly after it was updated elsewhere in your system, a case a pure timer-based cache handles poorly.
A write-triggered invalidation is more engineering work than a simple timer, but it's the difference between a cache that's occasionally slightly behind and one that's occasionally wrong in a way that actually matters to the person relying on the agent's answer.
Match the caching approach to the data before you turn it on:
- Cache the static parts of the prompt, such as system instructions and tool definitions, where your provider supports it.
- Set tool result expiry per tool, based on how often the underlying data changes, not one blanket rule.
- Skip caching for financial data, current account state, or anything where a stale answer could lead to a wrong action.
- Invalidate a cached result when the underlying data changes through a normal write path, not only on a timer.
- Review expiry settings whenever a workflow's usage pattern changes meaningfully, such as before a launch or seasonal spike.
A worked example: a cache that looked fine until it wasn't
Say a team caches a tool's inventory count for fifteen minutes to reduce load on a database that was struggling under agent-driven query volume. For months this works well, inventory rarely moves fast enough for fifteen minutes to matter. Then a flash sale drives a burst of purchases, and for the length of that fifteen-minute window, the agent keeps confidently telling customers an item is in stock well after it actually sold out.
The fix wasn't to remove the cache, the database still needed the relief it provided, it was to add a write-triggered invalidation specifically on the inventory table, so a cached count clears immediately the moment a purchase changes it, while everything else about the caching strategy stayed the same. The lesson generalizes: a caching decision that was correct under normal conditions can become wrong under exactly the conditions, a sudden spike in real activity, where getting it right matters most, which is worth remembering the next time a cache expiry seems safe simply because nothing's gone wrong with it yet. Review cache expiry settings whenever a workflow's usage pattern changes meaningfully, not only when it's first set up. A calendar reminder tied to major product launches or seasonal spikes is a reasonable trigger for that review, even without a formal process around it.
What Good Looks Like
Sound caching for an agent system uses prompt caching for genuinely static content, sets per-tool expiry based on how often the underlying data actually changes, and skips caching entirely for anything financial or action-critical.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is prompt caching worth setting up even for a small agent system?
Usually yes, since it typically requires little engineering effort and carries no data freshness risk, only the static parts of a prompt are cached. It's a reasonable first step before tackling the harder, riskier decisions around caching tool results.
How do we decide how long to cache a specific tool's results?
Base it on how often the underlying data actually changes and how costly a stale answer would be if it were wrong, not on a single default across every tool. A product description can tolerate a longer cache than a current inventory count, and an account balance shouldn't be cached at all.
Should we ever cache data an agent uses to take an action?
Be very cautious here. If a stale cached value could lead the agent to take a wrong action, sending a discount that no longer applies, confirming an appointment slot that's already taken, the latency savings usually aren't worth the risk, and a live lookup is the safer default.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
Why Your Agent Loop Feels Slow, and How to Fix It
A diagnostic guide to finding where latency actually comes from in an agentic system, and which fixes help each cause instead of masking it.
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
Choosing a Distributed Lock: Redis, Redlock, Postgres, or etcd
A comparison of single-node Redis locks, Redlock, Postgres advisory locks, and etcd for coordinating work across multiple application instances.
Finding Your Agent Stack's Breaking Point Before Customers Do
A worked example of benchmarking an agent system's throughput, so you know where it actually breaks under load instead of guessing until it does.
How to Know If Your Agent Is Actually Working
Building an evaluation framework for an AI agent, from the first small test set through catching quality regressions before customers do.