Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

Caching Context So Your Agents Don't Pay for It Twice

Caching in an agentic system isn't one decision, it's several, and each one trades a different kind of speed and cost against a different kind of staleness risk. Treating "add caching" as a single project tends to produce either an under-cached system that's still slow and expensive, or an over-cached one quietly serving stale data.

Prompt and context caching: the easiest win, and its limits

Many model providers support caching the static parts of a prompt, system instructions and tool definitions that don't change between calls, so you're not paying to reprocess them every single turn. This is close to a free win where it's supported, since it doesn't touch data freshness at all; it only helps the parts of context that are genuinely static, not the tool results that change from call to call.

Tool result caching: where the real tradeoff lives

Caching the output of a tool call, a lookup that returns a product's specifications, for example, can cut both latency and cost meaningfully if that data doesn't change often. The tradeoff is straightforward: cache too long and the agent works from stale data; cache too short and you've added complexity for little benefit. Set expiry per tool based on how often the underlying data actually changes, not a single blanket rule across every tool.

A product catalog entry might be safe to cache for an hour. A customer's current account balance is not safe to cache at all.

Also decide what identifies a cached result. A tool result that depends on who is asking, such as an order lookup for a specific customer, must be keyed by that user as well as by the query, or one customer's data can be served to another. Test this deliberately by requesting the same lookup as two different users and confirming each gets their own answer. Shared, user-independent data such as product specifications can safely use a simpler key.

When caching is the wrong tool for the job

For anything financial, anything involving current account state, or anything where a stale answer could lead to a wrong action, skip caching and accept the latency cost of a live lookup. The failure mode of a stale cache in these cases isn't just a slightly outdated answer, it can be an agent confidently taking an action based on data that was already wrong when it read it.

Invalidate deliberately, not just on a timer

Where possible, invalidate a cached tool result immediately when the underlying data changes through a normal write path, rather than relying solely on a time-based expiry to eventually catch up. This matters most for data an agent might read and act on shortly after it was updated elsewhere in your system, a case a pure timer-based cache handles poorly.

A write-triggered invalidation is more engineering work than a simple timer, but it's the difference between a cache that's occasionally slightly behind and one that's occasionally wrong in a way that actually matters to the person relying on the agent's answer.

Match the caching approach to the data before you turn it on:

  • Cache the static parts of the prompt, such as system instructions and tool definitions, where your provider supports it.
  • Set tool result expiry per tool, based on how often the underlying data changes, not one blanket rule.
  • Skip caching for financial data, current account state, or anything where a stale answer could lead to a wrong action.
  • Invalidate a cached result when the underlying data changes through a normal write path, not only on a timer.
  • Review expiry settings whenever a workflow's usage pattern changes meaningfully, such as before a launch or seasonal spike.

A worked example: a cache that looked fine until it wasn't

Say a team caches a tool's inventory count for fifteen minutes to reduce load on a database that was struggling under agent-driven query volume. For months this works well, inventory rarely moves fast enough for fifteen minutes to matter. Then a flash sale drives a burst of purchases, and for the length of that fifteen-minute window, the agent keeps confidently telling customers an item is in stock well after it actually sold out.

The fix wasn't to remove the cache, the database still needed the relief it provided, it was to add a write-triggered invalidation specifically on the inventory table, so a cached count clears immediately the moment a purchase changes it, while everything else about the caching strategy stayed the same. The lesson generalizes: a caching decision that was correct under normal conditions can become wrong under exactly the conditions, a sudden spike in real activity, where getting it right matters most, which is worth remembering the next time a cache expiry seems safe simply because nothing's gone wrong with it yet. Review cache expiry settings whenever a workflow's usage pattern changes meaningfully, not only when it's first set up. A calendar reminder tied to major product launches or seasonal spikes is a reasonable trigger for that review, even without a formal process around it.

Executive Capability Standard

What Good Looks Like

Sound caching for an agent system uses prompt caching for genuinely static content, sets per-tool expiry based on how often the underlying data actually changes, and skips caching entirely for anything financial or action-critical.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Check whether your model provider supports prompt or context caching and whether you're currently using it.
2. Do Manually:Manually set a reasonable cache expiry on your single slowest, least-sensitive tool call and measure the difference.
3. Delegate:Assign an engineer to review cache expiry settings per tool and flag anything cached that shouldn't be.
4. Automate:Build cache invalidation into your normal write paths so a cached tool result updates immediately when the underlying data changes.
5. Buy:Bring in infrastructure advisory support if your caching layer needs to span multiple regions or data stores.

How to Get Started

Frequently Asked Questions

Is prompt caching worth setting up even for a small agent system?

Usually yes, since it typically requires little engineering effort and carries no data freshness risk, only the static parts of a prompt are cached. It's a reasonable first step before tackling the harder, riskier decisions around caching tool results.

How do we decide how long to cache a specific tool's results?

Base it on how often the underlying data actually changes and how costly a stale answer would be if it were wrong, not on a single default across every tool. A product description can tolerate a longer cache than a current inventory count, and an account balance shouldn't be cached at all.

Should we ever cache data an agent uses to take an action?

Be very cautious here. If a stale cached value could lead the agent to take a wrong action, sending a discount that no longer applies, confirming an appointment slot that's already taken, the latency savings usually aren't worth the risk, and a live lookup is the safer default.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides