Cutting the Cost of Running LLM Agents at Scale
The first instinct when an agent's bill grows is to switch to a cheaper model. That helps, but it's usually the smallest lever available. Most of the spend in an agentic system comes from how many tokens flow through the loop, not from the per-token price, and the token volume is almost entirely something your own architecture controls.
Where do the tokens in an agent loop actually go?
Break down spend by stage: the system prompt and tool definitions sent on every call, the conversation history carried forward each turn, and the tool results fed back into context. Say a support agent calls a lookup tool that returns a full customer record when the task only needed the plan tier; that one unused field, multiplied across every call, can be a meaningful share of the monthly bill.
Most teams find this breakdown surprising: it's rarely the model's per-token price that's the problem, it's the volume of context repeated on every single call.
Trim what goes into every call, not just what comes out
Tool definitions and system instructions get sent with every request in the loop, so a bloated tool schema or an overly long instruction set is a recurring cost, not a one-time one. Cut tool descriptions down to what the model needs to decide when to call them, and return only the fields a tool call actually needs from downstream, rather than passing through a full API response.
Summarizing or truncating conversation history after a few turns, instead of carrying the full transcript forward indefinitely, is another lever that compounds the longer a session runs.
How should you route agent steps by task difficulty?
Not every step in an agent loop needs your most capable model. Classifying a request, choosing which tool to call, or extracting a field from a document can often run on a smaller, cheaper model, reserving the expensive one for the step that actually requires deeper reasoning. This routing decision is where most of the real savings live, more than the price difference between any two frontier models.
Set a hard spend cap before you optimize, not after
A misbehaving agent stuck in a retry loop, or a prompt injection that gets it calling tools repeatedly, can turn a normal day's spend into a very abnormal one within hours. Put a per-session and a daily spend ceiling in place with an automatic cutoff, separate from any cost optimization work, so a bug in the loop can't turn into a budget incident.
This is worth doing before you touch anything else on this list, because it bounds your downside while you're still working out where the real savings are, and it costs almost nothing to implement compared with the other changes here.
A worked example: where the savings actually came from
Say a team's agent spend has grown to a level that's starting to worry the founders, and the instinct is to swap to a cheaper model across the board. Before doing that, they break the bill down by stage and find that a single tool, one that looks up a customer's full account history to answer a simple "what's my current plan" question, accounts for a large share of total tokens, because it returns the entire history instead of just the current plan field.
Trimming that one tool's output to the three fields the agent actually uses, and adding a short cache since plan changes are rare within a single session, cuts the bill by more than switching every call to a cheaper model would have, and it does so without touching answer quality anywhere in the flow. The cheaper-model swap still happens afterward, for the steps that genuinely don't need the stronger model, but it becomes the second lever, not the first.
The lesson generalizes past this one tool: any time a bill grows faster than usage does, the first question worth asking is whether something in the loop is passing more context than the task in front of it actually needs, before assuming the fix has to involve a different model.
A practical order of attack for cutting agent spend:
- Break spend down by stage: the system prompt and tool definitions, the conversation history carried forward, and the tool results fed back into context.
- Cut tool descriptions and instructions to what the model needs, since they are sent with every request in the loop.
- Return only the fields a tool call needs from downstream instead of passing a full API response through.
- Route classification, tool choice and field extraction to a smaller model, and reserve the expensive one for steps that need deeper reasoning.
- Put per-session and daily spend ceilings with automatic cutoff in place before optimizing, so a loop bug cannot become a budget incident.
What Good Looks Like
Efficient agent spend means you can break your bill down by stage, tool schemas and outputs are trimmed to what a task needs, and a hard spend cap catches runaway loops before they become a budget incident.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is switching to a cheaper model the fastest way to cut costs?
It's the fastest to implement, but usually not the biggest lever. Trimming tool schemas, tool outputs, and conversation history first often cuts token volume enough that the model choice matters less, and you keep the quality of your strongest model where it counts.
How much should tool results be trimmed before they go back to the model?
Down to only the fields the next reasoning step actually needs. If a lookup tool returns twenty fields and the agent only ever uses three, filter the other seventeen out in the tool itself, not by asking the model to ignore them.
Do smaller models hurt output quality if we route to them?
Not for the steps they're suited to, like classification or field extraction. Quality risk shows up when a smaller model is asked to do multi-step reasoning it wasn't chosen for, so route by what the step requires rather than trying to save money on every step equally.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
Why Your Agent Loop Feels Slow, and How to Fix It
A diagnostic guide to finding where latency actually comes from in an agentic system, and which fixes help each cause instead of masking it.
A FinOps Checklist for Teams Before Their First Big Cloud Bill
The cost-optimization checklist to run before your cloud bill becomes a board topic, plus the five mistakes that quietly undo every fix on the list.
Three Ways to Cut Cloud Spend, and When Each One Works
Rightsizing, committed-use discounts, and architecture changes all cut cloud spend differently. Here's how to pick the right one for your situation.
Build vs. Buy for Your Security Tooling Stack
A decision framework for when to build DevSecOps tooling in-house versus buying a platform, based on team size, maintenance burden and audit needs.