Setting Spend Caps Before Your Agents Set Them for You
An agent stuck in a retry loop, or one that's been manipulated into calling a tool repeatedly, doesn't fail the way a normal application bug does. It keeps running, keeps calling paid APIs, and keeps racking up cost, sometimes for hours before anyone notices, because nothing about the failure looks like a crash. Rate limiting and spend caps exist specifically to bound that scenario.
Should you limit requests, tokens, or dollars?
A request-count limit is the simplest to implement but doesn't account for the fact that one agent turn can vary enormously in cost depending on how many tool calls and how much context it involves. A token-based limit tracks cost more directly. A hard dollar cap per session or per day is the most direct protection against a genuinely runaway loop, and it's worth having even if you also track the other two.
Set limits at more than one level
Per-session limits catch a single stuck conversation. Per-user daily limits catch a pattern across many sessions from the same source. An account-wide or system-wide daily cap catches the scenario where the first two limits are individually reasonable but something is happening at a scale you didn't design for, a lot of users hitting their per-session limit at once, for instance.
Layer the limits so each one catches a different failure:
- A per-session limit catches a single stuck conversation before it runs for the whole length of the exchange.
- A per-user daily limit catches a pattern across many sessions from the same source.
- An account-wide or system-wide daily cap catches activity at a scale you did not design for.
- A separate, lower limit on destructive or expensive tool calls covers the case where a loop keeps calling a write tool.
- A hard dollar cap per session with automatic cutoff bounds the worst case whatever caused the loop.
What should an agent do when it hits a limit?
The agent should degrade gracefully when it hits a limit, telling the user clearly that it's reached a usage boundary, rather than failing silently or returning a confusing error. For destructive or expensive tool calls specifically, consider a lower, separate limit than the one you'd set for read-only lookups, since a loop that keeps calling a write tool is a more expensive and more dangerous failure than one stuck reading the same record repeatedly.
For example, a support agent might allow generous read-only lookups per session but only a few write actions, with a clear message once the write limit is reached: the agent says it has hit a usage boundary and offers to hand the conversation to a person. The message matters as much as the limit, because a silent cutoff looks like a bug and generates support tickets, while an explicit one tells both the user and your own alerting exactly what happened.
Watch for the specific pattern that signals a real problem
A single session hitting a much higher tool-call count than your typical session, or the same tool being called with nearly identical arguments several times in a row, is the signature of a stuck loop or a manipulation attempt, not normal usage. Alert on this pattern specifically, rather than only on total spend crossing a threshold, since spend alerts fire after the damage is already done.
Build the alert to fire on the pattern within minutes, not as part of a daily spend report reviewed the next morning. The gap between when a loop starts and when someone notices it is exactly the window where the damage actually accumulates.
A worked example: what a stuck loop actually looks like in the logs
Say a customer's message contains an unusual formatting quirk that confuses a tool's argument parser, and the tool returns an error the model interprets as "try again with different phrasing" rather than "this input has a structural problem." The model rephrases and calls the same tool again, gets a similar error, rephrases again, and repeats this pattern a few dozen times within one turn before either hitting a cap or, without one, simply continuing until the user gives up waiting.
In the logs, this shows up as the same tool name called with slightly different argument text, one after another, in a single session, well outside the range of tool calls a normal turn would ever need. A per-session call-count cap catches this specific case cleanly, stopping the loop within a handful of attempts instead of letting it run for the full length of the conversation.
The actual fix, once found, was in the tool's error message, changing it from a generic "invalid input" to one that specifically named the formatting quirk, which gave the model something concrete to correct rather than a vague signal to keep guessing at. The rate cap didn't solve the underlying bug, but it kept a confused model from turning a small parsing issue into a large, unnoticed bill while the actual fix was still being worked out.
What Good Looks Like
Solid rate limiting for an agentic system caps spend at the session, user, and system level, degrades gracefully when a limit is hit, and specifically watches for the repeated-call pattern that signals a stuck loop rather than only tracking total spend.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's the most effective single guardrail against a runaway agent bill?
A hard dollar cap per session with automatic cutoff, checked separately from your general monitoring. It's the one guardrail that directly bounds the worst case, regardless of what caused the loop, whether it was a bug, a bad prompt, or a deliberate manipulation attempt.
Should read-only and write tool calls have the same rate limit?
No. A write tool called repeatedly in a loop is both more expensive and more dangerous than a read tool doing the same thing, so it's worth setting a tighter, separate limit on destructive or state-changing actions specifically.
How do we tell a legitimate heavy user from a stuck loop?
Look at the pattern, not just the volume. A legitimate heavy session usually shows varied tool calls working toward a resolution; a stuck loop usually shows the same tool called with nearly identical arguments repeatedly. Alerting on that repetition pattern catches problems earlier than a raw volume threshold does.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Setting Rate Limits Without Breaking Your Best Customers
A decision guide for setting per-tier rate limits and spend caps that protect your infrastructure without throttling the customers you most want to keep.
Setting Rate Limits and Spend Caps Without Breaking Real Users
How to set request-rate limits and spend caps for AI model serving separately, with a worksheet for setting your first cap.
Setting Rate Limits and Spend Caps That Don't Break Real Usage
How to set rate limits and spend caps that stop abuse and runaway costs without throttling your actual customers, with a worked example.
The Rate-Limit Gaps a 30-Minute Audit Usually Finds
A short, practical checklist for finding the rate-limiting and spend-cap gaps that let one bad actor or one buggy client burn through your budget.
Setting Spend Caps on a RAG Pipeline Without Breaking It
Per-tenant quotas, graceful degradation instead of hard rejection, and separating ingestion from query traffic: a practical guide to RAG rate limits.
Setting Rate Limits That Protect Budget, Not Just Uptime
A practical checklist for designing rate limits and spend caps that stop runaway costs and abuse without breaking legitimate customer usage.