Why Your Agent Loop Feels Slow, and How to Fix It
Agent latency usually comes from one slow hop in the loop, not from the model alone, so trace each step before you tune anything. A single response can involve a model call, three tool calls to different MCP servers, a second reasoning call and a final answer call, each with its own latency profile.
Trace the request before you tune anything
Instrument every hop in the agent loop separately: time to first token from the model, time spent in each tool call, and time spent re-serializing tool results back into the next model call. Without this breakdown, teams routinely spend weeks optimizing model inference when the real bottleneck was a tool calling a downstream API with no timeout and no cache.
A trace that shows three seconds in model reasoning and four seconds waiting on a single MCP tool call tells you exactly where to spend your next engineering week.
The model is rarely the biggest lever
Once you trace it, it's common to find that most of the wall-clock time sits in tool calls, not in the model itself: a synchronous call to a third-party API with no timeout, a database query the tool wasn't built to run efficiently, or several tool calls that could run in parallel but are wired up to run one after another.
Fixing those issues, parallelizing independent tool calls and setting hard timeouts with graceful fallbacks, often does more for perceived speed than switching models.
As a decision rule, fix latency in this order: remove or parallelize work first, shrink context second, and change models last. For example, a chat assistant that must show progress within two or three seconds, with a trace showing most of the time in one tool call, should get a timeout and a cache on that call before anyone evaluates a different model. A background report with a ten second budget may need no change at all, and speeding it up would be a poor use of engineering time.
When the model genuinely is the bottleneck
If your trace shows the model reasoning step itself is the slow part, look at how many tool results you're feeding back into context before the next call. A tool that returns a full record when the agent only needed one field forces the model to read, and pay for, far more tokens than the task required. Trim tool outputs to what the next reasoning step actually needs.
Streaming the final response back to the user, rather than waiting for the whole agent loop to finish before showing anything, also changes how slow the same underlying latency feels.
Set a latency budget before you optimize
Decide up front what "fast enough" means for the task: a background report can tolerate ten seconds, a chat response the user is waiting on probably needs to start showing progress within two or three. Different tasks in the same product can have different budgets, and treating them all the same either over-engineers the report or under-delivers the chat.
Write the budget down per task type and treat a trace that blows past it the same way you'd treat a failed test, something to investigate before the next release, not an occasional annoyance to revisit someday. Teams that skip this step tend to optimize whichever latency complaint was loudest last week instead of the hop that actually matters most across all their traffic.
Work through a slow agent response in this order:
- Trace every hop separately: time to first token, time in each tool call, and time spent passing tool results back into the next model call.
- Look at tool calls first, including synchronous third-party calls with no timeout, inefficient queries, and independent calls wired to run one after another.
- Run independent tool calls in parallel, and set hard timeouts with graceful fallbacks.
- If the model step is the slow part, trim tool outputs to the fields the next reasoning step actually needs.
- Stream the final response so users see progress before the whole loop finishes.
A worked example: chasing the wrong bottleneck
Say a team notices their support agent takes six seconds to answer and assumes the model is slow, so they switch to a faster model and see almost no change. A trace would have shown the real breakdown: half a second for the model's first reasoning step, four seconds waiting on a single tool call to an internal ticketing API with no timeout, and another second and a half for a second model call to write the final answer.
The actual fix, adding a two-second timeout with a fallback message and caching the ticketing lookup for thirty seconds, cut the typical response to under two seconds. The model swap, which the team tried first, would have shaved maybe a few hundred milliseconds off a six-second problem. Tracing before tuning would have found the real four-second hop on day one instead of after a model migration that didn't help.
What Good Looks Like
A well-tuned agent system has per-hop latency tracing in place, a documented budget for what "fast enough" means per task, and independent tool calls that run in parallel rather than in series.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should we cache tool call results to speed up the agent?
For data that doesn't change often, yes, with a short expiry and a way to force a fresh call when it matters. Caching a customer's current account balance is risky; caching a product catalog entry for a few minutes usually isn't.
Does a bigger or newer model fix latency problems?
Sometimes, but only if the trace shows the model itself is the bottleneck. Swapping models without tracing first is a common way to spend money and still have a slow agent, because most latency in practice comes from tool calls, not model inference.
How many tool calls should one agent turn make?
As few as the task genuinely requires. Every extra call adds a network round trip and more context for the model to read. If a task needs five separate lookups, check whether any of them can run in parallel instead of one after another.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
How to Actually Compare API Gateways on Latency
Why most API gateway latency comparisons are misleading, and a more honest way to benchmark the tradeoffs that actually matter for your own traffic.
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
Finding Your Agent Stack's Breaking Point Before Customers Do
A worked example of benchmarking an agent system's throughput, so you know where it actually breaks under load instead of guessing until it does.
How to Know If Your Agent Is Actually Working
Building an evaluation framework for an AI agent, from the first small test set through catching quality regressions before customers do.
Auditing Security on Your MCP and Agent Tool Stack
A step-by-step way for a CTO to audit which tools an AI agent can reach, what each one can do, and where the access is broader than it should be.