Watching What Your Agents Actually Do in Production
Traditional application monitoring answers questions like "is the server up" and "how long did the request take." An agent can be fast, return a 200, and still have done the wrong thing, called the wrong tool, misread a result, given a confidently wrong answer. Observability for agentic systems has to answer a different question: not just whether it ran, but whether it reasoned correctly.
What should we actually log for every agent turn?
Log the full decision trail: which tools the model considered, which one it called, the arguments it passed, the result it got back, and the reasoning that led to the next step. Without this, debugging a bad outcome means asking the model to explain itself after the fact, which produces a plausible-sounding story rather than what actually happened.
Log this even for turns that look successful, not just for ones that error, because the most useful traces for improving the agent are often the ones where it technically succeeded but took a roundabout path to get there. Those are the cases that tell you where a tool description or a prompt is unclear before that ambiguity causes an actual failure.
What's worth alerting on, versus what's just noise?
Alert on things that indicate the agent is confused or unsafe, not just things that indicate it's busy: a rising rate of tool call errors, a rising rate of the agent asking the same clarifying question repeatedly, or a jump in how often it falls back to a generic response. Raw call volume alerts tend to fire constantly and get ignored; behavioral drift alerts tend to catch real problems early.
Set the threshold relative to the agent's own recent baseline rather than a fixed number chosen once, since normal usage patterns shift over weeks and months and a fixed threshold either goes stale or gets tuned out entirely by whoever's on call.
How do we tell a good trace from a bad one at a glance?
A healthy trace shows a short, direct path: the model picks a tool, gets a result, and either answers or takes one more well-justified step. An unhealthy trace shows the model calling the same tool repeatedly with slightly different arguments, which usually means it didn't get the information it needed the first time and doesn't know it. Building a habit of spot-checking a handful of traces weekly, not just when something breaks, surfaces this pattern early.
What does a dashboard for this actually need to show?
Keep the main dashboard to a handful of numbers that answer "is the agent behaving normally right now": tool error rate, fallback rate, average number of tool calls per turn, and the rate of repeated identical tool calls within a single turn. Resist the urge to put every metric you're capable of tracking on one screen; a dashboard nobody can read at a glance during an incident isn't actually observability, it's just a lot of collected data.
Pair the dashboard with a way to drill from any one of those numbers straight into the actual traces behind it, since the number tells you something changed but the trace is what tells you why.
A useful test for each dashboard number is whether someone knows what to do when it moves. If the fallback rate rises, the next step is to open a sample of fallback traces and sort them by cause: a dependency timing out, a tool description the model misreads, or a genuine reasoning failure. If nobody can name the next step for a metric, move it off the main screen into a secondary view. That keeps the dashboard readable during an incident, when the first job is deciding where to look.
A dashboard that shows whether the agent is behaving normally needs only these numbers:
- Tool error rate, compared against the agent's own recent baseline instead of a fixed threshold chosen once.
- Fallback rate, meaning how often the agent drops to a generic response instead of completing the task.
- Average number of tool calls per turn, which rises when the agent takes a roundabout path.
- Rate of repeated identical tool calls within a single turn, a common sign the agent did not get the information it needed.
- A drill-down from each number into the actual traces behind it, since the number shows something changed and the trace shows why.
A worked example: what a rising fallback rate actually meant
Say the fallback rate on a scheduling agent creeps up over a couple of weeks, from its usual low baseline to noticeably higher, with no deploy and no obvious cause. Pulling ten of the recent fallback traces shows a pattern: nearly all of them involve a specific calendar tool timing out, not the model refusing the task. The calendar provider had quietly lowered its own rate limits, and the agent's fallback response, while technically correct, was masking what was really an infrastructure problem, not a reasoning one.
Without dashboard drill-down into real traces, this would likely have been chalked up to "the model got worse at scheduling" and investigated in entirely the wrong place, in the prompt, instead of in the actual dependency that had changed.
What Good Looks Like
Good observability for an agentic system captures the full tool-call decision trail per turn, alerts on behavioral drift rather than raw volume, and makes it easy to spot-check a normal trace against an unhealthy one.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How long should we retain full agent decision traces?
Keep full decision traces long enough to investigate a delayed complaint, often 30 to 90 days depending on your support cycle. Redact sensitive fields or store them separately from the reasoning trail, so retention policy does not force you to delete the debugging context you still need.
Should customers be able to see the agent's reasoning trace?
For high-stakes actions, a simplified version, such as which sources it used, builds trust and gives them a way to flag a wrong step. The full internal trace, including tool arguments and intermediate reasoning, is usually better kept internal for debugging.
What's a sign our alerting is set up wrong?
If every alert gets treated as informational instead of actionable within a week of turning it on, the thresholds are probably too sensitive. Alerts should be rare enough that a page actually means something is behaviorally different from normal.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
Why Your Agent Loop Feels Slow, and How to Fix It
A diagnostic guide to finding where latency actually comes from in an agentic system, and which fixes help each cause instead of masking it.
Finding Your Agent Stack's Breaking Point Before Customers Do
A worked example of benchmarking an agent system's throughput, so you know where it actually breaks under load instead of guessing until it does.
How to Know If Your Agent Is Actually Working
Building an evaluation framework for an AI agent, from the first small test set through catching quality regressions before customers do.
Auditing Security on Your MCP and Agent Tool Stack
A step-by-step way for a CTO to audit which tools an AI agent can reach, what each one can do, and where the access is broader than it should be.