Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

How to Know If Your Agent Is Actually Working

Traditional software tests are deterministic: the same input produces the same output, or the test fails. An agent built on a language model doesn't work that way, and teams that try to test it with pass or fail assertions end up with a suite that's either too brittle to survive a model update or too loose to catch a real regression.

How do you start an agent evaluation set?

Pull twenty to fifty real examples from actual usage, or from your own best guess at common cases if you haven't launched yet, and write down what a correct outcome looks like for each one. Resist the urge to write hundreds of cases before you've run any of them; a small set you actually review by hand teaches you more than a huge set nobody has time to look at.

How do you score agent outcomes instead of exact wording?

Grade whether the agent called the right tool, whether its final action was correct, and whether it asked for clarification when it should have, rather than comparing its response text word for word against a reference answer. Two different but equally correct phrasings shouldn't fail a test; a wrong tool call should, even if the response text sounds confident.

For example, a refund agent's test case might specify the correct tool as the order lookup, the correct final action as declining the refund because the order is outside the return window, and clarification as not needed. Each part can be checked by code, and none depends on how the reply is phrased. When a case fails, the record of which part failed points straight at the prompt, the tool description or the model, which makes each failure faster to diagnose than a single overall score would.

Run the evaluation set every time something changes

Re-run your evaluation set whenever you change a prompt, a tool definition, or the underlying model, the same discipline you'd apply to a regression test suite for ordinary code. This is the single most effective way to catch a change that quietly makes the agent worse at something it used to handle well, a class of bug that's otherwise invisible until a customer hits it.

Add a model-graded layer once the task gets subjective

For tasks where correctness isn't a simple yes or no, tone, helpfulness, whether a summary captured the important points, use a second model call to grade the output against a rubric you write. This scales further than manual review, but treat its scores as a signal to investigate, not as ground truth on their own; spot-check the grader's judgments against your own periodically.

Use a different, ideally stronger, model for grading than the one being evaluated where you can, since a model judging its own outputs against a rubric it would also use to generate them tends to be more lenient than a genuinely independent check would be.

Build an evaluation framework in this order:

  1. Collect twenty to fifty real examples, or your best guess at common cases, and write down what a correct outcome looks like for each.
  2. Grade whether the agent called the right tool, took the right final action and asked for clarification when it should have, not whether the wording matches.
  3. Re-run the set whenever a prompt, a tool definition or the underlying model changes.
  4. Add a model-graded layer for subjective tasks, using a written rubric, and spot-check the grader against your own judgment.
  5. Turn each real regression into a permanent case so the set grows one failure at a time.

A worked example: catching a regression before a customer does

Say a team updates their agent's system prompt to make its responses more concise, based on feedback that answers were running too long. The change reads well in a quick manual check, three or four example conversations all look tighter and clearer. Running the full evaluation set afterward tells a different story: on eleven of the forty cases, the more concise version now skips a required disclosure the longer version used to include, because the instruction to be brief and the requirement to state the disclosure were competing for space in the model's response.

Without the evaluation set, this regression would most likely have surfaced through a customer complaint or, worse, a compliance issue, weeks after the prompt change shipped. With it, the team caught the tradeoff the same day and rewrote the prompt to make the disclosure non-negotiable rather than something concision could quietly crowd out.

That eleven-out-of-forty case also became a new, permanent entry in the evaluation set itself, so the same tradeoff between brevity and the required disclosure gets checked automatically on every future prompt change, rather than depending on someone remembering this specific incident the next time concision comes up as a goal. This is how an evaluation set is supposed to grow over time, one real regression at a time, rather than being designed exhaustively up front by guessing at every case that might eventually matter, and it's a large part of why the set gets more valuable the longer a team maintains it.

Executive Capability Standard

What Good Looks Like

A working evaluation framework grades an agent's tool choices and actions against real examples, not exact response text, and runs automatically on every prompt, tool, or model change rather than only when someone remembers to check.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull twenty real or realistic examples of tasks your agent handles and write down what a correct outcome looks like for each.
2. Do Manually:Run the agent against that set by hand after your next prompt change and review every result yourself.
3. Delegate:Assign an engineer to own the evaluation set, adding new cases whenever a real failure surfaces in production.
4. Automate:Wire the evaluation suite into CI so it runs automatically on every prompt, tool, or model change before it ships.
5. Buy:Bring in outside help to design a model-graded evaluation layer if your tasks are too subjective for simple pass or fail scoring.

How to Get Started

Frequently Asked Questions

How big does an evaluation set need to be to be useful?

Smaller than most teams assume. Twenty to fifty well-chosen, real examples that you actually review by hand catch more regressions early on than a much larger set nobody has time to look at closely. Grow the set as you find gaps, not before.

Should we grade agent responses with exact string matching?

No. Grade the outcome, the tool called, the action taken, whether clarification was requested, rather than the exact wording of the response. Exact matching fails correct answers that are phrased differently and misses wrong answers that happen to sound confident.

How often should the evaluation suite run?

On every change to a prompt, a tool definition, or the underlying model, the same way a test suite runs on every code change. Running it only occasionally means a regression can sit in production for weeks before anyone notices the pattern.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides