Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Testing an AI Feature When 'Correct' Isn't a Fixed Answer

A normal service either returns the right response or it doesn't, and a unit test can tell the difference. A feature backed by a model doesn't work that way: two reasonable outputs can both be correct, and the one that passed yesterday might read worse today after a prompt change nobody flagged as risky.

This is how to build an evaluation setup that catches regressions in that kind of system, instead of relying on someone noticing an answer looks off.

Separate what must be exactly right from what's judged on quality

Some parts of an AI-backed response are deterministic and testable the normal way: did it call the right function, did it return valid JSON, did it stay within a token or cost budget. Test those with ordinary assertions, they don't need anything fancier.

The part that's actually hard, is the answer good, helpful, on-topic, needs a different kind of check. Mixing these into one vague 'does it work' test hides which failures are structural bugs and which are quality regressions, and those need different people to fix them.

Build a golden set before you build the eval pipeline

An evaluation framework is only as good as the set of inputs it runs against. Pull real examples from production, including the ones that broke, and keep expanding this set every time something goes wrong in a new way. A golden set of thirty to fifty cases, covering your actual edge cases, catches more real regressions than a generic benchmark ever will.

Run every model or prompt change against this set before it ships, not as an afterthought after users report something odd. The set should grow every time a new failure mode shows up, so it never stops mattering.

Use a second model as a judge, but verify the judge periodically

For subjective quality, a common working pattern is having a separate model score the output against a rubric, helpfulness, correctness, tone. This scales far better than a human reviewing every case, but the judge itself can drift or develop blind spots.

Spot-check the judge's scores against human judgment periodically, especially after changing which model does the judging. Treat the automated score as a strong signal to investigate, not an unquestioned final answer, particularly on cases near the pass or fail boundary.

A worked example: a prompt change that looked like an improvement

Say a team tightens a prompt to make responses shorter, and the golden set's average helpfulness score barely moves. Shipped without further review, support tickets about incomplete answers rise the following week.

Digging into the golden set results after the fact shows the average masked a real problem: most cases got slightly better, but a specific category, multi-step questions, got noticeably worse and dragged. An average score across all cases missed the pattern, which is why breaking eval results down by category matters more than one overall number.

Where evaluation setups fall short in practice

  • One aggregate score that hides which category of input is actually failing
  • A golden set that never grows, so it stops catching new failure modes
  • No check on cost or latency alongside quality, so a change ships that's technically 'better' but too slow or expensive to use
  • Evaluation run only before a big release instead of on every change

Make evaluation part of every change, not a pre-launch ritual

The teams that keep AI-backed features reliable run evaluation on every prompt or model change, the same way a normal CI pipeline runs tests on every pull request. Teams with this discipline tend to look a lot like teams with high deployment frequency generally, shipping small changes often because each one is cheap to verify1.

The alternative, evaluating only before major releases, means regressions accumulate silently between checks and get discovered by users instead of by your own pipeline.

Version your golden set the same way you version code

A golden set that lives as a loose folder of examples, edited in place whenever someone feels like it, loses its value as a regression baseline. If a case gets quietly removed because it's inconvenient, or its expected answer gets edited to match whatever the model currently outputs, the set stops measuring anything real.

Treat it like test fixtures: reviewed changes, a visible diff when a case is added or altered, and a reason recorded for why. This is what lets you trust a score improvement is real progress and not the set having been adjusted to make the number look better.

Executive Capability Standard

What Good Looks Like

A reliable AI-backed feature has deterministic checks for what must be exactly right, a growing golden set for subjective quality, and evaluation that runs on every change rather than only before major releases.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull twenty real production examples, including ones that went wrong, and write down what a good response actually looks like for each.
2. Do Manually:Run your current prompt or model against that set by hand and score the results yourself before building any tooling around it.
3. Delegate:Give one engineer ownership of the golden set, with the job of adding a new case every time a real failure shows up in production.
4. Automate:Wire evaluation into CI so every prompt or model change runs against the golden set automatically before it ships.
5. Buy:Bring in an ML evaluation specialist once your feature set has grown past what one team can keep golden sets current for by hand.

How to Get Started

Frequently Asked Questions

How big does a golden set need to be to catch real regressions?

Thirty to fifty well-chosen cases, covering your actual edge cases and past failures, catches more than a much larger set of generic examples. Size matters less than whether the cases represent what actually goes wrong in production.

Can we skip human review entirely if we have a good automated judge?

Not entirely. Spot-checking the judge against human judgment, especially near pass or fail boundaries and after changing the underlying model, is what keeps the automated score trustworthy instead of quietly drifting from what a person would actually think.

Should evaluation run in CI or as a separate scheduled job?

In CI, on every change that touches a prompt, model choice, or retrieval logic, the same way you'd run tests on any other code change. A scheduled job catches drift over time but misses the chance to block a bad change before it ships.

What's a common mistake teams make when they first set up an eval pipeline?

Building one aggregate quality score and stopping there. A single number hides which category of input is actually regressing, and the fix is almost always to break results down by case type from the start rather than adding that breakdown later, after a real regression slips through unnoticed.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides