Testing an AI Feature When 'Correct' Isn't a Fixed Answer
A normal service either returns the right response or it doesn't, and a unit test can tell the difference. A feature backed by a model doesn't work that way: two reasonable outputs can both be correct, and the one that passed yesterday might read worse today after a prompt change nobody flagged as risky.
This is how to build an evaluation setup that catches regressions in that kind of system, instead of relying on someone noticing an answer looks off.
Separate what must be exactly right from what's judged on quality
Some parts of an AI-backed response are deterministic and testable the normal way: did it call the right function, did it return valid JSON, did it stay within a token or cost budget. Test those with ordinary assertions, they don't need anything fancier.
The part that's actually hard, is the answer good, helpful, on-topic, needs a different kind of check. Mixing these into one vague 'does it work' test hides which failures are structural bugs and which are quality regressions, and those need different people to fix them.
Build a golden set before you build the eval pipeline
An evaluation framework is only as good as the set of inputs it runs against. Pull real examples from production, including the ones that broke, and keep expanding this set every time something goes wrong in a new way. A golden set of thirty to fifty cases, covering your actual edge cases, catches more real regressions than a generic benchmark ever will.
Run every model or prompt change against this set before it ships, not as an afterthought after users report something odd. The set should grow every time a new failure mode shows up, so it never stops mattering.
Use a second model as a judge, but verify the judge periodically
For subjective quality, a common working pattern is having a separate model score the output against a rubric, helpfulness, correctness, tone. This scales far better than a human reviewing every case, but the judge itself can drift or develop blind spots.
Spot-check the judge's scores against human judgment periodically, especially after changing which model does the judging. Treat the automated score as a strong signal to investigate, not an unquestioned final answer, particularly on cases near the pass or fail boundary.
A worked example: a prompt change that looked like an improvement
Say a team tightens a prompt to make responses shorter, and the golden set's average helpfulness score barely moves. Shipped without further review, support tickets about incomplete answers rise the following week.
Digging into the golden set results after the fact shows the average masked a real problem: most cases got slightly better, but a specific category, multi-step questions, got noticeably worse and dragged. An average score across all cases missed the pattern, which is why breaking eval results down by category matters more than one overall number.
Where evaluation setups fall short in practice
- One aggregate score that hides which category of input is actually failing
- A golden set that never grows, so it stops catching new failure modes
- No check on cost or latency alongside quality, so a change ships that's technically 'better' but too slow or expensive to use
- Evaluation run only before a big release instead of on every change
Make evaluation part of every change, not a pre-launch ritual
The teams that keep AI-backed features reliable run evaluation on every prompt or model change, the same way a normal CI pipeline runs tests on every pull request. Teams with this discipline tend to look a lot like teams with high deployment frequency generally, shipping small changes often because each one is cheap to verify1.
The alternative, evaluating only before major releases, means regressions accumulate silently between checks and get discovered by users instead of by your own pipeline.
Version your golden set the same way you version code
A golden set that lives as a loose folder of examples, edited in place whenever someone feels like it, loses its value as a regression baseline. If a case gets quietly removed because it's inconvenient, or its expected answer gets edited to match whatever the model currently outputs, the set stops measuring anything real.
Treat it like test fixtures: reviewed changes, a visible diff when a case is added or altered, and a reason recorded for why. This is what lets you trust a score improvement is real progress and not the set having been adjusted to make the number look better.
What Good Looks Like
A reliable AI-backed feature has deterministic checks for what must be exactly right, a growing golden set for subjective quality, and evaluation that runs on every change rather than only before major releases.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How big does a golden set need to be to catch real regressions?
Thirty to fifty well-chosen cases, covering your actual edge cases and past failures, catches more than a much larger set of generic examples. Size matters less than whether the cases represent what actually goes wrong in production.
Can we skip human review entirely if we have a good automated judge?
Not entirely. Spot-checking the judge against human judgment, especially near pass or fail boundaries and after changing the underlying model, is what keeps the automated score trustworthy instead of quietly drifting from what a person would actually think.
Should evaluation run in CI or as a separate scheduled job?
In CI, on every change that touches a prompt, model choice, or retrieval logic, the same way you'd run tests on any other code change. A scheduled job catches drift over time but misses the chance to block a bad change before it ships.
What's a common mistake teams make when they first set up an eval pipeline?
Building one aggregate quality score and stopping there. A single number hides which category of input is actually regressing, and the fix is almost always to break results down by case type from the start rather than adding that breakdown later, after a real regression slips through unnoticed.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Catching a Breaking API Change Before It Ships, Not After
How consumer-driven contract testing catches breaking changes between services before deploy, and how to set it up without slowing every release down.
Stress Testing Without Taking Down the System You're Trying to Protect
How to run stress tests aggressive enough to find real breaking points without risking the production system or the customers depending on it.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
Ephemeral Test Environments: When Per-Branch Stacks Pay Off
How to size, seed, and, most importantly, tear down per-branch test environments so they save engineering time instead of quietly burning cloud budget.
What an AI Code Reviewer Catches in a Distributed System, and What It Misses
Which distributed-systems failure modes AI code review catches well, which still need a senior engineer, and how to configure and roll out the tool.