Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

Building a Golden Set to Catch RAG Regressions Before Users Do

A golden set is a fixed collection of real queries with known-good expected results that you run against every change to your RAG pipeline, so regressions are caught before users see them. It's built from actual query logs, scored on retrieval and generation separately, and gated against your own previous baseline.

Why should a golden set start with real queries?

A team writing test queries from scratch tends to write queries that are easier than what real users ask: cleanly phrased, on-topic, unambiguous. Pull your golden set from actual query logs instead, including the awkwardly phrased ones and the ones that returned poor results, since those are exactly the cases most likely to regress silently in the future. Aim for a set that reflects the real distribution of query types your system sees, not just the ones that are easy to write expected answers for.

Include a handful of queries the system should decline to answer, too, ones that fall outside your corpus entirely. A pipeline that confidently generates a plausible-sounding answer for a question it has no real data on is a failure mode a golden set built only from answerable queries will never catch.

Define correct before you automate anything

For retrieval, correct usually means a specific set of chunk IDs should appear in the top results. For generation, correct is softer, usually a list of facts that must appear in the answer, and facts that must not. Write these definitions down manually for each query in the set before building any scoring automation, since agreeing on what correct means is a harder problem than building the tooling to check it, and skipping this step produces a set that scores things nobody actually agreed matter.

Measure retrieval and generation separately

A high generation score can mask a retrieval problem if the underlying model already knows the answer from its training data and doesn't actually need the retrieved context to get it right. Score retrieval on its own with recall or mean reciprocal rank against your known-good chunk IDs, and score generation separately with a groundedness check against what was actually retrieved. Only looking at end-to-end answer quality hides exactly the failure mode a golden set exists to catch, and it's the split that tends to reveal which half of the pipeline actually needs the next round of work.

For example, suppose a change to chunk size lifts end-to-end answer quality slightly, so the team ships it. Scored separately, retrieval recall fell on queries that mention specific product codes, but the model answered them from its training data anyway. The generation score hid the problem until the pipeline met a question the model didn't already know. A split score would have flagged the drop at review time, when reverting was cheap, and the affected queries would then be added to the set as permanent regression cases.

Should you gate on a fixed score or your own baseline?

An absolute quality threshold stops being meaningful as your corpus grows and query patterns shift, since a harder, more diverse corpus naturally scores lower even with a well-tuned pipeline. Compare each change against your previous version's score on the same golden set instead, and treat any meaningful drop as a blocker. This also makes the gate resilient to your corpus changing over time in ways that have nothing to do with a specific code or configuration change.

Grow the set from real failures, not just at the start

Every production complaint that turns out to be a genuine retrieval or generation miss is a candidate for a new golden set entry. Building this habit into your team's process, add the failing query, its correct answer, and a regression test in the same pass as the fix, keeps the set relevant instead of frozen at whatever the pipeline looked like on day one, and it's usually the cheapest source of new test cases you have.

A minimal golden set workflow looks like this:

  1. Pull real queries from your logs, including awkward ones, poorly answered ones, and a few the corpus can't answer at all.
  2. Write down what correct means for each query: expected chunk IDs for retrieval, and facts that must and must not appear for generation.
  3. Score retrieval and generation separately, using recall or mean reciprocal rank for one and a groundedness check for the other.
  4. Compare every change against the previous version's score on the same set, and treat a meaningful drop as a blocker.
  5. Add each confirmed production miss to the set in the same pass as its fix.

Decide who owns disagreements about what counts as correct

Two engineers will occasionally disagree about whether a given answer is acceptable, especially for open-ended queries where more than one phrasing could be considered right. Without a named owner for the golden set, these disagreements either stall the set's growth or get resolved inconsistently case by case. Give one person, or a small rotating group, final say on borderline calls, and record the reasoning briefly next to the entry so the next disagreement over similar territory has precedent to work from, instead of relitigating the same judgment call every time it resurfaces.

Executive Capability Standard

What Good Looks Like

The evaluation standard is a golden set built from real query logs with separately defined retrieval and generation correctness, gated on regression against the pipeline's own previous baseline rather than a fixed score.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull a sample of real queries from your logs and manually review how the pipeline currently handles each one.
2. Do Manually:Build an initial golden set of a few dozen queries by hand, with correctness defined for both retrieval and generation.
3. Delegate:Assign an engineer to own growing the golden set from real production failures as part of the normal bug-fix process.
4. Automate:Wire the golden set into CI so a regression against the previous baseline blocks a merge instead of shipping unnoticed.
5. Buy:Bring in an ML evaluation specialist once the pipeline is complex enough that scoring generation quality reliably becomes its own ongoing project.

How to Get Started

Frequently Asked Questions

How many queries does a useful golden set need?

Fewer than teams expect to start, a well-chosen set of a few dozen queries covering your main query types and known edge cases catches most regressions. A huge set built all at once tends to be lower quality per entry than a smaller set built carefully and grown over time from real production failures.

Should the golden set include queries we know the pipeline currently fails?

Yes, and label them explicitly as known failures rather than excluding them. Tracking known failures alongside passing cases lets you see whether a change accidentally fixes one of them, and prevents a future contributor from assuming the set only contains cases the pipeline currently handles correctly.

Can we reuse the same golden set across different RAG features built on the same pipeline?

Only if the features share the same corpus and the same notion of a correct answer. A support search feature and an internal analytics feature built on the same retrieval infrastructure usually need separate golden sets, since what counts as a correct result differs enough between them that a shared set would test neither one well.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides