Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

The Signals That Tell You a RAG Pipeline Is Degrading

A RAG pipeline can return a 200 status code, respond in 300 milliseconds, and still be giving users worse answers than it did last month. Standard uptime and latency dashboards won't catch that. Here's what actually will.

How do you track retrieval quality in production?

You usually don't have ground truth for every live query, but you can track proxies: the similarity score of the top retrieved result, the score gap between the first and fifth result, and how often a query returns zero results above your relevance threshold. A gradual drop in average top-result similarity across a week is an early signal that something upstream, a chunking change, a stale index, an embedding model mismatch, is degrading match quality before anyone files a complaint.

It's worth tracking this per collection or per document source separately, not only as one blended average across the whole system, since a single poorly performing source can drag down the aggregate just enough to hide underneath a broader number that still looks acceptable.

Watch for embedding drift between index and query time

If your document embeddings and query embeddings are ever produced by different model versions, even briefly during a rollout, retrieval quality degrades in a way that's easy to miss because nothing errors. Track the embedding model version used for both the index and each incoming query as metadata, and alert if they ever mismatch. This single check catches one of the more common silent-failure modes in production RAG.

This matters most during a migration to a new embedding model, when it's tempting to switch query-time embedding first and backfill the index later to avoid downtime. That approach guarantees a mismatch window; a safer pattern is writing both old and new embeddings during the transition and cutting over query traffic only once the new index is fully backfilled.

How do you measure index freshness lag?

If your ingestion pipeline updates the index on a schedule or in response to source system changes, measure the actual lag between a document changing and that change being queryable. A support team publishing an urgent fix to a help article expects users to get the updated answer soon after, not after an unmonitored batch job that silently stopped running three days ago. Alert on lag exceeding your target, not just on the ingestion job's own success or failure status.

Break this out by source system if you ingest from more than one, since a single broken connector to one source can sit quietly for weeks while the aggregate ingestion pipeline continues reporting healthy overall.

Add a lightweight groundedness signal on generated answers

You don't need a full evaluation framework running on every request to catch the worst cases. A cheap heuristic, flagging responses where the model's answer contains claims that don't appear anywhere in the retrieved chunks, catches a meaningful share of hallucinated or unsupported answers for closer review. Sample a percentage of production traffic through this check rather than running it on everything if cost is a concern.

Route anything the heuristic flags to a lightweight human review queue rather than blocking the response, at least at first. This builds a labeled set of real failure cases over time, which is far more useful for improving the pipeline than a synthetic test set built before you had production traffic to learn from.

Tier your alerts so real degradation doesn't drown in noise

Not every one of these signals deserves a page. A single query with a low similarity score is normal noise; a sustained drop in the average across a day is worth investigating; an embedding-version mismatch is worth paging on immediately, since it silently degrades every query until it's fixed. Sort your signals into tiers like these explicitly rather than routing everything to the same alert channel, or the team will start ignoring all of it once the noisy ones train them to.

Sort your signals into tiers like these:

  • Treat a single low-similarity query as normal noise and record it without paging anyone.
  • Investigate a sustained drop in average top-result similarity over a day, compared with your own historical baseline.
  • Page immediately on an embedding-version mismatch between index and query, since it silently degrades every query until fixed.
  • Alert when index freshness lag grows, for example when a scheduled ingestion job has silently stopped running.

Set SLOs like you would for any other service, with an error budget

Once you have these signals, treat them the way you'd treat any other service-level objective. A team targeting 99.9% availability has roughly 8.76 hours of budget per year1; a similar mindset, an explicit budget for degraded retrieval quality rather than an implicit "we'll notice eventually," gives the team a concrete threshold for when to stop shipping features and go fix the pipeline. For general infrastructure monitoring underneath these RAG-specific signals, see our Datadog vs. New Relic vs. Dynatrace comparison.

Executive Capability Standard

What Good Looks Like

The observability standard is dedicated retrieval-quality, embedding-version, index-freshness, and groundedness signals with their own alerting, tracked separately from generic service uptime and latency.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Review your current dashboards and identify which of the five signals above you have zero visibility into today.
2. Do Manually:Spot-check similarity scores and groundedness on a sample of production queries by hand weekly until you know what normal looks like.
3. Delegate:Have an engineer build the embedding-version-mismatch and index-freshness-lag alerts first, since those two catch the most common silent failures.
4. Automate:Wire all five signals into your existing monitoring stack with alerts tied to a defined threshold, not just a dashboard someone has to remember to check.
5. Buy:Bring in an observability or ML infrastructure specialist once the pipeline is complex enough that building and maintaining these signals is a meaningful ongoing task.

How to Get Started

Frequently Asked Questions

Do we need a full evaluation framework just to monitor production?

Not to start. A handful of cheap proxy signals, similarity scores, embedding version matching, index freshness lag, catches most silent degradation. A full evaluation framework with a golden test set is worth building once you're iterating on the pipeline often enough that these lighter signals stop being specific enough.

How do we know if a drop in similarity scores means anything?

Compare against your own historical baseline rather than an absolute number, since similarity scores aren't comparable across embedding models or even across different query types on the same model. A meaningful signal is a sustained week-over-week drop on the same query mix, not a single day's fluctuation.

Should groundedness checks run on every single request?

Usually not necessary. Sampling a percentage of production traffic, and always sampling anything a user flags as unhelpful, gives you enough signal to catch systemic problems without the added latency and cost of checking every response in real time.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides