The Signals That Tell You a RAG Pipeline Is Degrading
A RAG pipeline can return a 200 status code, respond in 300 milliseconds, and still be giving users worse answers than it did last month. Standard uptime and latency dashboards won't catch that. Here's what actually will.
How do you track retrieval quality in production?
You usually don't have ground truth for every live query, but you can track proxies: the similarity score of the top retrieved result, the score gap between the first and fifth result, and how often a query returns zero results above your relevance threshold. A gradual drop in average top-result similarity across a week is an early signal that something upstream, a chunking change, a stale index, an embedding model mismatch, is degrading match quality before anyone files a complaint.
It's worth tracking this per collection or per document source separately, not only as one blended average across the whole system, since a single poorly performing source can drag down the aggregate just enough to hide underneath a broader number that still looks acceptable.
Watch for embedding drift between index and query time
If your document embeddings and query embeddings are ever produced by different model versions, even briefly during a rollout, retrieval quality degrades in a way that's easy to miss because nothing errors. Track the embedding model version used for both the index and each incoming query as metadata, and alert if they ever mismatch. This single check catches one of the more common silent-failure modes in production RAG.
This matters most during a migration to a new embedding model, when it's tempting to switch query-time embedding first and backfill the index later to avoid downtime. That approach guarantees a mismatch window; a safer pattern is writing both old and new embeddings during the transition and cutting over query traffic only once the new index is fully backfilled.
How do you measure index freshness lag?
If your ingestion pipeline updates the index on a schedule or in response to source system changes, measure the actual lag between a document changing and that change being queryable. A support team publishing an urgent fix to a help article expects users to get the updated answer soon after, not after an unmonitored batch job that silently stopped running three days ago. Alert on lag exceeding your target, not just on the ingestion job's own success or failure status.
Break this out by source system if you ingest from more than one, since a single broken connector to one source can sit quietly for weeks while the aggregate ingestion pipeline continues reporting healthy overall.
Add a lightweight groundedness signal on generated answers
You don't need a full evaluation framework running on every request to catch the worst cases. A cheap heuristic, flagging responses where the model's answer contains claims that don't appear anywhere in the retrieved chunks, catches a meaningful share of hallucinated or unsupported answers for closer review. Sample a percentage of production traffic through this check rather than running it on everything if cost is a concern.
Route anything the heuristic flags to a lightweight human review queue rather than blocking the response, at least at first. This builds a labeled set of real failure cases over time, which is far more useful for improving the pipeline than a synthetic test set built before you had production traffic to learn from.
Tier your alerts so real degradation doesn't drown in noise
Not every one of these signals deserves a page. A single query with a low similarity score is normal noise; a sustained drop in the average across a day is worth investigating; an embedding-version mismatch is worth paging on immediately, since it silently degrades every query until it's fixed. Sort your signals into tiers like these explicitly rather than routing everything to the same alert channel, or the team will start ignoring all of it once the noisy ones train them to.
Sort your signals into tiers like these:
- Treat a single low-similarity query as normal noise and record it without paging anyone.
- Investigate a sustained drop in average top-result similarity over a day, compared with your own historical baseline.
- Page immediately on an embedding-version mismatch between index and query, since it silently degrades every query until fixed.
- Alert when index freshness lag grows, for example when a scheduled ingestion job has silently stopped running.
Set SLOs like you would for any other service, with an error budget
Once you have these signals, treat them the way you'd treat any other service-level objective. A team targeting 99.9% availability has roughly 8.76 hours of budget per year1; a similar mindset, an explicit budget for degraded retrieval quality rather than an implicit "we'll notice eventually," gives the team a concrete threshold for when to stop shipping features and go fix the pipeline. For general infrastructure monitoring underneath these RAG-specific signals, see our Datadog vs. New Relic vs. Dynatrace comparison.
What Good Looks Like
The observability standard is dedicated retrieval-quality, embedding-version, index-freshness, and groundedness signals with their own alerting, tracked separately from generic service uptime and latency.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need a full evaluation framework just to monitor production?
Not to start. A handful of cheap proxy signals, similarity scores, embedding version matching, index freshness lag, catches most silent degradation. A full evaluation framework with a golden test set is worth building once you're iterating on the pipeline often enough that these lighter signals stop being specific enough.
How do we know if a drop in similarity scores means anything?
Compare against your own historical baseline rather than an absolute number, since similarity scores aren't comparable across embedding models or even across different query types on the same model. A meaningful signal is a sustained week-over-week drop on the same query mix, not a single day's fluctuation.
Should groundedness checks run on every single request?
Usually not necessary. Sampling a percentage of production traffic, and always sampling anything a user flags as unhelpful, gives you enough signal to catch systemic problems without the added latency and cost of checking every response in real time.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Datadog vs New Relic vs Dynatrace: Cloud Observability Platforms Compared
Compare Datadog, New Relic, and Dynatrace for cloud observability: log ingestion costs, distributed tracing, APM overhead, and MTTR compression.
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Zero Trust for a RAG Pipeline Means No Service Gets a Free Pass
A decision framework for applying zero trust to a production RAG pipeline: verifying every service and user call, not just the ones at the edge.
Where RAG Latency Actually Goes, and How to Budget It
Break a RAG request into its four latency stages, find out which one is actually slow, and set a budget for each before you start tuning blindly.
How Vector Search Throughput Degrades as Your Index Grows
Throughput doesn't fall off gradually as a vector index grows. Here's why it degrades in steps, and how sharding, replicas, and quantization each help.
Building a Golden Set to Catch RAG Regressions Before Users Do
A step-by-step approach to building a RAG evaluation set from real queries, scoring retrieval and generation separately, and gating on regressions.