An On-Call Runbook for When Retrieval Quality Drops
When retrieval quality drops, first triage whether it's a full outage or a quality degradation, then check the three most common causes: an embedding model version mismatch, a silently stopped ingestion pipeline, and a recent deploy that changed chunking, ranking, or prompts. Nothing throws an error and dashboards look healthy, which makes this page harder than downtime.
How do you triage a retrieval quality drop quickly?
Is this a full outage, the vector database or embedding provider unreachable, or a quality degradation, the system is up and responding but answers are worse than usual? These call for different playbooks entirely, and confusing one for the other wastes the first, most valuable minutes of an incident. A quick check of error rates alongside a manual test query usually answers this in under a minute and determines everything that follows.
Resist the urge to skip straight to root-causing before this triage step is done. An engineer who starts debugging embedding drift when the real problem is a fully unreachable vector database has wasted the incident's most valuable early minutes on the wrong branch of the runbook.
Check the three most common causes first
In order of how often they're actually the cause: an embedding model version mismatch between the index and incoming queries, an ingestion pipeline that stopped running silently and left the index stale, and a recent deploy that changed chunking, ranking, or prompt logic without an adequate quality check beforehand. Checking these three before anything more exotic resolves the large majority of real incidents quickly, since exotic causes are, definitionally, rare.
How do you disable just the affected collection?
A scoped kill switch, disabling or falling back for one specific collection or tenant, is far safer under pressure than an all-or-nothing switch that takes down the entire service to contain a problem confined to one part of it. Confirm before an incident, not during one, that this scoped control actually exists and that the on-call engineer knows exactly how to use it without having to read documentation for the first time while users are already affected.
Decide the fallback posture for the duration of the incident
Whatever fallback behavior your system supports, cached results, keyword search, an honest degraded-service message, decide explicitly whether to activate it for the affected area while you investigate the root cause, rather than leaving users exposed to bad answers during the entire triage process. This is a judgment call worth making deliberately and quickly, not something to default into by accident partway through debugging.
For example, an on-call engineer is paged because answer quality dropped after a release. The triage query returns results, error rates are flat, and the deploy that shipped a new chunking rule is an hour old. Rather than debug ranking first, they use the scoped switch to route the affected collection to keyword search, confirm answers improve, and then investigate the chunking change with users no longer exposed to bad answers. Containing first and root-causing second keeps the incident short, and the triggering query joins the golden set afterward.
Write down who to notify and when
Define a clear threshold for when a retrieval quality incident becomes a customer-facing notification rather than an internal ticket, since this decision under pressure, made inconsistently incident to incident, either over-communicates minor blips or under-communicates real customer impact. A downtime budget shrinks fast at higher availability targets1, so know in advance how much of that budget a given incident is consuming and at what point that crosses into needing external communication.
Close the loop with a blameless review and a new test case
Every real incident is a signal that your existing tests and monitoring missed something. Add the query or scenario that triggered it to your golden evaluation set as a permanent regression test, and review, without assigning blame, what monitoring signal should have caught it earlier. An incident that doesn't produce a new test case or a monitoring improvement is a missed opportunity to make the next one less likely.
Keep the runbook somewhere the on-call engineer will actually find it
A detailed runbook buried in an internal wiki nobody browses during an active incident is functionally the same as having no runbook at all. Link it directly from the alert that pages the on-call engineer, and keep it short enough to actually follow under pressure, a checklist of concrete steps rather than a long narrative document explaining the pipeline's full architecture. The best-written runbook is worthless if the person paged for an incident can't find it in the first sixty seconds, so treat its location as part of the design, not an afterthought once the content is done.
A short on-call checklist for a retrieval quality page:
- Run a manual test query and check error rates to decide between a full outage and a quality degradation.
- Check for an embedding model version mismatch, a silently stopped ingestion pipeline, and a recent deploy affecting chunking, ranking, or prompts.
- Use the scoped kill switch to disable or fall back for only the affected collection or tenant.
- Decide deliberately whether to activate a fallback, such as cached results, keyword search, or a degraded-service message.
- Apply your notification threshold to decide whether customers hear about it.
- Add the triggering query to the golden evaluation set and review, without blame, which monitoring signal was missing.
What Good Looks Like
The incident response standard is a written runbook with a fast outage-versus-degradation triage step, the three most common causes checked first, a scoped kill switch, and every real incident feeding a new regression test.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How do we tell a real quality drop from normal variance in retrieval scores?
Compare against your own historical baseline for the same query mix rather than a single day's numbers, since similarity and quality scores naturally fluctuate day to day. A sustained, multi-hour drop that correlates with a recent deploy or ingestion failure is a much stronger signal than a brief dip that could be ordinary noise.
Should the same engineer who deployed a change be the one who responds if it causes an incident?
It helps if they're available, since they likely have the most immediate context, but the runbook shouldn't depend on it. Anyone on call needs to be able to work through the triage steps and use the scoped kill switch without waiting for a specific person, since incidents don't wait for convenient staffing.
What's the fastest way to confirm an embedding model version mismatch during an incident?
Check the version metadata logged alongside the index build against the version your live query path is currently using, assuming you've instrumented this as suggested elsewhere. Without that instrumentation already in place, confirming this during an active incident is much slower, which is itself a strong argument for adding it ahead of time.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Writing an Incident Response Runbook People Actually Follow at 3 A.M.
A worksheet approach to writing incident runbooks that hold up under real pressure, when the person on call is tired, stressed, and reading fast.
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Writing an Incident Runbook People Actually Follow at 2 A.M.
How to write an incident response runbook that a half-awake, stressed engineer can actually follow, instead of one that only reads well in review.
What Actually Belongs in an Incident Response Runbook
What a useful incident response runbook actually contains: the first five minutes, a named commander, a communication cadence, and a scheduled postmortem.
Writing an Incident Runbook People Will Actually Follow
How to write an incident response runbook engineers actually reach for during a real outage, instead of one that sits unread until the next audit.
Zero Trust for a RAG Pipeline Means No Service Gets a Free Pass
A decision framework for applying zero trust to a production RAG pipeline: verifying every service and user call, not just the ones at the edge.