Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

An On-Call Runbook for When Retrieval Quality Drops

When retrieval quality drops, first triage whether it's a full outage or a quality degradation, then check the three most common causes: an embedding model version mismatch, a silently stopped ingestion pipeline, and a recent deploy that changed chunking, ranking, or prompts. Nothing throws an error and dashboards look healthy, which makes this page harder than downtime.

How do you triage a retrieval quality drop quickly?

Is this a full outage, the vector database or embedding provider unreachable, or a quality degradation, the system is up and responding but answers are worse than usual? These call for different playbooks entirely, and confusing one for the other wastes the first, most valuable minutes of an incident. A quick check of error rates alongside a manual test query usually answers this in under a minute and determines everything that follows.

Resist the urge to skip straight to root-causing before this triage step is done. An engineer who starts debugging embedding drift when the real problem is a fully unreachable vector database has wasted the incident's most valuable early minutes on the wrong branch of the runbook.

Check the three most common causes first

In order of how often they're actually the cause: an embedding model version mismatch between the index and incoming queries, an ingestion pipeline that stopped running silently and left the index stale, and a recent deploy that changed chunking, ranking, or prompt logic without an adequate quality check beforehand. Checking these three before anything more exotic resolves the large majority of real incidents quickly, since exotic causes are, definitionally, rare.

How do you disable just the affected collection?

A scoped kill switch, disabling or falling back for one specific collection or tenant, is far safer under pressure than an all-or-nothing switch that takes down the entire service to contain a problem confined to one part of it. Confirm before an incident, not during one, that this scoped control actually exists and that the on-call engineer knows exactly how to use it without having to read documentation for the first time while users are already affected.

Decide the fallback posture for the duration of the incident

Whatever fallback behavior your system supports, cached results, keyword search, an honest degraded-service message, decide explicitly whether to activate it for the affected area while you investigate the root cause, rather than leaving users exposed to bad answers during the entire triage process. This is a judgment call worth making deliberately and quickly, not something to default into by accident partway through debugging.

For example, an on-call engineer is paged because answer quality dropped after a release. The triage query returns results, error rates are flat, and the deploy that shipped a new chunking rule is an hour old. Rather than debug ranking first, they use the scoped switch to route the affected collection to keyword search, confirm answers improve, and then investigate the chunking change with users no longer exposed to bad answers. Containing first and root-causing second keeps the incident short, and the triggering query joins the golden set afterward.

Write down who to notify and when

Define a clear threshold for when a retrieval quality incident becomes a customer-facing notification rather than an internal ticket, since this decision under pressure, made inconsistently incident to incident, either over-communicates minor blips or under-communicates real customer impact. A downtime budget shrinks fast at higher availability targets1, so know in advance how much of that budget a given incident is consuming and at what point that crosses into needing external communication.

Close the loop with a blameless review and a new test case

Every real incident is a signal that your existing tests and monitoring missed something. Add the query or scenario that triggered it to your golden evaluation set as a permanent regression test, and review, without assigning blame, what monitoring signal should have caught it earlier. An incident that doesn't produce a new test case or a monitoring improvement is a missed opportunity to make the next one less likely.

Keep the runbook somewhere the on-call engineer will actually find it

A detailed runbook buried in an internal wiki nobody browses during an active incident is functionally the same as having no runbook at all. Link it directly from the alert that pages the on-call engineer, and keep it short enough to actually follow under pressure, a checklist of concrete steps rather than a long narrative document explaining the pipeline's full architecture. The best-written runbook is worthless if the person paged for an incident can't find it in the first sixty seconds, so treat its location as part of the design, not an afterthought once the content is done.

A short on-call checklist for a retrieval quality page:

  1. Run a manual test query and check error rates to decide between a full outage and a quality degradation.
  2. Check for an embedding model version mismatch, a silently stopped ingestion pipeline, and a recent deploy affecting chunking, ranking, or prompts.
  3. Use the scoped kill switch to disable or fall back for only the affected collection or tenant.
  4. Decide deliberately whether to activate a fallback, such as cached results, keyword search, or a degraded-service message.
  5. Apply your notification threshold to decide whether customers hear about it.
  6. Add the triggering query to the golden evaluation set and review, without blame, which monitoring signal was missing.
Executive Capability Standard

What Good Looks Like

The incident response standard is a written runbook with a fast outage-versus-degradation triage step, the three most common causes checked first, a scoped kill switch, and every real incident feeding a new regression test.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Write down what your team currently does, informally, when a retrieval quality issue gets reported, and compare it against the steps above.
2. Do Manually:Draft the runbook by hand and walk through it once as a tabletop exercise with the on-call rotation before an actual incident happens.
3. Delegate:Assign an owner to keep the runbook current as the pipeline's architecture and common failure modes change over time.
4. Automate:Build the scoped kill switch and the embedding-version mismatch check into tooling the on-call engineer can use directly, not just read about.
5. Buy:Bring in an incident response or reliability specialist once this pipeline is critical enough that response time during a real incident has real business cost.

How to Get Started

Frequently Asked Questions

How do we tell a real quality drop from normal variance in retrieval scores?

Compare against your own historical baseline for the same query mix rather than a single day's numbers, since similarity and quality scores naturally fluctuate day to day. A sustained, multi-hour drop that correlates with a recent deploy or ingestion failure is a much stronger signal than a brief dip that could be ordinary noise.

Should the same engineer who deployed a change be the one who responds if it causes an incident?

It helps if they're available, since they likely have the most immediate context, but the runbook shouldn't depend on it. Anyone on call needs to be able to work through the triage steps and use the scoped kill switch without waiting for a specific person, since incidents don't wait for convenient staffing.

What's the fastest way to confirm an embedding model version mismatch during an incident?

Check the version metadata logged alongside the index build against the version your live query path is currently using, assuming you've instrumented this as suggested elsewhere. Without that instrumentation already in place, confirming this during an active incident is much slower, which is itself a strong argument for adding it ahead of time.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides