What Should Happen When Your Vector Search Call Fails
When a vector search call fails, the system should fall back deliberately: serve cached or slightly stale results, run keyword search over the same documents, or return an honest degraded-search message, chosen before the outage. A failed call is inevitable, and most teams make this choice by accident during the first incident.
Which fallback should you choose before an outage?
The reasonable options are serving cached or slightly stale results, falling back to plain keyword search over the same documents, or returning an honest message that search is temporarily degraded. Which one is right depends on the feature: a wrong-but-plausible cached answer might be fine for a low-stakes internal tool and genuinely harmful for something like a support answer involving account-specific details. Make this choice deliberately per feature, in writing, rather than letting whatever the code happens to do during the first real outage become the de facto policy.
Three fallback options are worth weighing for each feature:
- Serve cached or slightly stale results, which keeps the feature responsive but can show outdated content.
- Fall back to plain keyword search over the same documents, which lowers relevance but still returns something useful.
- Return an honest message that search is temporarily degraded, which is simplest and avoids a wrong-but-plausible answer.
How do you tell a timeout from a sustained error?
A vector database that's slow under load and one that's genuinely unreachable call for different handling. A single slow response is often worth a bounded retry with a short timeout, since it may just clear on its own. A sustained pattern of failures calls for an immediate fallback rather than retrying repeatedly and adding latency on top of an outage that retries won't fix. Track these as separate signals so your handling logic, and your alerting, can distinguish them.
Don't let the fallback quietly mask a real problem
If a keyword-search fallback covers every vector database outage smoothly enough that users barely notice, that's good for user experience and bad for visibility, since nobody gets paged and the underlying outage can persist far longer than it should. Alert on fallback activation itself as its own signal, separate from alerting on the underlying vector database's health, so a graceful degradation doesn't become an invisible one.
Fail closed on request errors, fail open on transient ones
A malformed request, an invalid filter, a top-k value outside allowed bounds, should return a clear error immediately rather than being silently absorbed by fallback logic meant for infrastructure failures. Reserve the fallback path for genuinely transient problems: a network blip, a timeout, a temporary capacity issue. Conflating the two means a client-side bug in a caller's code gets treated the same as an actual outage, hiding a problem that a clear error message would have surfaced immediately.
Exercise the fallback path in production occasionally
Fallback code that only runs during a real outage tends to rot, since it's the least-exercised path in the system and small changes elsewhere in the codebase can silently break it without anyone noticing until it's actually needed. Deliberately route a small percentage of production traffic through the fallback path on a regular basis, even when the primary path is healthy, to confirm it still works the way you expect. Downtime budgets shrink fast at higher availability targets1, so a fallback that fails the one time it's called for is a meaningfully worse outcome than not having one.
Give the generation step its own fallback too
Retrieval failing isn't the only failure worth planning for: the generation model itself can time out, rate-limit, or return an error independently of whether the vector search succeeded. Decide separately what happens then, showing the retrieved sources directly without a generated summary is often a reasonable fallback, since the underlying information is still available even if the model that would normally synthesize it isn't. Treat retrieval failure and generation failure as two distinct scenarios with two distinct fallback behaviors, not one blended error path.
Write the postmortem question into the runbook itself
After any real incident that triggers the fallback path, the natural next question is whether the fallback behaved the way it was designed to and whether users were actually served something reasonable during the outage. Put that question directly into your incident runbook as an explicit step, rather than leaving it to come up informally in a retrospective days later, so the fallback design gets validated, or corrected, against real incident data instead of only against the assumptions it was built with.
What Good Looks Like
The error recovery standard is a deliberately chosen fallback per feature, decided in advance rather than during an incident, with fallback activation alerted on separately and the fallback path exercised regularly so it works when it's actually needed.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is a keyword search fallback worth building if we don't already have one?
For a customer-facing feature where degraded-but-present search beats an outright error message, yes, it's usually worth the investment. For a lower-stakes internal tool, a clear "search is temporarily unavailable" message might be simpler and honest enough that the added complexity of a keyword fallback isn't worth building and maintaining.
How long should a retry wait before falling back?
Short enough that the added latency stays within your feature's acceptable response time, typically a bounded retry of one or two attempts with a brief timeout rather than an open-ended retry loop. If the retries would push total latency past what's tolerable for the feature, skip straight to the fallback instead.
Should users be told when they're seeing a fallback result instead of a normal one?
For anything where the fallback result quality noticeably differs from normal, yes, a brief note that results may be limited is more honest than silently serving a degraded experience as if it were normal. For a fallback that's genuinely close in quality, like a well-tuned keyword search, this matters less.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Designing Retry and Fallback Logic That Doesn't Make Outages Worse
How to design retries, backoff, and circuit breakers so error recovery logic protects a struggling dependency instead of overwhelming it further.
The Retry Logic That Makes Outages Worse, Not Better
How naive retry and fallback logic can amplify an outage instead of recovering from it, and the specific patterns that actually help.
Designing Retry Logic That Doesn't Make Things Worse
How to decide when a retry actually helps, why fixed intervals cause outages, and how to design fallback logic that degrades instead of breaking.
Where Retry Logic Quietly Drains Your Infrastructure Budget
How poorly designed error handling and retry logic turns a small outage into a large bill, and the specific patterns worth fixing first.
Designing Retry and Fallback Logic That Doesn't Undermine Your Access Controls
A checklist for building retry, idempotency, and fallback logic for zero trust APIs, plus the specific pitfalls that quietly weaken access control.
What Happens When a Tool Call Fails Mid-Task
A decision guide for designing fallback logic in an agent loop, so a single failed tool call degrades gracefully instead of derailing the whole task.