Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

A Go-Live Checklist for Shipping a RAG Pipeline

Shipping a RAG pipeline for the first time is different from shipping a normal service, because the thing that can silently break isn't the code, it's the retrieval quality. A deploy can pass every health check and still hand users worse answers than yesterday.

Here's what to check before you call it live.

How do you warm a vector index before go-live?

If your vector index builds in-memory structures on startup, or if your managed vector database needs to load a new collection into cache, the first requests after a deploy or a reindex can be dramatically slower than steady state. Run a warm-up pass, a batch of representative queries against the new index, before routing real traffic to it. This is easy to skip because it doesn't show up in a staging environment with a small test corpus, only at production scale.

How long the warm-up takes depends on index size and the underlying engine, so measure it once for your actual corpus rather than assuming a fixed number of seconds is enough. A warm-up pass that finishes in five seconds against a small test collection can take several minutes against your full production index.

Why should you pin your embedding model version?

Query and document embeddings only compare meaningfully if they came from the same model version. If your embedding provider ships a silent model update, or if your own code auto-upgrades to "latest," queries embedded with the new model can drift out of alignment with a document index still built on the old one, degrading retrieval without any error being thrown. Pin the exact model version in configuration, and treat a model upgrade as a full reindex event, not a config change you deploy casually.

This applies to any other component in the chunking or preprocessing path that can change behavior silently too, a text extraction library upgrade that parses PDFs slightly differently, for instance. Pin dependency versions for the whole ingestion path, not just the embedding model, and treat any upgrade to that path the same way you'd treat an embedding model change.

Run a canary on retrieval quality, not just uptime

A standard canary deploy checks error rates and latency. For a RAG pipeline, add a retrieval-quality check: run a fixed set of test queries with known-good expected chunks against the new deployment, and compare recall against the previous version before shifting full traffic. Teams in DORA's top performance cluster manage on-demand deploys, shipping multiple times a day1, and that pace only stays safe with an automated quality gate like this one, not manual spot checks.

Test with production-shaped data before you call it validated

A staging environment with a hundred sample documents will pass every check and still fail to predict production behavior, since chunking edge cases, duplicate near-identical content, and adversarial formatting only show up at real scale and real document diversity. Before go-live, run your canary query set and your full ingestion pipeline against a copy of production data, or as close to it as your data handling policies allow, rather than a hand-picked demo corpus that happens to work well.

Decide who signs off, and on what, before you shift full traffic

Once the retrieval-quality canary and the warm-up pass both look clean, someone still has to decide whether to shift the remaining traffic. Write down in advance what a pass looks like, a specific recall threshold against your test set, not a subjective "looks fine" review made under deploy-day pressure, and who has authority to hold the rollout if the numbers come back borderline. Deciding this ahead of time removes the incentive to wave through a marginal result because the deploy window is closing.

Write down the actual rollback plan

"Roll back the deploy" isn't enough if the new version also changed the index schema or chunking strategy, since the old code may not be able to read the new index format. Before go-live, confirm whether your rollback plan is a simple code revert or whether it also requires restoring a previous index snapshot. If it's the latter, test that restore process once before you need it under pressure, not during an incident.

Include a decision rule for how long to wait before rolling back versus fixing forward. A retrieval-quality regression caught within minutes of a deploy is usually safer to roll back than to patch live, while a smaller issue discovered hours later, after real ingestion has already happened on the new version, might be safer to fix forward to avoid losing that data.

Work through this list before you shift full traffic:

  1. Warm the new index with a batch of representative queries before routing real traffic, and measure how long that takes on your full corpus.
  2. Pin the exact embedding model version so queries and documents always come from the same model.
  3. Run a canary that compares retrieval recall on a fixed set of known-good queries against the previous version.
  4. Test with a copy of production-shaped data, not a small set of sample documents.
  5. Write down the pass threshold and who can hold the rollout before you shift full traffic.
  6. Confirm whether rollback is a code revert or also needs an index snapshot restore, and test that restore once.
Executive Capability Standard

What Good Looks Like

A production-ready deploy process warms the index, pins the embedding model version explicitly, canaries on retrieval quality in addition to uptime, and has a tested rollback path that accounts for index schema changes.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Document your current deploy process end to end and mark which steps, if any, actually verify retrieval quality rather than just service health.
2. Do Manually:Run the warm-up and canary query set by hand before your next few deploys until the pattern is second nature to the team.
3. Delegate:Assign ownership of the deploy checklist, including the rollback test, to a specific engineer rather than leaving it as tribal knowledge.
4. Automate:Build the warm-up pass and retrieval-quality canary into your CI/CD pipeline so no deploy skips them under time pressure.
5. Buy:Bring in a platform or DevOps specialist if your team is shipping RAG changes often enough that manual gating is now the bottleneck.

How to Get Started

Frequently Asked Questions

Do we need a separate staging vector database, or can we use a smaller test collection in production?

Use a separate environment if you can. A small test collection in production shares the same cluster resources and access controls as real data, which defeats the purpose of isolating risk. If cost is the blocker, a smaller-tier managed instance for staging is usually still cheaper than an incident.

How do we test a rollback without affecting live traffic?

Restore your last index snapshot into a separate, unlabeled collection and run your canary query set against it. This confirms the restore process actually works and that the old code can read the restored format, without touching the collection serving production traffic.

What's the single most common go-live mistake with RAG pipelines?

Treating an embedding model upgrade as a routine deploy instead of a full reindex event. Queries embedded with a new model version drift out of alignment with an old document index in a way that degrades answers gradually rather than throwing an obvious error.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides