Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

What to Check First in a RAG Pipeline Security Audit

A RAG pipeline is not just an application to audit. It's a data pipeline (the ingestion and embedding jobs), a database (the vector store), and an attack surface (anything an attacker can get indexed and later retrieved). Most security reviews start at the API layer and never make it to the other two.

Here's an order that catches the parts teams usually skip.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What can your retriever actually touch?

Start with the blast radius, not the code. List every collection or index in your vector store, then list who or what can query each one: end users through the app, background jobs, internal tools, and any service account with a broad API key. It's common to find a single service-role key from an early prototype that still has read access to every collection months after the team moved to per-tenant scoping.

For each collection, write down what happens if a user in tenant A can retrieve a chunk that belongs to tenant B. If the answer is "that would be bad," fix the scoping problem before anything else on this list matters.

Don't stop at query-time access. Ingestion pipelines often run with elevated database credentials to write new vectors, and a compromised ingestion job can silently poison or delete an entire collection. Give ingestion its own scoped credential, separate from anything a user-facing request path uses.

Treat retrieved content as untrusted input

Anything that ends up in your vector store, a support ticket, a scraped page, an uploaded PDF, can later be retrieved and placed directly into a prompt. If an attacker can get a document indexed through a public form, a shared inbox, or a crawled site, they can plant instructions in it. Hidden text on a PDF page telling the model to ignore its instructions and reveal the system prompt is a real pattern, not a hypothetical.

Audit your ingestion path for anything that accepts external content, and check whether your prompt template separates retrieved text from instructions clearly enough that the model treats it as data, not as commands.

Verify encryption and access control on the vector store itself

Vector databases get treated as an implementation detail, so they're often the last system to get the same encryption-at-rest and network policy review as the primary database. Check three things: whether embeddings and their source text are encrypted at rest, whether the vector store's endpoint is reachable from outside your VPC or service mesh, and whether collection-level access control exists or whether any authenticated caller can query everything.

If you're on a managed vector database, read its shared responsibility model directly. Some products encrypt storage by default; others leave it opt-in.

Are your retrieval queries logged at all?

Ask your team a simple question: if a customer asks what data of theirs the system accessed last month, can you answer it? Many RAG deployments log the final model response but not the retrieval step, so there's no record of which chunks were pulled for which query. The federal remediation windows agencies work under for known vulnerabilities1 are a useful reference for how fast you'd need to move if a logging gap turned into an actual disclosure question.

At minimum, log the query, the collection queried, the caller's identity, and the chunk IDs returned, even if you don't log full chunk content.

Close the vendor risk gap

Your embedding provider and your vector database vendor both see your data, often the full source text alongside the vectors. Confirm each vendor's retention policy, whether they train on customer data by default, and whether a signed data processing agreement is in place if you handle regulated data. This is a short check per vendor that gets skipped because it doesn't feel like "real" security work.

Ask specifically about subprocessors, since an embedding API is sometimes itself a thin wrapper around a third vendor's model. A data processing agreement with your direct vendor doesn't automatically cover a subprocessor two layers down, and that's exactly where source text tends to actually get processed.

If you're comparing compliance automation platforms to keep vendor reviews current, see our Vanta vs. Drata vs. Secureframe comparison.

Have a plan ready before you need it

An audit is only useful if it changes what happens during an actual incident. Write a short runbook for the specific failure this pipeline invites: a tenant-scoping bug that exposed another customer's chunks, or a prompt injection that got a model to leak part of a retrieved document. Name who gets paged, what gets disabled first, usually the affected collection rather than the whole service, and who owns customer notification if the exposure turns out to be real.

Test this runbook once with a tabletop exercise, walking through the steps out loud without touching production. It surfaces gaps a written checklist alone won't.

Run the audit in this order:

  1. List every collection in the vector store and every user, job, tool, and service account that can query or write to it.
  2. Treat retrieved content as untrusted input, since anything indexed can later be placed into a prompt, including planted instructions.
  3. Check that embeddings and source text are encrypted at rest and that the vector store endpoint is not reachable from outside your network.
  4. Confirm retrieval queries are logged, so you can answer what data of a given customer the system accessed.
  5. Review each vendor's retention policy, default training on customer data, and data processing agreement.
  6. Write a short runbook for the failures this pipeline invites, naming who gets paged and what gets disabled first.
Executive Capability Standard

What Good Looks Like

The audit standard is a written map of every collection, who can query it, what's logged, and which vendors touch the data, reviewed on a fixed cadence rather than only after an incident.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read through your ingestion code and vector database configuration to list every collection and every caller that can query it.
2. Do Manually:Walk through the five checks above by hand once, documenting findings in a shared doc before you fix anything.
3. Delegate:Assign a senior engineer to own the audit checklist and re-run it on a fixed schedule, not just when someone remembers.
4. Automate:Wire retrieval logging and access-control checks into CI so a new collection or a broadened API key fails a build instead of shipping quietly.
5. Buy:Bring in a compliance automation platform or an outside security review once customer data is at stake and you need continuous evidence, not a one-time checklist.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Can someone reconstruct the original text from an embedding vector?

Partially, in some cases. Embedding inversion techniques can recover fragments or close paraphrases of source text from a vector, especially with access to the embedding model itself. Treat embeddings as sensitive data, not as an anonymized representation, and apply the same access controls you'd apply to the source documents.

How often should we re-run this audit?

Re-run it whenever you add a new data source to the index, change who can query a collection, or switch embedding or vector database vendors. Outside those triggers, a semiannual review catches drift from smaller changes that add up, like a new internal tool quietly getting broad query access.

Do we need our vector database vendor to be SOC 2 certified?

If you handle customer data or anything covered by a compliance framework you've committed to, treat the vendor the same as any other data processor. If you're early-stage and pre-compliance, at minimum get their security documentation in writing before you index anything sensitive.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Security patch remediation SLAs (CISA federal mandates, used as industry norm). CISA Binding Operational Directives 19-02 and 22-01 (CISA briefing hosted at NIST CSRC), 2022.

Related Guides