What to Check First in a RAG Pipeline Security Audit
A RAG pipeline is not just an application to audit. It's a data pipeline (the ingestion and embedding jobs), a database (the vector store), and an attack surface (anything an attacker can get indexed and later retrieved). Most security reviews start at the API layer and never make it to the other two.
Here's an order that catches the parts teams usually skip.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
What can your retriever actually touch?
Start with the blast radius, not the code. List every collection or index in your vector store, then list who or what can query each one: end users through the app, background jobs, internal tools, and any service account with a broad API key. It's common to find a single service-role key from an early prototype that still has read access to every collection months after the team moved to per-tenant scoping.
For each collection, write down what happens if a user in tenant A can retrieve a chunk that belongs to tenant B. If the answer is "that would be bad," fix the scoping problem before anything else on this list matters.
Don't stop at query-time access. Ingestion pipelines often run with elevated database credentials to write new vectors, and a compromised ingestion job can silently poison or delete an entire collection. Give ingestion its own scoped credential, separate from anything a user-facing request path uses.
Treat retrieved content as untrusted input
Anything that ends up in your vector store, a support ticket, a scraped page, an uploaded PDF, can later be retrieved and placed directly into a prompt. If an attacker can get a document indexed through a public form, a shared inbox, or a crawled site, they can plant instructions in it. Hidden text on a PDF page telling the model to ignore its instructions and reveal the system prompt is a real pattern, not a hypothetical.
Audit your ingestion path for anything that accepts external content, and check whether your prompt template separates retrieved text from instructions clearly enough that the model treats it as data, not as commands.
Verify encryption and access control on the vector store itself
Vector databases get treated as an implementation detail, so they're often the last system to get the same encryption-at-rest and network policy review as the primary database. Check three things: whether embeddings and their source text are encrypted at rest, whether the vector store's endpoint is reachable from outside your VPC or service mesh, and whether collection-level access control exists or whether any authenticated caller can query everything.
If you're on a managed vector database, read its shared responsibility model directly. Some products encrypt storage by default; others leave it opt-in.
Are your retrieval queries logged at all?
Ask your team a simple question: if a customer asks what data of theirs the system accessed last month, can you answer it? Many RAG deployments log the final model response but not the retrieval step, so there's no record of which chunks were pulled for which query. The federal remediation windows agencies work under for known vulnerabilities1 are a useful reference for how fast you'd need to move if a logging gap turned into an actual disclosure question.
At minimum, log the query, the collection queried, the caller's identity, and the chunk IDs returned, even if you don't log full chunk content.
Close the vendor risk gap
Your embedding provider and your vector database vendor both see your data, often the full source text alongside the vectors. Confirm each vendor's retention policy, whether they train on customer data by default, and whether a signed data processing agreement is in place if you handle regulated data. This is a short check per vendor that gets skipped because it doesn't feel like "real" security work.
Ask specifically about subprocessors, since an embedding API is sometimes itself a thin wrapper around a third vendor's model. A data processing agreement with your direct vendor doesn't automatically cover a subprocessor two layers down, and that's exactly where source text tends to actually get processed.
If you're comparing compliance automation platforms to keep vendor reviews current, see our Vanta vs. Drata vs. Secureframe comparison.
Have a plan ready before you need it
An audit is only useful if it changes what happens during an actual incident. Write a short runbook for the specific failure this pipeline invites: a tenant-scoping bug that exposed another customer's chunks, or a prompt injection that got a model to leak part of a retrieved document. Name who gets paged, what gets disabled first, usually the affected collection rather than the whole service, and who owns customer notification if the exposure turns out to be real.
Test this runbook once with a tabletop exercise, walking through the steps out loud without touching production. It surfaces gaps a written checklist alone won't.
Run the audit in this order:
- List every collection in the vector store and every user, job, tool, and service account that can query or write to it.
- Treat retrieved content as untrusted input, since anything indexed can later be placed into a prompt, including planted instructions.
- Check that embeddings and source text are encrypted at rest and that the vector store endpoint is not reachable from outside your network.
- Confirm retrieval queries are logged, so you can answer what data of a given customer the system accessed.
- Review each vendor's retention policy, default training on customer data, and data processing agreement.
- Write a short runbook for the failures this pipeline invites, naming who gets paged and what gets disabled first.
What Good Looks Like
The audit standard is a written map of every collection, who can query it, what's logged, and which vendors touch the data, reviewed on a fixed cadence rather than only after an incident.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Useful once you need continuous evidence that vector store access controls and vendor agreements stay in place, not just a one-time checklist.
Fits the same role as Vanta here: ongoing audit evidence for the access-control and vendor-risk items above, instead of a point-in-time review.
Frequently Asked Questions
Can someone reconstruct the original text from an embedding vector?
Partially, in some cases. Embedding inversion techniques can recover fragments or close paraphrases of source text from a vector, especially with access to the embedding model itself. Treat embeddings as sensitive data, not as an anonymized representation, and apply the same access controls you'd apply to the source documents.
How often should we re-run this audit?
Re-run it whenever you add a new data source to the index, change who can query a collection, or switch embedding or vector database vendors. Outside those triggers, a semiannual review catches drift from smaller changes that add up, like a new internal tool quietly getting broad query access.
Do we need our vector database vendor to be SOC 2 certified?
If you handle customer data or anything covered by a compliance framework you've committed to, treat the vendor the same as any other data processor. If you're early-stage and pre-compliance, at minimum get their security documentation in writing before you index anything sensitive.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Security patch remediation SLAs (CISA federal mandates, used as industry norm). CISA Binding Operational Directives 19-02 and 22-01 (CISA briefing hosted at NIST CSRC), 2022.
Related Guides
Vanta vs Drata vs Secureframe: Best SOC 2 Automation Platform
Comparing Vanta, Drata, and Secureframe: API evidence collection, auditor networks, true costs, and when each platform is the wrong choice.
What Actually Belongs in Your RAG Audit Log (and What Doesn't)
A framework for deciding what a production RAG system's audit log should capture, how long to keep it, and when to redact retrieved content.
Mapping SOC 2 Controls to a RAG Pipeline's Real Components
SOC 2 auditors ask about access, change management, and vendors in the abstract. Here's what each control actually maps to in a RAG pipeline.
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Why Your RAG Infrastructure Drifted From What Terraform Says It Should Be
A walkthrough of how production RAG infrastructure drifts from its IaC definitions, and the governance practices that catch it before an incident does.
Zero Trust for a RAG Pipeline Means No Service Gets a Free Pass
A decision framework for applying zero trust to a production RAG pipeline: verifying every service and user call, not just the ones at the edge.