Enforcing Document-Level Permissions in Multi-Tenant RAG
A multi-tenant RAG system has a specific way to fail quietly: retrieval works perfectly, the answers are fluent and relevant, and one of them includes a sentence lifted from a document the requesting user was never supposed to see. Here's how to prevent that at the retrieval layer itself, not just at the application layer around it.
Should you filter before or after the vector search?
Pre-filtering restricts the candidate set to permitted documents before running the nearest-neighbor search, which guarantees correct results but needs your vector index to support metadata filtering efficiently at search time. Post-filtering runs a broader search first and drops disallowed results afterward, which is simpler to implement but can return fewer than the requested top-k results if many of the nearest matches turn out to be restricted.
For a system where most users can see most content, post-filtering is often good enough. For a system where access is narrow and specific, like a document shared with exactly one other user, pre-filtering is usually worth the added index complexity.
Measure the actual shortfall from post-filtering before assuming it's a problem. Requesting a larger candidate set than you need, then filtering and truncating to the real top-k, is a cheap way to absorb some of the loss without committing to pre-filtering's added complexity right away.
Store permission metadata on the chunk, not just the document
A document-level permission check requires a second lookup for every chunk returned, adding latency and a dependency on another system being available. Storing the owning tenant, and any relevant access list, as metadata directly on each chunk lets the vector database filter at query time without that extra round trip. It does mean permission data now lives in two places, the source system and the index, which brings its own consistency question.
How do you avoid an N+1 permission check?
Checking each of the top-k results against an external permission service in real time, one call per result, adds meaningful latency and introduces a race condition if permissions change between the check and the response being returned. Encoding permission state redundantly in the index, then refreshing it when the source system's permissions change, avoids both problems at the cost of some staleness between a permission change and the index catching up.
Handle permission changes without a full reindex
A document's access changing, a user losing access, a document moving between tenants, doesn't change its content or its vector, only the filter attached to it. Build a lightweight metadata update path that patches the permission fields on existing chunks in place, separate from your full ingestion and embedding pipeline. Routing every permission change through a full reindex wastes embedding cost on content that hasn't actually changed and adds unnecessary lag before the permission change takes effect.
Offboarding is the case that exposes this gap fastest. When a user loses access to a whole tenant, that's potentially thousands of chunks needing a metadata update at once, and a system that can only apply permission changes through a full reindex will leave that access open for however long the next scheduled reindex is away.
Test with a negative case, not just a positive one
It's easy to test that a user with access gets the right results and stop there. The test that actually matters is confirming a user without access gets zero results from a restricted collection, not degraded or partial results. Build this into your test suite as an explicit case, a specific document, a specific unauthorized user, an assertion that it never appears, and run it on every change to the permission or filtering logic.
Verify document-level permissions with these checks:
- Confirm that a user without access gets zero results from a restricted collection, using a specific document and a specific unauthorized user.
- Store the owning tenant and any access list as metadata on each chunk, so the vector database can filter at query time.
- Patch permission metadata in place when access changes, rather than routing every change through a full reindex.
- Check generated answers for leakage, where a permitted chunk quotes or summarizes a restricted document.
- Use database row-level security as a second layer where it is available, so a dropped filter is caught.
Watch for leakage through the generated answer, not just the retrieved chunks
A subtler failure mode: retrieval correctly filters out a restricted chunk, but a different, permitted chunk still references information from it, a summary document that quotes a restricted one, for example, and the model repeats that reference in its answer. Filtering at the vector search layer doesn't catch this, since the offending content technically lives inside a chunk the user is allowed to see. This is worth a specific test case of its own, checking not just what gets retrieved but what leaks through indirectly.
What Good Looks Like
The access control standard is permission metadata stored on each chunk and enforced at the retrieval layer itself, refreshed without a full reindex, and verified with an explicit negative test that unauthorized users get zero results.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Can row-level security in the underlying database handle this instead of application code?
If your vector database sits on top of a database with native row-level security, like Postgres with pgvector, it's worth using that as a second layer of defense rather than relying only on application-level filtering. It won't replace the need for correct metadata, but it catches the case where a filter gets dropped by a bug in application code.
What's the biggest mistake teams make with RAG access control?
Assuming the application layer's authorization checks automatically extend to the retrieval layer. A developer adds a new internal tool that queries the vector store directly, bypassing the application's normal access checks entirely, and nobody notices until an audit or an incident. Enforce filtering at the retrieval layer itself, not only in the application calling it.
How do we handle a document that's shared with only some members of a tenant?
Store an explicit access list on the chunk rather than only a tenant ID, and filter on membership in that list. Tenant-level filtering alone can't express partial sharing within a tenant, so if that's a real requirement, the permission model needs to support it from the start rather than being bolted on later.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Zero Trust for a RAG Pipeline Means No Service Gets a Free Pass
A decision framework for applying zero trust to a production RAG pipeline: verifying every service and user call, not just the ones at the edge.
Where RAG Latency Actually Goes, and How to Budget It
Break a RAG request into its four latency stages, find out which one is actually slow, and set a budget for each before you start tuning blindly.
Designing Role-Based Access Control That Survives Your Next Reorg
A worksheet approach to mapping roles to permissions so access control doesn't quietly rot every time your team's structure changes.
How Vector Search Throughput Degrades as Your Index Grows
Throughput doesn't fall off gradually as a vector index grows. Here's why it degrades in steps, and how sharding, replicas, and quantization each help.
Building a Golden Set to Catch RAG Regressions Before Users Do
A step-by-step approach to building a RAG evaluation set from real queries, scoring retrieval and generation separately, and gating on regressions.