Where RAG Systems Actually Leak Data Over the Network
A RAG system has more network paths than a typical application: the retriever calls the vector database, the vector database's embedding step might call an outside provider, the generation step calls an LLM API, and logging pipes everything to a third place. Each of those hops is somewhere data can leave your network boundary without anyone deciding it should.
Isolating this properly means mapping every hop your query and retrieved content travel through, not just locking down the front door.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Why map every hop before you isolate anything?
Draw the actual path a single query takes: from the user-facing API, to the embedding call, often an external API unless you self-host, to the vector database, to any reranking service, to the generation model, and back. Include the logging and observability pipeline in this map too, since it usually receives a copy of the query and sometimes the retrieved content, which makes it a hop with real data exposure even though nobody built it to serve queries. Most teams isolate the vector database itself carefully and then discover the embedding call goes over the public internet to a third-party API with the full query text in the request body. You can't isolate a hop you haven't mapped.
Should the vector database be reachable from the public internet?
A production vector database should sit inside a private network with no public endpoint, reachable only from the services that need it, through a VPC peering connection or a private link, not an IP allowlist on a public address. An IP allowlist is a network policy that depends on nobody's IP changing and nobody misconfiguring it; a private network path removes the exposure by design instead of relying on a rule staying correct forever. This matters even for internal tooling: a debugging script that queries the vector database directly, bypassing your application layer, often gets exempted from the private-network rule because it's just for engineers, and that exemption is exactly the kind of exception an incident later traces back through.
Decide, per hop, whether the data leaving is acceptable
Some external calls are unavoidable: if you use a hosted embedding or generation model, query text leaves your network by design, and that's a vendor and data-handling decision, not a network configuration one. The isolation question is whether every other hop, the ones that don't need to leave, actually stays contained. Retrieved document content going to a logging pipeline outside your network boundary, for instance, is usually accidental rather than a deliberate choice anyone signed off on. Write this decision down per hop: which ones are allowed to leave the network boundary and why, versus which ones are misconfigurations waiting to be found by whoever runs the next security review.
A network isolation checklist for a RAG stack
- Does the vector database have a public endpoint, even one behind an IP allowlist?
- Does every external API call, embedding, reranking, generation, go through a documented, reviewed path?
- Is service-to-service traffic between the retriever, vector database, and generation step encrypted, not just isolated?
- Does your logging or observability pipeline receive retrieved content, and does it live inside the same network boundary?
- Has anyone actually traced a live query through the network, or only read the architecture diagram?
Revisit isolation after every new integration, not just at launch
Network isolation tends to be solid at launch and erode afterward: someone adds a new reranking service, a new logging destination, or a debug endpoint that's supposed to be temporary, and each addition is a new hop nobody re-mapped. Treat network isolation as something you check whenever a new service joins the RAG pipeline, not a diagram you drew once and archived.
A lightweight way to catch this drift is to assign someone to trace one live query through the whole pipeline each quarter, from the user-facing request through every hop back to the response, and compare what they find against the last diagram. New hops show up in that exercise long before they show up in a security review, because whoever added the integration usually isn't the one asking whether it changed the network boundary.
What Good Looks Like
Good network isolation for a RAG system means you can trace every hop a single query takes, end to end, and state which ones are allowed to leave your network boundary and why.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Network isolation for sensitive data flows is a control Drata expects evidence for; documenting the hop map above gives you something concrete to show instead of reconstructing it during the audit.
Vanta covers the same control category, and either is easier to satisfy when the network diagram already exists than when you're drawing it for the first time under audit pressure.
Frequently Asked Questions
Should a production vector database ever have a public endpoint?
No. It should sit behind a private network path, reachable only by the services that need it, whether that's VPC peering, a private link, or an equivalent your cloud provider offers. An IP allowlist on a public endpoint is a policy that can be misconfigured; a private network path removes the exposure by design.
Is it a network isolation problem if my embedding provider is a public API?
Not by itself. Using a hosted embedding or generation API means query text leaves your network on purpose, which is a vendor and data-handling decision. The isolation question is whether every hop that doesn't need to leave your network, like your vector database or logging pipeline, actually stays contained.
How often should we re-check network isolation for a RAG system?
Every time a new service joins the pipeline, not just at launch. A new reranking service, logging destination, or debug endpoint each adds a hop that needs the same review the original architecture got, and isolation quietly erodes when that review gets skipped.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Where VPC Peering Breaks Down and How to Isolate Blast Radius Instead
Why VPC peering alone doesn't isolate anything, and a more reliable way to contain blast radius between services and environments.
What a Misconfigured VPC Peering Connection Actually Breaks
A walkthrough of a real VPC peering misconfiguration, what it exposed, and the four checks that would have caught it before it shipped.
The VPC Peering Mistake That Opens Your Whole Network
How VPC peering misconfigurations quietly expose more of your network than intended, and four concrete checks that catch the mistake before an audit does.
A Practical Checklist for VPC Peering and Network Isolation
The specific network isolation mistakes that quietly undermine a zero-trust architecture, and a checklist for catching them in your VPC peering setup.
VPC Peering Looks Like Isolation Until You Check the Routes
VPC peering can silently become transitive, undoing the isolation you thought you had. Here is how to audit what can actually reach what in your network.
Four Network Isolation Checks Most VPC Peering Setups Skip
Four specific checks for VPC peering and network isolation setups, plus the pitfalls that let a segmentation boundary look correct while quietly failing.