Hardening the Containers Behind Your RAG Inference Stack
The containers running a RAG pipeline's inference workloads, embedding, reranking, sometimes a self-hosted generation model, tend to get less security attention than the application containers around them, partly because they're often built from a base image supplied by a model framework or vendor that the team never fully audits.
Hardening them means treating inference containers with the same scrutiny as anything else that touches your corpus and, often, your model weights.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How do you audit the base image, not just your own layer?
Inference containers are frequently built on a base image from a machine learning framework that bundles a large dependency tree, and that base image is exactly what most teams skip auditing, assuming the vendor handled it. Scan the full image, including the base layers, for known vulnerabilities, and check the base image's own update cadence: an inference framework that hasn't published a base image update in a long time is a dependency risk even if your own application code is clean.
What least privilege should inference containers run with?
An embedding or generation container typically needs to read model weights and write nothing back except its response, but default configurations often run these containers with broader filesystem and network access than that. Run inference containers as a non-root user, mount model weights read-only, and restrict outbound network access to only the endpoints the workload actually calls, so a compromised inference container can't reach the vector database or the rest of your internal network on its own.
For example, imagine an embedding container that runs as root, has a writable weights directory, and can open outbound connections to any address. If a malicious input or a compromised dependency gives an attacker code execution inside it, that attacker can alter the weights, read the traffic passing through, and probe the internal network from a trusted position. Now imagine the same container running as a non-root user, with weights mounted read-only and outbound access limited to the one endpoint it needs. The same compromise now has almost nothing to reach. The workload is identical in both cases; only the permissions differ, and tightening them before an incident costs far less than explaining them after one.
Treat model weights as a sensitive artifact, not a static file
If you self-host a model, its weights are a valuable and sometimes proprietary artifact, and a compromised container with write access to the weights directory could tamper with them undetected. Store weights in read-only storage mounted into the container, verify their checksum against a known-good value at startup, and log any attempt to write to that path, since a legitimate inference workload should never need to.
Isolate inference workloads from each other, not just from the outside
If your embedding and generation workloads run as separate containers, don't assume they're isolated from each other just because they're separate. Check whether they share a network namespace or an overly permissive internal network policy that lets one call the other directly with no reason to. A compromised reranking container that can reach your generation model directly, bypassing your application's own access controls, has more reach than it should.
A container hardening checklist for RAG inference workloads
- Has the full image, including base layers, been scanned for known vulnerabilities?
- Does the container run as a non-root user with model weights mounted read-only?
- Is outbound network access restricted to only the endpoints the workload actually calls?
- Are model weight checksums verified at startup, with write attempts logged?
- Are inference containers isolated from each other on the network, not just from the outside world?
Rescan on a schedule, since the risk isn't static
A base image that was clean when you built your container six months ago isn't necessarily clean today; new vulnerabilities get disclosed against dependencies that haven't changed at all. Rescan inference images on a recurring schedule, not only when you rebuild them, so a newly disclosed vulnerability in a framework you haven't touched in months still gets caught.
Don't let inference containers become the exception to your normal process
It's common for inference containers to sit outside a team's normal container security process because they were set up quickly by whoever was building the RAG pipeline, not by whoever owns platform security. Bring them under the same review, scanning, and hardening process as every other production container instead of letting them stay a one-off exception, since the exception is exactly where the gaps tend to accumulate unnoticed.
This is worth confirming explicitly, not assuming: ask whoever owns your container security process whether inference workloads are actually included in their scope today, or whether they were quietly excluded because they looked different from a typical web service. The answer is often no, and it's a short conversation to fix once someone asks.
What Good Looks Like
Good container hardening for a RAG inference stack means every inference image, including its base layers, is scanned on a schedule, runs with least privilege, and model weights are protected as a sensitive artifact.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
CrowdStrike's workload protection can flag a compromised inference container behaving outside its expected network or filesystem pattern, which matters more here since inference workloads often run with broader default permissions than they need.
Tenable's vulnerability scanning is well suited to catching the base-image dependency risk in inference containers specifically, since that layer is the one teams most often assume a vendor already handled.
Frequently Asked Questions
Do inference containers need different security treatment than application containers?
Yes, mainly because they're often built on a vendor-supplied base image with a large dependency tree that teams assume is already secure and skip auditing. Scan the full image including base layers, and don't treat a model framework's base image as exempt from the same scrutiny as your own code.
How should model weights be protected inside a container?
Mount them read-only, verify their checksum against a known-good value at startup, and log any write attempt to that path, since a legitimate inference workload never needs to modify its own weights. A compromised container with write access could tamper with them undetected otherwise.
Should embedding and generation containers be able to talk to each other directly?
Only if there's a real reason for it, and usually there isn't. Check whether separate inference containers share a network namespace or an overly permissive internal policy, since a compromised one that can reach another directly, bypassing your application's access controls, has more reach than it should.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
What to Check First in a RAG Pipeline Security Audit
A practical order of operations for auditing a production RAG pipeline: data exposure, prompt injection, access control, logging, and vendor risk.
Zero Trust for a RAG Pipeline Means No Service Gets a Free Pass
A decision framework for applying zero trust to a production RAG pipeline: verifying every service and user call, not just the ones at the edge.
Hardening Containers: The Checks That Actually Stop Real Attacks
Which container hardening steps actually reduce risk, versus the ones that mostly look good on a checklist without stopping much.
Where RAG Latency Actually Goes, and How to Budget It
Break a RAG request into its four latency stages, find out which one is actually slow, and set a budget for each before you start tuning blindly.
How Vector Search Throughput Degrades as Your Index Grows
Throughput doesn't fall off gradually as a vector index grows. Here's why it degrades in steps, and how sharding, replicas, and quantization each help.