Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

Hardening the Containers Behind Your RAG Inference Stack

The containers running a RAG pipeline's inference workloads, embedding, reranking, sometimes a self-hosted generation model, tend to get less security attention than the application containers around them, partly because they're often built from a base image supplied by a model framework or vendor that the team never fully audits.

Hardening them means treating inference containers with the same scrutiny as anything else that touches your corpus and, often, your model weights.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you audit the base image, not just your own layer?

Inference containers are frequently built on a base image from a machine learning framework that bundles a large dependency tree, and that base image is exactly what most teams skip auditing, assuming the vendor handled it. Scan the full image, including the base layers, for known vulnerabilities, and check the base image's own update cadence: an inference framework that hasn't published a base image update in a long time is a dependency risk even if your own application code is clean.

What least privilege should inference containers run with?

An embedding or generation container typically needs to read model weights and write nothing back except its response, but default configurations often run these containers with broader filesystem and network access than that. Run inference containers as a non-root user, mount model weights read-only, and restrict outbound network access to only the endpoints the workload actually calls, so a compromised inference container can't reach the vector database or the rest of your internal network on its own.

For example, imagine an embedding container that runs as root, has a writable weights directory, and can open outbound connections to any address. If a malicious input or a compromised dependency gives an attacker code execution inside it, that attacker can alter the weights, read the traffic passing through, and probe the internal network from a trusted position. Now imagine the same container running as a non-root user, with weights mounted read-only and outbound access limited to the one endpoint it needs. The same compromise now has almost nothing to reach. The workload is identical in both cases; only the permissions differ, and tightening them before an incident costs far less than explaining them after one.

Treat model weights as a sensitive artifact, not a static file

If you self-host a model, its weights are a valuable and sometimes proprietary artifact, and a compromised container with write access to the weights directory could tamper with them undetected. Store weights in read-only storage mounted into the container, verify their checksum against a known-good value at startup, and log any attempt to write to that path, since a legitimate inference workload should never need to.

Isolate inference workloads from each other, not just from the outside

If your embedding and generation workloads run as separate containers, don't assume they're isolated from each other just because they're separate. Check whether they share a network namespace or an overly permissive internal network policy that lets one call the other directly with no reason to. A compromised reranking container that can reach your generation model directly, bypassing your application's own access controls, has more reach than it should.

A container hardening checklist for RAG inference workloads

  • Has the full image, including base layers, been scanned for known vulnerabilities?
  • Does the container run as a non-root user with model weights mounted read-only?
  • Is outbound network access restricted to only the endpoints the workload actually calls?
  • Are model weight checksums verified at startup, with write attempts logged?
  • Are inference containers isolated from each other on the network, not just from the outside world?

Rescan on a schedule, since the risk isn't static

A base image that was clean when you built your container six months ago isn't necessarily clean today; new vulnerabilities get disclosed against dependencies that haven't changed at all. Rescan inference images on a recurring schedule, not only when you rebuild them, so a newly disclosed vulnerability in a framework you haven't touched in months still gets caught.

Don't let inference containers become the exception to your normal process

It's common for inference containers to sit outside a team's normal container security process because they were set up quickly by whoever was building the RAG pipeline, not by whoever owns platform security. Bring them under the same review, scanning, and hardening process as every other production container instead of letting them stay a one-off exception, since the exception is exactly where the gaps tend to accumulate unnoticed.

This is worth confirming explicitly, not assuming: ask whoever owns your container security process whether inference workloads are actually included in their scope today, or whether they were quietly excluded because they looked different from a typical web service. The answer is often no, and it's a short conversation to fix once someone asks.

Executive Capability Standard

What Good Looks Like

Good container hardening for a RAG inference stack means every inference image, including its base layers, is scanned on a schedule, runs with least privilege, and model weights are protected as a sensitive artifact.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your inference framework's documentation on its base image contents and update cadence, since that's usually the least-audited part of the container.
2. Do Manually:Manually review container run configurations, user, filesystem mounts, network policy, for your inference workloads at least once, checking each against least privilege.
3. Delegate:Give one engineer ownership of inference container hardening as a distinct responsibility from general application container security, since the risk profile is different.
4. Automate:Set up recurring vulnerability scans on inference images, including base layers, so a newly disclosed vulnerability gets caught without a manual rebuild.
5. Buy:Use a container security platform that scans base images and enforces runtime policies instead of stitching together your own scanning and enforcement tooling.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Do inference containers need different security treatment than application containers?

Yes, mainly because they're often built on a vendor-supplied base image with a large dependency tree that teams assume is already secure and skip auditing. Scan the full image including base layers, and don't treat a model framework's base image as exempt from the same scrutiny as your own code.

How should model weights be protected inside a container?

Mount them read-only, verify their checksum against a known-good value at startup, and log any write attempt to that path, since a legitimate inference workload never needs to modify its own weights. A compromised container with write access could tamper with them undetected otherwise.

Should embedding and generation containers be able to talk to each other directly?

Only if there's a real reason for it, and usually there isn't. Check whether separate inference containers share a network namespace or an overly permissive internal policy, since a compromised one that can reach another directly, bypassing your application's access controls, has more reach than it should.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides