What to Fix in a Container Image Before It Ships
Container hardening gets treated as a one-time checklist run before a security review, when the more useful version of it is a set of checks that run on every image, every time, because the base image you hardened last quarter has almost certainly drifted since.
The checks below aren't exotic. Most teams already have the tooling to run every one of them; what's usually missing is making them a required, automated gate instead of a manual pass someone remembers to do before a big release.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Why shouldn't a container run as root?
A container running as root doesn't grant root on the host by default, but it removes a layer of defense if a container escape vulnerability is ever found, and it makes a compromised container more useful to an attacker than it needs to be. Set a non-root user explicitly in the image rather than relying on the base image's default, since defaults change between versions and silently reintroducing root is easy to miss in a routine base image bump.
Scan for known exploited vulnerabilities before you push, not after
A vulnerability scan that runs after an image is already in production is a detection mechanism, not a prevention one. Scan at build time and block the pipeline on anything matching a known exploited vulnerability, since those carry the shortest remediation clock of any category: federal guidance gives a 14-day window for a known exploited vulnerability with a CVE assigned in 2021 or later1, and that clock starts whether or not your pipeline caught it before shipping.
Keep the base image small and the update cadence tight
Every package in the base image is something that can carry a vulnerability, whether your application uses it or not. A minimal base image reduces that surface directly, and a short, automated cadence for rebuilding on top of an updated base image means a patched vulnerability upstream actually reaches your running containers instead of sitting unaddressed in an image nobody has rebuilt in months.
Why shouldn't you bake secrets into an image layer?
A secret added in one layer and removed in a later one is still recoverable from the image history, because layers are cumulative, not a clean overwrite. Use runtime secret injection or a secrets manager instead of copying credentials into the build context at any point, and audit existing images for this specific mistake, since it's one of the most common ways a credential ends up somewhere nobody intended.
Set resource limits so one container can't starve its neighbors
A container with no CPU or memory limit can consume enough of a shared host's resources to degrade or crash unrelated containers on the same node, turning what should be an isolated problem into a shared outage. Set explicit limits as a default part of your deployment configuration, not an afterthought added after the first incident where one runaway container took down several others.
Make the checklist a pipeline gate, not a habit
A checklist that depends on an engineer remembering to run it manually before a release will eventually be skipped, usually during exactly the rushed release where it mattered most. Wire each of these checks into the build pipeline as a required gate: non-root enforcement, vulnerability scan results, a secret-scanning pass over image layers, and a check for explicit resource limits in the deployment manifest. A gate that blocks the pipeline is a check that actually runs every time; a habit is a check that runs most of the time, and the gap between those two is where incidents live.
Treat a failed gate as information, not an obstacle to route around
The first weeks after turning these checks into hard gates usually produce a wave of blocked builds, as existing images catch up to the new bar. Resist the temptation to add a blanket exception list to get releases flowing again; instead, fix the specific images that fail and let the exception list, if you need one at all, stay short, named, and reviewed on a schedule so it doesn't quietly become the new default path around the gate you just built.
What Good Looks Like
A hardened container image runs as a non-root user, is scanned for known exploited vulnerabilities at build time with the pipeline blocked on hits, carries no secrets in any layer, and is rebuilt on a cadence tight enough that upstream patches actually reach production.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
CrowdStrike's runtime protection is the layer that catches an issue after an image is already deployed, which matters since new vulnerabilities in existing packages surface continuously after build time.
Tenable's vulnerability scanning is built specifically to catch known exploited vulnerabilities before an image ships, which is the check with the least room for delay given how short federal remediation windows are for that category.
Frequently Asked Questions
Is scanning at build time enough, or do we also need runtime scanning?
Both matter for different reasons. Build-time scanning stops a known bad image from shipping. Runtime scanning catches vulnerabilities discovered after an image is already deployed, which is common since new vulnerabilities in existing packages are found continuously, not only at the moment you built the image.
How often should base images be rebuilt?
Frequently enough that a security patch in the base image actually reaches your running containers in a reasonable window. A monthly or slower cadence leaves a meaningful gap; many teams rebuild weekly or trigger a rebuild automatically whenever the base image itself is updated upstream.
What's the most common hardening mistake teams make?
Treating hardening as a one-time review before a security audit rather than an ongoing set of checks that run on every build. A hardened image from last quarter, built on a base image that has since had several security updates, is not still a hardened image today.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Security patch remediation SLAs (CISA federal mandates, used as industry norm). CISA Binding Operational Directives 19-02 and 22-01 (CISA briefing hosted at NIST CSRC), 2022.
Related Guides
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
Auditing Security on Your MCP and Agent Tool Stack
A step-by-step way for a CTO to audit which tools an AI agent can reach, what each one can do, and where the access is broader than it should be.
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
Why Your Agent Loop Feels Slow, and How to Fix It
A diagnostic guide to finding where latency actually comes from in an agentic system, and which fixes help each cause instead of masking it.
Finding Your Agent Stack's Breaking Point Before Customers Do
A worked example of benchmarking an agent system's throughput, so you know where it actually breaks under load instead of guessing until it does.
How to Know If Your Agent Is Actually Working
Building an evaluation framework for an AI agent, from the first small test set through catching quality regressions before customers do.