Container Runtime Hardening: Rootless, Read-Only, and Fewer Capabilities
Scanning tells you what is inside the image. These four settings decide what it can do once it is running — and what breaks when you turn them on.
Scanning an image answers a question about its contents. It says nothing about what the process is permitted to do after it starts. Two images can contain identical, fully patched packages and still differ enormously in blast radius, because the difference is in the runtime flags.
A container escaping into its host almost never starts as a kernel exploit. It starts as a container that was running as root, with a writable filesystem, holding a capability nobody had a reason to give it.
Four settings close most of that gap.
1. Not root
Declare it in the image and enforce it at the platform:
RUN addgroup --system app && adduser --system --ingroup app app
USER app
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsNonRoot: true is the enforceable half — the kubelet rejects a pod whose user is 0 rather than trusting the Dockerfile to have done it. Set both, because the Dockerfile alone is a request and the security context is a rule.
Where the platform supports it, user namespaces add a second layer: the container's root maps to an unprivileged uid outside, so even a genuine escape lands as a no-privileged user rather than the host's root. It is the single most valuable setting that most clusters still have switched off.
2. Read-only root filesystem
An attacker who has landed in a container wants to write: a reverse shell, a cron job, an appended line in .bashrc. Make that harder by default:
securityContext:
readOnlyRootFilesystem: true
This is the setting that breaks things first, and it is worth knowing why before you roll it out. Applications write to /tmp, log to a directory they own, or build cache under the working directory. All of those now fail — which is useful information about your application, delivered as an error.
The fix is to declare the writable paths explicitly:
volumeMounts:
- name: tmp
mountPath: /tmp
volumes:
- name: tmp
emptyDir: {}
Apply it to one deployment at a time and read the errors. In my experience the errors fall into two buckets: something genuinely needs to write, and something was writing only because nobody ever asked whether it should.
3. Capabilities, dropped wholesale
Linux capabilities split the root privilege into named pieces. Containers are granted a short list by default — NET_RAW and NET_ADMIN among them — and almost no application needs any of it.
Start by dropping everything, then add back only what a real error proves you need:
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
add: ["NET_BIND_SERVICE"]
NET_BIND_SERVICE is the common one: a process binding port 80 or 443 without being root needs it and nothing else. allowPrivilegeEscalation: false additionally blocks the setuid and ptrace tricks that turn a small foothold into a larger one.
4. seccomp, and AppArmor if your runtime has it
seccomp filters which syscalls a process may make. The default profile already blocks the obviously dangerous ones, and turning it on explicitly means it cannot be quietly skipped:
securityContext:
seccompProfile:
type: RuntimeDefault
RuntimeDefault is the right answer almost always. Writing a custom profile is a project, not a setting, and an inaccurate one produces failures that look like application bugs. If your nodes run AppArmor, apparmor.dock/default on a Debian-based node is a similar low-effort addition.
Apply them in this order
runAsNonRoot— least disruptive, highest value.- Drop all capabilities — usually invisible until something needs
NET_BIND_SERVICE. readOnlyRootFilesystem— expect to addemptyDirmounts.seccompProfile: RuntimeDefault— should be a non-event.
Doing them in the other order means you debug three problems at once and cannot tell which setting caused which.
What to leave for its own change window
Custom seccomp profiles, mandatory SELinux policies on every node, and rootless container daemons on day one — that last one is valuable but changes how your images and build tooling behave, so it wants its own window.
Summary
Image scanning and runtime configuration answer different questions: one asks what is in the container, the other asks what it can do. Run as a non-root user and enforce it, make the root filesystem read-only with declared writable paths, drop every capability and add back only what an error proves necessary, and enable the default seccomp profile. Four settings, applied one at a time, and the realistic escape routes through a container narrow considerably.
SDP Clouds Team
DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.
More about us →