Overview
I used Docker for four years before I understood what it actually does. The mental model I had was "lightweight VM," which is wrong in every direction that matters. A container is a Linux process that's been lied to about what machine it's on, plus a resource limit that the kernel enforces.
That's the whole thing. Two kernel features — namespaces and cgroups — and a filesystem layered on top. Once you see it, a lot of Docker behavior stops being mysterious.
Namespaces: the isolation part
A namespace wraps a global system resource and makes it look private to the processes inside. There are eight of them on a modern Linux system:
| Namespace | Isolates | Since |
|---|---|---|
mnt | Filesystem mount points | 2.4.19 |
pid | Process IDs | 2.6.24 |
net | Network interfaces, routes, ports | 2.6.24 |
ipc | System V IPC, POSIX message queues | 2.6.19 |
uts | Hostname and domain name | 2.6.19 |
user | User and group IDs | 3.8 |
cgroup | cgroup root directory | 4.6 |
time | Boot and monotonic clocks | 5.6 |
You can create a namespace with the unshare command:
# Run bash in a new UTS namespace
sudo unshare --uts /bin/bash
# Inside this shell:
hostname my-container
hostname
# my-container
exit
# Back in your shell:
hostname
# your-actual-hostname
You just changed the hostname of the system from inside a shell, and it didn't affect anything outside. That's the entire mechanism behind a container having its own hostname. It's a UTS namespace.
The process and network namespaces are the interesting ones:
sudo unshare --pid --fork --mount-proc /bin/bash
# Inside:
ps aux
# PID 1 is bash, plus maybe one or two others
# You cannot see any process outside this namespace
The --mount-proc is needed because /proc is a mount that's shared. Without remounting it, ps would still show all the host's processes even though the PID namespace is new.
cgroups: the limit part
Namespaces control what a process can see. cgroups control what it can use. CPU, memory, disk I/O, number of processes — all enforced by the kernel, not by any userspace process.
cgroup v2 (the current version, unified hierarchy) is organized as a filesystem tree:
/sys/fs/cgroup/
├── system.slice/
│ ├── sshd.service/
│ └── cron.service/
├── user.slice/
└── my-container/
├── cgroup.controllers
├── memory.max
├── cpu.max
└── pids.max
To limit a process tree to 512MB of memory:
sudo mkdir /sys/fs/cgroup/my-test
echo "+memory +cpu +pids" | sudo tee /sys/fs/cgroup/my-test/cgroup.subtree_control
echo "512M" | sudo tee /sys/fs/cgroup/my-test/memory.max
echo "200000 100000" | sudo tee /sys/fs/cgroup/my-test/cpu.max # 2 CPUs
echo "100" | sudo tee /sys/fs/cgroup/my-test/pids.max
echo $$ | sudo tee /sys/fs/cgroup/my-test/cgroup.procs
# Any process in this shell is now limited
The cpu.max format is "quota period" in microseconds. 200000/100000 means 2 CPU-seconds per 1 CPU-second, which is two cores' worth.
If a process exceeds memory.max, the kernel kills it. The OOM killer runs, the process dies, and if you're running Docker, the container exits with code 137 — which is 128 + 9 (SIGKILL). That's why that specific exit code comes up when a container runs out of memory.
Putting them together: what Docker actually does
When you run docker run nginx, roughly this happens:
- Docker creates a new set of namespaces:
pid,net,mnt,uts,ipc. Optionallyuserandcgroup. - It creates a cgroup directory and writes the memory, CPU, and PID limits you specified.
- It sets up a filesystem — either a single directory (overlayfs, or a bind mount), which becomes the container's
/. - It calls
clone()with flags for each namespace, forks the container's initial process (PID 1 inside the namespace), and moves it into the cgroup. - The process runs. From its perspective, it's alone on a machine with one filesystem, one network interface, and a hostname it set itself.
You can do all of this manually with unshare, ip, and cgroup writes. It's tedious but not complicated, and doing it once makes Docker's behavior obvious.
Reproducing a container by hand
# Create isolated namespaces and a cgroup
sudo mkdir -p /sys/fs/cgroup/mycontainer
echo "+memory +cpu +pids" | sudo tee /sys/fs/cgroup/cgroup.subtree_control 2>/dev/null
echo "256M" | sudo tee /sys/fs/cgroup/mycontainer/memory.max
# Unshare everything at once, run bash
sudo unshare \
--uts --ipc --pid --net --mount --fork --mount-proc \
--cgroup /sys/fs/cgroup/mycontainer \
/bin/bash
# Inside, you're in a fresh namespace
hostname isolated
ip link set lo up
ip link show
# only the loopback interface
The --cgroup flag takes a path and moves the process into that cgroup. Now any process you spawn from this shell is memory-limited to 256MB.
Try running something that allocates memory:
Python3 -c "x = 'a' * (300 * 1024 * 1024)"
# Killed
That's what an OOM-killed container looks like from inside.
What namespaces don't do
| Isolation | What it covers | What it doesn't |
|---|---|---|
| Filesystem | Mount points | A process with CAP_SYS_ADMIN can still see host mounts |
| Process | Visible PIDs | Kernel threads are still shared |
| Network | Interfaces and routes | No firewalling; eBPF on the host sees everything |
| User | UID/GID mapping | Kernel CVEs can escape |
The big one: containers are not a security boundary the way VMs are. The kernel is shared. A container escape is a kernel exploit that gets you root on the host. This is why people run gVisor or Firecracker (which AWS Lambda uses) when they need real isolation.
Docker does reduce the attack surface significantly: no CAP_SYS_ADMIN by default, a seccomp filter blocking ~40 syscalls, AppArmor/SELinux profiles. But "container = sandbox" is a simplification that gets people into trouble.
Namespaces you can see right now
Every process on your system is in some namespace. You can see them:
# List namespaces for the current process
ls -l /proc/self/ns/
# Compare with PID 1
ls -l /proc/1/ns/
# See all network namespaces
sudo lsns -t net
# Check which namespace PID X is in
sudo readlink /proc/<pid>/ns/net
Two processes with the same namespace inode number are in the same namespace. This is how you find which container a process belongs to when you're debugging on the host.
docker inspect --format '{{.State.Pid}}' my-container
# 28431
sudo ls -l /proc/28431/ns/
The net namespace inode in the container's process matches what lsns -t net shows for the container. That's how you can jump from a container to the host's view of it, or vice versa.
Why this matters in practice
Four situations where knowing namespaces and cgroups helps:
- Debugging a container that "can't reach the network." The container has its own net namespace. If you're expecting it to share the host's, you need
--network host. Otherwise, it's in a veth pair with an IP that only exists inside the namespace. - Understanding memory limits. A Java process in a container with 2GB limit and no JVM
-Xmxflag will try to use the host's memory and get OOM-killed. The JVM doesn't read cgroups by default (in older versions). Same for Node.js and its heap. - Performance tuning.
cpu.maxthrottling is enforced by the kernel scheduler. If a container is CPU-throttled, you'll see it incpu.statundernr_throttled, and it looks like a sluggish application rather than a hard limit. - Security review. Understanding that containers aren't VMs changes how you think about multi-tenant isolation, and pushes you toward user namespaces or rootless containers when the threat model warrants it.
Docker, Podman, containerd, and Kubernetes all do the same things at the kernel level. Kubernetes is orchestration, not a separate isolation mechanism. Once you've seen namespaces and cgroups, kubectl pods stop being magical.
Reading list
man namespaces, man cgroups, and man unshare are the primary sources, and they're actually readable. The clone(2) man page lists every namespace flag with a one-line description.
If you want to play with this without Docker in the way, unshare and nsenter are the two commands to learn. Both ship with util-linux on every distro. Ten minutes with them and the mental model clicks.
