Overview

I used Docker for four years before I understood what it actually does. The mental model I had was "lightweight VM," which is wrong in every direction that matters. A container is a Linux process that's been lied to about what machine it's on, plus a resource limit that the kernel enforces.

That's the whole thing. Two kernel features — namespaces and cgroups — and a filesystem layered on top. Once you see it, a lot of Docker behavior stops being mysterious.

Namespaces: the isolation part

A namespace wraps a global system resource and makes it look private to the processes inside. There are eight of them on a modern Linux system:

NamespaceIsolatesSince
mntFilesystem mount points2.4.19
pidProcess IDs2.6.24
netNetwork interfaces, routes, ports2.6.24
ipcSystem V IPC, POSIX message queues2.6.19
utsHostname and domain name2.6.19
userUser and group IDs3.8
cgroupcgroup root directory4.6
timeBoot and monotonic clocks5.6

You can create a namespace with the unshare command:

# Run bash in a new UTS namespace
sudo unshare --uts /bin/bash

# Inside this shell:
hostname my-container
hostname
# my-container

exit

# Back in your shell:
hostname
# your-actual-hostname

You just changed the hostname of the system from inside a shell, and it didn't affect anything outside. That's the entire mechanism behind a container having its own hostname. It's a UTS namespace.

The process and network namespaces are the interesting ones:

sudo unshare --pid --fork --mount-proc /bin/bash

# Inside:
ps aux
# PID 1 is bash, plus maybe one or two others
# You cannot see any process outside this namespace

The --mount-proc is needed because /proc is a mount that's shared. Without remounting it, ps would still show all the host's processes even though the PID namespace is new.

cgroups: the limit part

Namespaces control what a process can see. cgroups control what it can use. CPU, memory, disk I/O, number of processes — all enforced by the kernel, not by any userspace process.

cgroup v2 (the current version, unified hierarchy) is organized as a filesystem tree:

/sys/fs/cgroup/
├── system.slice/
│   ├── sshd.service/
│   └── cron.service/
├── user.slice/
└── my-container/
    ├── cgroup.controllers
    ├── memory.max
    ├── cpu.max
    └── pids.max

To limit a process tree to 512MB of memory:

sudo mkdir /sys/fs/cgroup/my-test
echo "+memory +cpu +pids" | sudo tee /sys/fs/cgroup/my-test/cgroup.subtree_control
echo "512M" | sudo tee /sys/fs/cgroup/my-test/memory.max
echo "200000 100000" | sudo tee /sys/fs/cgroup/my-test/cpu.max  # 2 CPUs
echo "100" | sudo tee /sys/fs/cgroup/my-test/pids.max

echo $$ | sudo tee /sys/fs/cgroup/my-test/cgroup.procs
# Any process in this shell is now limited

The cpu.max format is "quota period" in microseconds. 200000/100000 means 2 CPU-seconds per 1 CPU-second, which is two cores' worth.

If a process exceeds memory.max, the kernel kills it. The OOM killer runs, the process dies, and if you're running Docker, the container exits with code 137 — which is 128 + 9 (SIGKILL). That's why that specific exit code comes up when a container runs out of memory.

Putting them together: what Docker actually does

When you run docker run nginx, roughly this happens:

  1. Docker creates a new set of namespaces: pid, net, mnt, uts, ipc. Optionally user and cgroup.
  2. It creates a cgroup directory and writes the memory, CPU, and PID limits you specified.
  3. It sets up a filesystem — either a single directory (overlayfs, or a bind mount), which becomes the container's /.
  4. It calls clone() with flags for each namespace, forks the container's initial process (PID 1 inside the namespace), and moves it into the cgroup.
  5. The process runs. From its perspective, it's alone on a machine with one filesystem, one network interface, and a hostname it set itself.

You can do all of this manually with unshare, ip, and cgroup writes. It's tedious but not complicated, and doing it once makes Docker's behavior obvious.

Reproducing a container by hand

# Create isolated namespaces and a cgroup
sudo mkdir -p /sys/fs/cgroup/mycontainer
echo "+memory +cpu +pids" | sudo tee /sys/fs/cgroup/cgroup.subtree_control 2>/dev/null
echo "256M" | sudo tee /sys/fs/cgroup/mycontainer/memory.max

# Unshare everything at once, run bash
sudo unshare \
  --uts --ipc --pid --net --mount --fork --mount-proc \
  --cgroup /sys/fs/cgroup/mycontainer \
  /bin/bash

# Inside, you're in a fresh namespace
hostname isolated
ip link set lo up
ip link show
# only the loopback interface

The --cgroup flag takes a path and moves the process into that cgroup. Now any process you spawn from this shell is memory-limited to 256MB.

Try running something that allocates memory:

Python3 -c "x = 'a' * (300 * 1024 * 1024)"
# Killed

That's what an OOM-killed container looks like from inside.

What namespaces don't do

IsolationWhat it coversWhat it doesn't
FilesystemMount pointsA process with CAP_SYS_ADMIN can still see host mounts
ProcessVisible PIDsKernel threads are still shared
NetworkInterfaces and routesNo firewalling; eBPF on the host sees everything
UserUID/GID mappingKernel CVEs can escape

The big one: containers are not a security boundary the way VMs are. The kernel is shared. A container escape is a kernel exploit that gets you root on the host. This is why people run gVisor or Firecracker (which AWS Lambda uses) when they need real isolation.

Docker does reduce the attack surface significantly: no CAP_SYS_ADMIN by default, a seccomp filter blocking ~40 syscalls, AppArmor/SELinux profiles. But "container = sandbox" is a simplification that gets people into trouble.

Namespaces you can see right now

Every process on your system is in some namespace. You can see them:

# List namespaces for the current process
ls -l /proc/self/ns/

# Compare with PID 1
ls -l /proc/1/ns/

# See all network namespaces
sudo lsns -t net

# Check which namespace PID X is in
sudo readlink /proc/<pid>/ns/net

Two processes with the same namespace inode number are in the same namespace. This is how you find which container a process belongs to when you're debugging on the host.

docker inspect --format '{{.State.Pid}}' my-container
# 28431
sudo ls -l /proc/28431/ns/

The net namespace inode in the container's process matches what lsns -t net shows for the container. That's how you can jump from a container to the host's view of it, or vice versa.

Why this matters in practice

Four situations where knowing namespaces and cgroups helps:

  • Debugging a container that "can't reach the network." The container has its own net namespace. If you're expecting it to share the host's, you need --network host. Otherwise, it's in a veth pair with an IP that only exists inside the namespace.
  • Understanding memory limits. A Java process in a container with 2GB limit and no JVM -Xmx flag will try to use the host's memory and get OOM-killed. The JVM doesn't read cgroups by default (in older versions). Same for Node.js and its heap.
  • Performance tuning. cpu.max throttling is enforced by the kernel scheduler. If a container is CPU-throttled, you'll see it in cpu.stat under nr_throttled, and it looks like a sluggish application rather than a hard limit.
  • Security review. Understanding that containers aren't VMs changes how you think about multi-tenant isolation, and pushes you toward user namespaces or rootless containers when the threat model warrants it.

Docker, Podman, containerd, and Kubernetes all do the same things at the kernel level. Kubernetes is orchestration, not a separate isolation mechanism. Once you've seen namespaces and cgroups, kubectl pods stop being magical.

Reading list

man namespaces, man cgroups, and man unshare are the primary sources, and they're actually readable. The clone(2) man page lists every namespace flag with a one-line description.

If you want to play with this without Docker in the way, unshare and nsenter are the two commands to learn. Both ship with util-linux on every distro. Ten minutes with them and the mental model clicks.