Skip to main content

Docker Fundamentals

What is a Container? (Demystified)

The Core Mental Model: A container is not a mini-virtual machine. There is no guest operating system kernel or hypervisor. A container is simply a standard Linux process isolated by the host kernel.

When you run docker run -d -p 80:80 nginx, the Linux host starts a regular process called nginx. However, the Docker daemon wraps that process inside three Linux kernel isolation primitives:

  1. Linux Namespaces: Controls what the process can SEE (its own private process tree, network interfaces, and filesystem).
  2. Control Groups (cgroups): Controls what the process can CONSUME (maximum CPU percentage, memory limits, and I/O rates).
  3. OverlayFS Union Filesystem: Layers read-only image layers under a thin read-write scratch layer.
Docker Under the Hood: Kernel Internals & Architecture
Demystifying Containers: "A Container is Just an Isolated Linux Process"
Containers are not virtual machines. There is no hypervisor and no guest OS. A container is a standard Linux process isolated by the kernel using three foundational primitives: Namespaces (what it can see), Cgroups (what it can consume), and OverlayFS (how its filesystem is layered).
๐Ÿง Shared Host Linux Kernel (Rings 0 & 3)Single kernel manages all containers๐Ÿ“ฆ Container Process (e.g. Node.js)๐Ÿ›ก๏ธ Linux NamespacesPID, NET, MNT, IPCโš–๏ธ Control GroupsMax 512MB RAM, 1 CPU๐Ÿ“ OverlayFS Union FilesystemContainer Layer (Read/Write - UpperDir)Image Layers (Read-Only - LowerDir / Base OS)
PID Namespace โ€” Process Isolation
The container process sees itself as PID 1, while on the host OS it has a normal PID (e.g. 14230). Cannot see host processes.

The 3 Foundations of Linux Container Isolation

1. Linux Namespaces (Visibility & Scoping)

Namespaces provide process-level virtualization by creating independent partitions for system resources:

NamespaceLinux FlagWhat It IsolatesContainer Behavior
PIDCLONE_NEWPIDProcess IDsThe container process sees itself as PID 1. It cannot see any other process running on the host or other containers.
NETCLONE_NEWNETNetwork stackContainer gets its own private lo loopback (127.0.0.1), IP routing table, and virtual ethernet pair (veth) attached to docker0 bridge.
MNTCLONE_NEWNSFilesystem mount pointsRoots the container into its private filesystem, hiding /home, /etc, and /var of the host.
IPCCLONE_NEWIPCInter-process communicationPrevents container processes from accessing shared memory segments, semaphores, or message queues of the host.
UTSCLONE_NEWUTSHostname and domainAllows setting a container-specific hostname (--hostname web-01) without modifying the host machine's name.
USERCLONE_NEWUSERUser and group IDsMaps container root (UID 0) to an unprivileged UID (e.g. UID 10001) on the host, preventing host root escalation.

2. Control Groups (cgroups) (Resource Guardrails)

While namespaces prevent a container from snooping on the host, cgroups prevent a "noisy neighbor" container from crashing the host:

  • docker run -m 512m --cpus="1.5" creates a cgroup directory in /sys/fs/cgroup/memory/docker/<container_id>.
  • If memory usage exceeds 512MB, the Linux kernel's Out-Of-Memory (OOM) killer terminates that container process without affecting the host or other containers.

3. OverlayFS (Layered Union Mount & Copy-on-Write)

Docker images are built as immutable, stacked layers using a union filesystem:

  • LowerDir (Read-Only): The immutable base OS (e.g., Alpine/Debian) and installed runtime packages.
  • UpperDir (Read/Write): A thin ephemeral layer created when the container starts. Any new files or edits are written here (Copy-on-Write).
  • MergedDir: The unified filesystem view that the container process actually sees.
OverlayFS Architecture:
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ MergedDir: Unified view presented to container process โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ UpperDir: Ephemeral Read-Write layer (/tmp, modified) โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ LowerDir Layer 3: COPY app.jar (Read-Only) โ”‚
โ”‚ LowerDir Layer 2: RUN apt-get install (Read-Only) โ”‚
โ”‚ LowerDir Layer 1: FROM debian:bookworm (Read-Only) โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

The Container Runtime Hierarchy (Under the Hood)

When you execute docker run, Docker delegates execution down a modular stack governed by Open Container Initiative (OCI) standards:

User CLI: "docker run"
โ”‚
โ–ผ REST API / Unix Socket (/var/run/docker.sock)
[ Docker Daemon (dockerd) ]
โ”‚ (High-level: manages networking, volumes, CLI parsing)
โ–ผ gRPC
[ containerd ]
โ”‚ (Supervises containers, image pulls, storage snapshots)
โ–ผ
[ containerd-shim ] โ”€โ”€(Surrogate parent process; enables daemonless containers)
โ”‚
โ–ผ CLI invocation
[ runc ] โ”€โ”€(Low-level OCI runtime: invokes clone(), unshare(), pivot_root)
โ”‚
โ–ผ (runc exits after spawning)
[ Container Process (e.g. nginx PID 1) ]
  • Why containerd-shim exists: The shim acts as a lightweight daemonless babysitter process. It holds stdout/stderr pipes open and waits on the container's exit code. This allows dockerd and containerd to crash or be upgraded without terminating running containers!

Containers vs Virtual Machines

Virtual Machines vs Docker Containers vs Kubernetes Orchestration
STACK LAYERS (TOP TO BOTTOM):
Container 1 / Container 2 / Container 3 (App Bin/Libs)
Container Runtime Engine (dockerd / containerd / runc)
Host Linux OS Kernel (Shared cgroups, Namespaces, Syscalls)
Physical Bare-Metal Hardware (CPU, RAM, NIC, SSD)
Docker Containers: Lightweight Kernel Process Isolation

Containers are not virtual machines! They are regular Linux processes isolated by kernel Namespaces, cgroups, and capabilities.

  • Boot Time: Sub-second (instant process fork).
  • Isolation: Kernel-level sandboxing (shared kernel).
  • Overhead: Near-zero โ€” only application memory used.

Image Naming and Tags

docker.io / library / ubuntu : 24.04
โ†‘ โ†‘ โ†‘ โ†‘
Registry Namespace Image Tag/version

# Examples:
ubuntu # docker.io/library/ubuntu:latest
nginx:1.25 # docker.io/library/nginx:1.25
mycompany/myapp:1.0.0 # docker.io/mycompany/myapp:1.0.0
123456789.dkr.ecr.us-east-1.amazonaws.com/myapp:v2 # AWS ECR

Tag Best Practices

TagUseRisk
latestDevelopment onlyUnpredictable โ€” changes silently
1.0.0 (semver)Production โœ…Immutable reference
sha256:abc123...Pinned exact version โœ…Most explicit, never changes
# Always tag with version + latest for production images
docker build -t myapp:1.2.3 -t myapp:latest .

# Pull by digest (guaranteed immutable)
docker pull ubuntu@sha256:45b23dee08af5e43a7fea6c4cf9c25ccf269ee113168c19722f87876677c5cb2

Container Lifecycle

docker create
Image โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ†’ Created
โ”‚
docker start โ”‚
โ†“
Running โ†โ”€โ”€โ”€โ”€ docker restart
โ”‚
docker pause โ†“
Paused
โ”‚
docker unpause โ†“
Running
โ”‚
docker stop (SIGTERM) โ”‚
docker kill (SIGKILL) โ†“
Stopped/Exited
โ”‚
docker rm โ†“
Removed (deleted)

Shortcut: docker run = docker create + docker start

Registries

Public Registries

RegistryURLNotes
Docker Hubhub.docker.comDefault, largest public registry
GitHub Container Registryghcr.ioIntegrated with GitHub Actions
Google Container Registrygcr.ioGoogle Cloud
Amazon ECR Publicpublic.ecr.awsAWS public images

Private Registries

RegistryNotes
Amazon ECRPrivate, per AWS account
Google Artifact RegistryPrivate, replaces GCR
Azure Container RegistryPrivate, Azure
HarborSelf-hosted, open source
Nexus RepositorySelf-hosted, enterprise
# Login to Docker Hub
docker login

# Login to AWS ECR
aws ecr get-login-password --region us-east-1 | \
docker login --username AWS --password-stdin \
123456789.dkr.ecr.us-east-1.amazonaws.com

# Tag for ECR
docker tag myapp:1.0.0 123456789.dkr.ecr.us-east-1.amazonaws.com/myapp:1.0.0

# Push
docker push 123456789.dkr.ecr.us-east-1.amazonaws.com/myapp:1.0.0

Key Concepts Summary

TermDefinition
ImageImmutable, layered snapshot of a filesystem + config. Blueprint.
ContainerRunning (or stopped) instance of an image. Has writable layer.
DockerfileText file with instructions to build an image.
RegistryRemote repository for storing and distributing images.
LayerRead-only filesystem diff. Multiple layers make up an image.
TagHuman-readable label pointing to a specific image version.
DigestContent-addressable SHA256 hash โ€” uniquely identifies an image.
VolumePersistent storage that survives container restarts.
NetworkVirtual network connecting containers.
Docker ComposeTool for defining multi-container apps in YAML.

Interview Questions

  1. What is a container and how does it differ from a virtual machine?
  2. What are Docker image layers and why do they matter for build performance?
  3. What is the difference between an image and a container?
  4. Why is using the latest tag bad practice in production?
  5. What are Linux namespaces and cgroups? How do they relate to containers?
  6. What happens to data in a container's writable layer when the container is removed?
  7. What is a container registry and name three examples.
  8. Explain the Docker client-daemon architecture.
๐Ÿ“–
Track Page Progress0 / 635 Read
Knowledge Base Completion0%