Skip to main content

Pods & Containers

A Pod is the smallest deployable unit in Kubernetes. It wraps one or more containers that share a network and storage.


What is a Pod?

Kubernetes Architecture, Control Plane & Pod Lifecycle Simulator
CONTROL PLANE NODES (MASTER):
kube-apiserver
etcd Key-Value Store
kube-scheduler
kube-controller-manager
WORKER NODES (COMPUTE):
kubelet Agent
kube-proxy
Container Runtime (containerd)
Subsystem Component Inspection
kube-apiserver
Role: API Gateway & Validation

Central hub. All kubectl calls & internal agents authenticate & validate YAML schemas via apiserver.

Key rules:

  • All containers in a Pod share the same IP address
  • Containers in a Pod can communicate via localhost
  • A Pod is always scheduled to one Node โ€” it doesn't span nodes
  • Pods are ephemeral โ€” they're not moved, they're recreated (new IP each time)

Pod YAML โ€” Complete Example

apiVersion: v1
kind: Pod
metadata:
name: my-api-pod
namespace: default
labels:
app: my-api # Used by Services to find this Pod
version: "1.0"
environment: production
annotations:
description: "Spring Boot API server"
git-commit: "abc123"

spec:
# โ”€โ”€โ”€ Init Containers (run to completion BEFORE app starts) โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
initContainers:
- name: wait-for-db
image: busybox:1.36
command: ['sh', '-c',
'until nc -z postgres 5432; do echo waiting for postgres; sleep 2; done']

- name: run-migrations
image: myapp-migrations:1.0.0
env:
- name: DB_URL
value: "jdbc:postgresql://postgres:5432/mydb"

# โ”€โ”€โ”€ Main Containers โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
containers:
- name: api
image: myapp:1.0.0
imagePullPolicy: Always # Always | IfNotPresent | Never

# Ports (documentation only โ€” doesn't actually open ports)
ports:
- name: http
containerPort: 8080
protocol: TCP
- name: management
containerPort: 9090

# โ”€โ”€โ”€ Environment Variables โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
env:
- name: SPRING_PROFILES_ACTIVE
value: "production"
- name: DB_PASSWORD
valueFrom:
secretKeyRef:
name: db-secret # Secret name
key: password # Key within the secret
- name: APP_NAME
valueFrom:
configMapKeyRef:
name: app-config
key: app-name
- name: MY_POD_NAME
valueFrom:
fieldRef: # Downward API โ€” inject Pod metadata
fieldPath: metadata.name
- name: MY_NODE_NAME
valueFrom:
fieldRef:
fieldPath: spec.nodeName
- name: MY_CPU_LIMIT
valueFrom:
resourceFieldRef:
resource: limits.cpu

# Inject entire ConfigMap or Secret as env vars
envFrom:
- configMapRef:
name: app-config # All keys become env vars
- secretRef:
name: app-secrets

# โ”€โ”€โ”€ Resource Requests & Limits โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
resources:
requests: # Minimum guaranteed resources
memory: "256Mi" # Scheduler uses this for placement
cpu: "250m" # 250 millicores = 0.25 CPU
limits: # Maximum allowed
memory: "512Mi" # OOMKilled if exceeded
cpu: "1000m" # Throttled if exceeded (NOT killed)

# โ”€โ”€โ”€ Health Probes โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
startupProbe: # Is the app done starting?
httpGet:
path: /actuator/health
port: 9090
failureThreshold: 30 # 30 ร— 10s = 5 min max startup
periodSeconds: 10

livenessProbe: # Is the app alive? Restart if not.
httpGet:
path: /actuator/health/liveness
port: 9090
initialDelaySeconds: 0 # startupProbe handles initial delay
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3

readinessProbe: # Is app ready for traffic? Remove from LB if not.
httpGet:
path: /actuator/health/readiness
port: 9090
periodSeconds: 5
timeoutSeconds: 3
failureThreshold: 2
successThreshold: 1

# โ”€โ”€โ”€ Volume Mounts โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
volumeMounts:
- name: config-vol
mountPath: /app/config
readOnly: true
- name: data-vol
mountPath: /app/data
- name: tmp-vol
mountPath: /tmp

# โ”€โ”€โ”€ Security Context โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
securityContext:
runAsNonRoot: true
runAsUser: 1001
runAsGroup: 1001
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]

# โ”€โ”€โ”€ Lifecycle Hooks โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
lifecycle:
postStart: # Runs after container starts (async)
exec:
command: ["/bin/sh", "-c", "echo Container started"]
preStop: # Runs before container stops (sync)
exec: # K8s waits for this before SIGTERM
command: ["/bin/sh", "-c", "sleep 5"] # Drain connections

# โ”€โ”€โ”€ Sidecar Container โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
- name: log-shipper
image: fluent/fluent-bit:2.2
volumeMounts:
- name: log-vol
mountPath: /var/log/app
readOnly: true

# โ”€โ”€โ”€ Volumes โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
volumes:
- name: config-vol
configMap:
name: app-config
- name: data-vol
persistentVolumeClaim:
claimName: my-pvc
- name: tmp-vol
emptyDir: {} # Ephemeral, shared between containers
- name: log-vol
emptyDir: {}

# โ”€โ”€โ”€ Pod-Level Settings โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
restartPolicy: Always # Always | OnFailure | Never
terminationGracePeriodSeconds: 30 # SIGTERM grace period

# Node selection
nodeSelector:
kubernetes.io/arch: amd64
node-type: compute

# Service account for RBAC / cloud provider permissions
serviceAccountName: my-app-sa

# Image pull secrets (for private registries)
imagePullSecrets:
- name: my-registry-secret

# DNS
dnsPolicy: ClusterFirst # ClusterFirst | Default | None

Pod Lifecycle States

Pending โ†’ Pod accepted, but not yet scheduled/pulled
Running โ†’ Pod bound to node, at least one container running
Succeeded โ†’ All containers completed successfully (for Jobs)
Failed โ†’ All containers terminated, at least one failed
Unknown โ†’ Pod state can't be determined (node comms lost)

Container States

Waiting โ†’ Not yet running (pulling image, init containers running)
Running โ†’ Executing
Terminated โ†’ Finished (exit code 0 = success, non-zero = failed)

Health Probes Deep Dive

Three types, each answers a different question:

startupProbe โ†’ "Has the application finished starting?"
(replaces initial delay โ€” handles slow-starting apps)

livenessProbe โ†’ "Is the application still alive?"
(Kubernetes restarts container if this fails)

readinessProbe โ†’ "Is the application ready to serve traffic?"
(Kubernetes removes from Service endpoints if this fails)

Probe Types

# HTTP GET โ€” most common for web services
httpGet:
path: /actuator/health
port: 8080
httpHeaders:
- name: X-Health-Check
value: "true"

# TCP Socket โ€” for non-HTTP services
tcpSocket:
port: 5432 # Just checks if port is open

# Exec Command โ€” run command inside container
exec:
command:
- sh
- -c
- "redis-cli ping | grep PONG"
# Exit code 0 = healthy, non-zero = unhealthy

# gRPC (K8s 1.24+)
grpc:
port: 50051
service: ""

Probe Timing

initialDelaySeconds: 30 # Wait N seconds before first probe
periodSeconds: 10 # Probe every N seconds
timeoutSeconds: 5 # Probe times out after N seconds
failureThreshold: 3 # Mark unhealthy after N consecutive failures
successThreshold: 1 # Mark healthy after N consecutive successes (readiness)

Spring Boot Probe Setup

# application.yml
management:
endpoint:
health:
probes:
enabled: true # Enables /actuator/health/liveness + /readiness
endpoints:
web:
exposure:
include: health
health:
livenessstate:
enabled: true
readinessstate:
enabled: true

Pod Termination Lifecycle & Zero-Downtime Graceful Shutdown

When Kubernetes deletes a Pod (during rolling updates, node drains, or HPA scale-in), two independent workflows execute concurrently:

  1. Network Control Plane (2โ€“5s): The EndpointSlice controller removes the Pod IP. kube-proxy updates iptables/IPVS, and Ingress controllers (Nginx, Envoy, ALB) purge upstream endpoints.
  2. Kubelet & Container Process (under 300ms): Kubelet sends SIGTERM directly to the container process. Web servers without delay close their listening sockets immediately.

If the application stops accepting connections before the network data plane completes endpoint removal, the Ingress router forwards live client requests to a dead socket, generating HTTP 502 Bad Gateway or Connection Reset by Peer.

Kubernetes + Spring Boot 3.x: Zero-Downtime Graceful Shutdown Engine
The 2-Layer architecture introduces a deliberate preStop sleep barrier to synchronize the network and application planes:
0s (Terminating)3s (Endpoints Drained)10s (SIGTERM Sent)25s (Tomcat Drained)45s (SIGKILL limit)Layer 2: preStop Hook [sleep 10s]Endpoints Removed (2-3s)Traffic Drain Buffer (No 502!)Layer 1: Spring GracefulSafety Headroom (20s)
Kubernetes Deployment Spec (`preStop` Hook)
spec:
  # Must be > preStop sleep + Spring timeout
  terminationGracePeriodSeconds: 45
  containers:
    - name: app
      lifecycle:
        preStop:
          exec:
            command: ["/bin/sh", "-c", "sleep 10"]
Spring Boot 3.x (`application.yml`)
server:
  shutdown: graceful # Closes socket, allows in-flight

spring:
  lifecycle:
    timeout-per-shutdown-phase: 30s # Max in-flight wait

The 2-Layer Zero-Downtime Solution

spec:
# Formula: terminationGracePeriodSeconds > preStop sleep + shutdown-timeout + buffer
terminationGracePeriodSeconds: 45
containers:
- name: app
image: my-app:v1.0.0
lifecycle:
preStop:
exec:
# Delay SIGTERM by 10s: allows Ingress & iptables to drain traffic first
command: ["/bin/sh", "-c", "sleep 10"]

Production Best Practices:

  • Exec Form in Dockerfile: Always declare ENTRYPOINT ["java", "-jar", "app.jar"]. Shell form (ENTRYPOINT java -jar app.jar) executes /bin/sh as PID 1, which ignores SIGTERM and causes Kubelet to forcibly kill the application with SIGKILL (kill -9).
  • Grace Period Math: Ensure terminationGracePeriodSeconds exceeds preStop sleep plus the application's graceful draining timeout (timeout-per-shutdown-phase).
  • Deep-dive guide: See Zero-Downtime Graceful Shutdown: Kubernetes & Spring Boot 3.x for full architecture and async worker configuration.

Resource Requests and Limits

resources:
requests: # What the Pod NEEDS (scheduler uses this for placement)
memory: "256Mi"
cpu: "250m" # 250 millicores = 0.25 of 1 CPU core
limits: # Maximum the Pod can USE
memory: "512Mi"
cpu: "1" # 1 full CPU core

What Happens at the Limit?

ResourceAt Limit
MemoryContainer is OOMKilled (Out of Memory). Restarted by K8s.
CPUContainer is throttled (slowed down). NOT killed.

CPU Units

1 CPU = 1000m (millicores)
0.5 CPU = 500m
0.25 CPU = 250m
2 CPUs = 2000m or just "2"

Memory Units

128Mi = 128 Mebibytes (binary, 1024-based) โ€” preferred in K8s
128M = 128 Megabytes (decimal, 1000-based)
1Gi = 1 Gibibyte

QoS Classes

ClassConditionEviction Priority
Guaranteedrequests == limits for all containersLast to be evicted
Burstablerequests < limits (at least one resource)Middle priority
BestEffortNo requests or limits setFirst to be evicted

Init Containers

Run sequentially to completion before the main containers start. If an init container fails, K8s retries until it succeeds.

initContainers:
# 1. Wait for database to be ready
- name: wait-for-postgres
image: busybox:1.36
command:
- sh
- -c
- |
until nc -z -w3 postgres 5432; do
echo "Waiting for postgres..."
sleep 3
done
echo "PostgreSQL is ready!"

# 2. Run database migrations
- name: db-migrate
image: myapp:1.0.0
command: ["java", "-jar", "/app/app.jar", "--spring.batch.job.enabled=migrate"]
env:
- name: SPRING_DATASOURCE_URL
value: jdbc:postgresql://postgres:5432/mydb

# Main container only starts after BOTH init containers succeed
containers:
- name: api
image: myapp:1.0.0

Multi-Container Pod Patterns

Sidecar Pattern

Enhance the main container without modifying it.

containers:
- name: api
image: myapp:1.0.0
volumeMounts:
- name: log-vol
mountPath: /app/logs

- name: log-shipper # Sidecar: ships logs to ELK
image: fluent/fluent-bit:2.2
volumeMounts:
- name: log-vol
mountPath: /var/log/app
readOnly: true

volumes:
- name: log-vol
emptyDir: {} # Shared between both containers

Ambassador Pattern

Proxy outbound connections for the main container.

containers:
- name: api
image: myapp:1.0.0
# API connects to localhost:5432 (ambassador)
env:
- name: DB_HOST
value: localhost # Connects to ambassador, not DB directly

- name: db-ambassador # Ambassador: handles auth, TLS, retries
image: envoy:latest
# Proxies connections from localhost:5432 โ†’ real DB

Adapter Pattern

Adapt the main container's output format for external systems.

containers:
- name: legacy-app
image: legacy-app:1.0.0 # Outputs proprietary log format

- name: log-adapter # Adapter: converts to standard format
image: log-converter:1.0.0
# Reads legacy format, outputs JSON for Elasticsearch

emptyDir โ€” Shared Volume Between Containers

volumes:
- name: shared-data
emptyDir: {} # Created when Pod starts, deleted when Pod dies

containers:
- name: writer
volumeMounts:
- mountPath: /output
name: shared-data

- name: reader
volumeMounts:
- mountPath: /input
name: shared-data # Same directory โ€” both containers see same files

Interview Questions

  1. What is a Pod in Kubernetes and why is it the basic unit rather than a container?
  2. What is the difference between a liveness probe and a readiness probe?
  3. What is a startup probe and when should you use it?
  4. What is the difference between resource requests and resource limits?
  5. What happens to a container when it exceeds its memory limit? Its CPU limit?
  6. What are init containers and what are they used for?
  7. How do two containers in the same Pod communicate?
  8. What is the Sidecar pattern? Give an example.
  9. What are the three Pod QoS classes and how are they determined?
  10. Why are Pods considered ephemeral? What replaces a dead Pod?
๐Ÿ“–
Track Page Progress0 / 635 Read
Knowledge Base Completion0%