Skip to main content

Load Balancing & Service Reliability

Comparative Architecture

To understand how a Load Balancer differs from a Reverse Proxy and an API Gateway, and how they coexist in a production network path, see the Reverse Proxy vs. Load Balancer vs. API Gateway Guide.


Load Balancing Algorithms

AlgorithmHowBest For
Round RobinRotate through servers in orderEqual capacity servers, stateless
Weighted Round RobinMore requests to higher-weight serversMixed capacity servers
Least ConnectionsRoute to server with fewest active connectionsVariable request durations
IP HashHash client IP β†’ same serverSession affinity, WebSocket
RandomRandom server selectionSimple, low overhead
Resource-basedRoute based on CPU/memoryCPU-intensive workloads

L4 vs L7 Load Balancing

L4 (Transport)L7 (Application)
Works atTCP/UDP levelHTTP/HTTPS level
Content awarenessNo (binary stream)Yes (URL, headers, cookies)
TLS terminationNoYes
Routing byIP:portURL path, headers, cookies
PerformanceHigherSlightly lower (parses headers)
ExamplesAWS NLB, HAProxy (L4 mode)AWS ALB, nginx, Envoy

Load Balancer Architecture

Load Balancer Architecture Path
Client (User)DNS ResolutionGlobal LB (Anycast)Regional LB (L7)Service Pools & Mesh

Global Load Balancer

Uses Anycast routing to advertise a single IP address globally. Directs packets over private fiber to the closest data center.

  • Layer 4 (TCP/UDP) routing via Border Gateway Protocol (BGP).
  • Absorbs massive distributed DDoS attacks at edge PoPs.
  • Passes traffic to local application balancers.

πŸ’‘ Click any diagram node to inspect details.


Health Checks

Types

TypeDescriptionUse
PassiveTrack errors/timeouts on live trafficReal traffic quality signal
ActivePeriodic synthetic request to health endpointProactive failure detection
HybridBoth passive + activeBest coverage
# nginx health check config
upstream backend {
server app1:8080;
server app2:8080;

check interval=3000 rise=2 fall=3 timeout=1000 type=http;
check_http_send "GET /actuator/health HTTP/1.0\r\n\r\n";
check_http_expect_alive http_2xx;
}

Kubernetes Probes

livenessProbe: # Is the container alive? Restart if fails.
httpGet:
path: /actuator/health/liveness
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
failureThreshold: 3

readinessProbe: # Is the container ready to serve traffic? Remove from LB if fails.
httpGet:
path: /actuator/health/readiness
port: 8080
periodSeconds: 5
failureThreshold: 2

startupProbe: # Is the app done starting? Don't liveness-kill a slow start.
httpGet:
path: /actuator/health
port: 8080
failureThreshold: 30
periodSeconds: 10

Failover Strategies

Failover Strategies Explorer
ClientLoad BalancerPrimary ServerACTIVE (100% Load)Standby ServerIDLE (Hot Standby)Heartbeat OK

Active-Passive Strategy

One active primary processes 100% of the traffic. A secondary (standby) is kept synchronized but idle.

Status Log:
🟒 Primary Server running. Standby synchronizing state.
  • RTO (Recovery Time): 1–5 minutes (DNS propagation / VIP sweep).
  • RPO (Data Loss): Depends on replication mode (Async vs Sync).
Multi-Region Active-Active Replication Failover
Region A (US East)Region B (EU West)US ClientEU ClientALB us-eastApp ServersDB us-eastALB eu-westApp ServersDB eu-west

Multi-Region Deployment

Normally, clients connect to their closest region. Databases replicate continuously across regions.

Traffic Status:
🟒 Region-aware routing functioning. Synchronizing cross-region databases.

Graceful Degradation

Design systems to provide reduced functionality rather than complete failure.

@CircuitBreaker(name = "recommendations", fallbackMethod = "defaultRecommendations")
public List<Product> getRecommendations(Long userId) {
return recommendationService.getPersonalized(userId);
}

// Fallback: show popular items instead of personalized
public List<Product> defaultRecommendations(Long userId, Exception ex) {
log.warn("Recommendation service unavailable, using popular items fallback");
return productService.getMostPopular(10);
}

Degradation Levels

Graceful Degradation Levels

Level 0: Full Functionality

All systems green. Normal latency and full database writes/reads.

FEATURE MATRIX STATUS:
Reads (Catalog, User profile)
AvailableDirectly from DB / local cache.
Writes (Orders, Checkouts)
AvailableTransactions executed synchronously.
AI Recommendations
AvailablePersonalized models executed.
ARCHITECTURE / IMPLEMENTATION PATTERN:
// Full latency path
public List<Product> getRecommendations(Long userId) {
    return recommendationService.getPersonalized(userId);
}

Chaos Engineering

Deliberately inject failures to find weaknesses before users do.

Principles

  1. Define steady state (normal behavior)
  2. Hypothesize what will happen during failure
  3. Introduce failure in controlled way
  4. Compare actual vs expected

Common Chaos Experiments

ExperimentWhat to Test
Kill random podService resilience, restart behavior
Introduce network latencyTimeout handling, circuit breakers
Drop packetsRetry logic, idempotency
Exhaust connection poolBackpressure, error handling
Spike CPU to 90%Autoscaling, latency under load
Kill a DB replicaFailover, replication handling

Chaos Monkey (Spring Boot)

// Chaos Monkey for Spring Boot (Netflix Chaos Monkey style)
@ChaosMonkey(
assaults = {LatencyAssault.class},
watcher = {ServiceWatcher.class}
)
@Service
public class OrderService { ... }

Disaster Recovery

RTO & RPO Targets

TierRTORPOStrategy
Tier 1 (Critical)< 1 hour~0 (zero data loss)Active-active, synchronous replication
Tier 2 (Important)< 4 hours< 1 hourActive-passive, async replication
Tier 3 (Standard)< 24 hours< 24 hoursBackup + restore
Tier 4 (Low)< 72 hours< 72 hoursPeriodic backup

DR Runbook Checklist

  • Identify failed component(s)
  • Assess data loss window (check last replication timestamp)
  • Activate DR environment
  • Point DNS to DR
  • Verify functionality with smoke tests
  • Notify stakeholders
  • Document incident timeline
  • Post-mortem within 48 hours

Zero-Downtime Deployments

Zero-Downtime Deployment Strategies
User TrafficRouter / LBBlue Cluster (v1)100% TrafficGreen Cluster (v2)0% Traffic (Staging)

Blue-Green Deployment

Two identical physical environments are maintained. Traffic is switched atomically by updating the load balancer rules.

Operational Log:
πŸ”΅ Running version 1.0.0 in Blue environment. Green is idle and ready for deployment.
# Kubernetes canary with Argo Rollouts
apiVersion: argoproj.io/v1alpha1
kind: Rollout
spec:
strategy:
canary:
steps:
- setWeight: 5 # 5% to canary
- pause: {duration: 10m}
- setWeight: 25
- pause: {duration: 10m}
- setWeight: 100

Anycast

CDNs use anycast to route users to the nearest PoP automatically.

Anycast Routing Mechanics
US ClientEU ClientShared IP: 104.16.82.100 (Announced globally via BGP)US Edge PoPIP: 104.16.82.100EU Edge PoPIP: 104.16.82.100

Anycast Routing

In Anycast, multiple physically separate servers announce the exact same IP address to the internet using **BGP (Border Gateway Protocol)**.

πŸ‡ΊπŸ‡Έ Routing client in US:

Internet routers identify the US Edge PoP as the shortest topological path. Packets arrive at the US server instantly.

πŸ’‘ Click on either client node to see where their traffic is routed.

Used by: Cloudflare, Google (8.8.8.8), root DNS servers.


Session Persistence (Sticky Sessions)

When application state is stored in-server memory, all requests from a user must go to the same server.

Session Persistence & Affinity Models
Alice (Client)Sticky LBServer AActive (Session: Alice)Server BActive (Session: Empty)

Sticky Sessions

The Load Balancer uses cookies or IP hashing to pin a user's session to a specific backend server.

System Impact:
🟒 Active. LB uses cookie `SERVERID=server-a` to keep Alice pinned to Server A.
Best Practice

Avoid sticky sessions when possible. Store session data in a distributed cache (Redis) so any server can handle any request β€” true horizontal scaling.


nginx Load Balancer Configuration

upstream api_servers {
least_conn; # algorithm

server 10.0.0.1:8080 weight=3;
server 10.0.0.2:8080 weight=2;
server 10.0.0.3:8080 backup; # only used when others are down

keepalive 32; # keep N idle upstream connections open
}

server {
listen 80;
server_name api.example.com;

location /api/ {
proxy_pass http://api_servers;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;

proxy_connect_timeout 5s;
proxy_send_timeout 60s;
proxy_read_timeout 60s;
}

location /static/ {
root /var/www;
expires 30d;
add_header Cache-Control "public, immutable";
}
}

Global Server Load Balancing (GSLB)

Distributes traffic across multiple data centers globally using DNS:

Global Server Load Balancing (GSLB) Modes
US ClientEU ClientGSLB Server(Route 53 / DNS)Data Center AUS East (203.0.113.10)Data Center BEU West (198.51.100.20)

Latency-Based Routing

GSLB resolves the DNS query to the IP of the data center topologically nearest to the client, minimizing round-trip times.

DNS Resolution Log:
🌐 US user resolved to 203.0.113.10 (DC A).
🌐 EU user resolved to 198.51.100.20 (DC B).

Interview Questions

Q: What load balancing algorithm would you use for WebSocket connections? Why?

A: Use least-connections (or weighted least-connections) because WebSockets are long-lived and uneven. It balances concurrent connection load better than round-robin.

Q: What is the difference between L4 and L7 load balancing?

A: L4 routes by IP/port and is fast/protocol-agnostic; L7 routes by HTTP attributes like path, host, or headers. L7 enables smarter routing and policy but adds more processing overhead.

Q: What is the difference between liveness and readiness probes in Kubernetes?

A: Liveness checks if container should be restarted; readiness checks if it can serve traffic now. Readiness protects rollout and dependency warm-up phases.

Q: What is blue-green deployment and how does it enable zero-downtime releases?

A: Run old and new stacks in parallel, then switch traffic atomically to the new stack. Rollback is fast by flipping traffic back to the old environment.

Q: What is canary deployment? How do you decide when to proceed vs rollback?

A: Canary sends a small traffic slice to new version first and expands gradually. Advance only if error/latency/business KPIs stay within guardrails; otherwise auto-rollback.

Q: What is chaos engineering and why is it important?

A: Chaos engineering injects controlled faults to validate resilience assumptions in production-like conditions. It exposes hidden coupling before real incidents do.

Q: What is RPO and RTO? How do you design a system to meet given targets?

A: RPO is acceptable data loss window; RTO is acceptable recovery time. Meet targets with replication frequency, backup strategy, failover automation, and regular disaster drills.

Q: How do you implement graceful degradation in a microservices system?

A: Prioritize core paths and shed optional features during overload/failure. Use timeouts, fallbacks, cached defaults, and feature flags to keep essential functionality alive.

Q: What is the difference between active-passive and active-active failover?

A: Active-passive keeps standby idle until failover, simplifying consistency but increasing failover delay. Active-active serves traffic in multiple sites continuously, improving availability but complicating conflict handling.

Q: How do you design a multi-region system that remains consistent during a regional outage?

A: Classify data by consistency need: strongly consistent writes via quorum/primary strategy, and eventually consistent data via async replication. Automate failover with clear write-routing and reconciliation plans.

CDN & Network Level Load Balancing Questions

Q1. What is the difference between L4 and L7 load balancing?

L4 (transport layer) load balancing operates on TCP/UDP β€” it sees source/destination IP and port but not application data. Fast, protocol-agnostic, but can't route by URL path or headers. L7 (application layer) operates on HTTP β€” it can route by URL, headers, cookies, method; terminate TLS; perform content-based routing; modify requests. More overhead but much more flexible.

Q2. What is the difference between round-robin and least-connections algorithms?

Round-robin cycles requests evenly across servers β€” good for stateless apps where each request takes similar time. Least-connections sends to the server with fewest active connections β€” better when request duration varies significantly (e.g., some requests take 100ms, others take 10s). Least-connections prevents piling requests onto a slow server that's busy with long operations.

Q3. What are sticky sessions and why are they problematic at scale?

Sticky sessions (session affinity) ensure all requests from one user go to the same backend server β€” needed when session state is in server memory. Problems: uneven load distribution, a server crash loses all its users' sessions, prevents true horizontal scaling. Solution: externalize session state to Redis (or a database), making all servers stateless and interchangeable.

Q4. How does anycast work and where is it used?

Anycast announces the same IP address from multiple geographic locations. BGP routing naturally sends packets to the topologically nearest location announcing that prefix. There's no DNS trick β€” the internet's routing infrastructure handles it. Used by: CDNs (serve from nearest PoP), public DNS resolvers (8.8.8.8 works from anywhere), root DNS servers, DDoS mitigation (absorb attacks at multiple locations).

Q5. What is the purpose of health checks in load balancing?

Health checks detect unhealthy backends before they cause user-facing errors. Active checks proactively send probes (HTTP GET to /health) and mark backends down before failures impact traffic. Passive checks detect failures from live traffic patterns. Without health checks, the LB would send requests to dead servers, causing errors for those users until the problem is noticed manually.

Q6. What is cache busting and why is it needed with CDNs?

When static assets (JS, CSS) are updated but clients have cached the old version, users see outdated code. Cache busting adds a content hash to filenames (app.a3f4b.js) so each new version has a unique URL β€” never cached as the old version. CDN serves app.a3f4b.js with very long TTL (immutable); after deployment, the HTML references app.c7d2a.js (new hash) β€” a fresh request.

Q7. How would you design a load balancer health check for a Spring Boot application?

Expose Spring Boot Actuator's /actuator/health endpoint. Configure it to check database connectivity, cache availability, and any critical dependencies. Return HTTP 200 when healthy, 503 when not. Configure the LB to send GET /actuator/health every 10–30s, mark unhealthy after 2–3 failures, mark healthy after 2 successes. Protect the endpoint β€” restrict access to LB's IP range or use a separate management port.

Q8. Explain CDN cache invalidation strategies.

Options: (1) TTL expiry β€” wait for TTL to expire (simplest; acceptable for slow-changing content); (2) Versioned URLs β€” change the URL on update (hash in filename); cache old versions expire naturally; (3) API purge β€” call CDN's purge API to immediately invalidate specific URLs or patterns (Cloudflare, Fastly, CloudFront all support this); (4) Cache tags/surrogate keys β€” tag cached objects and purge all objects with a tag (e.g., purge all objects tagged product:42).

πŸ“–
Track Page Progress0 / 635 Read
Knowledge Base Completion0%