Interview Setup
Interview Prompt
Design a load balancer that distributes incoming traffic across a pool of backend servers. Support health checking, graceful connection draining during deploys, session affinity where needed, and both Layer 4 (TCP) and Layer 7 (HTTP) routing modes.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Layer 4, Layer 7, or a hybrid architecture of both? | Layer 4 operates at the transport layer for raw throughput and protocol neutrality such as gaming and database proxies, whereas Layer 7 inspects application headers for path routing, TLS termination, and HTTP/2 multiplexing. |
| Are the backend servers stateful or stateless? | Stateful workloads require sticky sessions or consistent hashing, whereas stateless backends can leverage any balancing algorithm such as round robin or least connections. |
| What is the anticipated connection rate and network bandwidth? | Processing 1M new connections per second is a high stress case that may require kernel bypass technologies such as DPDK or optimized XDP paths rather than standard software load balancing, while aggregate bandwidth dictates Network Interface Card (NIC) sizing. |
| Is the system deployed within a single region or globally across multiple regions? | Multi-region deployments require Global Server Load Balancing (GSLB) using GeoDNS or Anycast BGP, whereas single-region topologies rely on local high-availability pairs. |
Scope
In scope
- Layer 4 vs Layer 7 load balancing architectures
- Consistent hashing and Maglev hashing for stable backend affinity
- Active and passive health check mechanisms
- Connection draining for graceful deployments
- Sticky sessions using cookie injection and IP hashing
- Global Server Load Balancing (GSLB) with Anycast and DNS
Out of scope (state explicitly)
- Building custom hardware load balancing Application-Specific Integrated Circuits (ASICs)
- Implementing a full service mesh control plane, although sidecar data planes are evaluated as an evolutionary stage
- Edge Distributed Denial of Service (DDoS) scrubbing centers, which belong to the WAF and CDN layer
Functional Requirements
In the system design interview, establish whether the architecture requires Layer 4 transport forwarding, Layer 7 application routing, or a layered hybrid of both. Health checking, dynamic autoscaling hooks, SSL termination, and session affinity are foundational discussion areas.
- Distribute traffic: Route incoming client requests efficiently across a pool of healthy backend servers.
- Health checks: Detect unhealthy or degraded backend instances and immediately halt traffic distribution to them.
- Multiple algorithms: Support round robin, weighted round robin, least connections, least response time, random selection, power of two choices, IP hashing, consistent hashing, and Maglev.
- Session affinity (sticky sessions): Pin requests originating from the same client session to the identical backend server when required.
- SSL and TLS termination: Decrypt HTTPS connections at the load balancer to offload cryptographic overhead from backend instances.
- Autoscaling integration: Dynamically register newly provisioned backends and deregister decommissioned nodes.
- Rate limiting: Protect downstream infrastructure by enforcing request rate ceilings per client IP or session.
- Content-based request routing: Route traffic based on HTTP URL path prefixes, headers, hostnames, or query parameters.
- Connection draining: Gracefully complete in-flight transactions before retiring backend servers from the pool.
Non-Functional Requirements
Fair workload distribution, minimal proxy latency, and high availability form the core criteria for evaluating this infrastructure component. The load balancer is the logical ingress tier for each region, so regional load balancer failure directly affects the availability of traffic entering that region.
- High Availability: 99.999% uptime because the load balancer serves as the primary ingress gateway, where any downtime causes a total system outage.
- Ultra-Low Latency: Add less than 1 millisecond of network routing overhead per request.
- High Throughput: Support millions of concurrent connections and over 100K requests per second at peak.
- Scalability: Horizontally scale the load balancer tier without introducing bottlenecks as backend capacity expands.
- Fault Tolerance: Prevent single points of failure through active-passive or active-active redundant clustering.
- Transparency: Maintain a stable service endpoint such as a DNS name or Virtual IP so clients do not need to know the backend topology.
Capacity Estimations
Evaluate traffic volume, connection concurrency, and backend instance counts before selecting a load balancer topology. The baseline design targets 100K requests per second and 500K concurrent connections, while the separate 1M requests per second and 10M connection figures are stress cases used to size the load balancer tier and NAT connection table.
| Metric | Calculation | Value |
|---|---|---|
| Peak client requests / sec | Baseline design target from requirements | 100K |
| Load balancer stress target / sec | Dedicated stress-sizing target | 1M |
| Concurrent connections | Baseline design target | 500K |
| Concurrent connections stress case | Dedicated NAT connection-table sizing case | 10M |
| Backend servers | Regional pool range | 50 to 500 |
| Backend servers stress fleet | Capacity scaling scenario | 1,000 |
| Bandwidth | Given network capacity target | 100 Gbps |
| Health check interval | Given | 5 seconds |
| Health checks / sec | 1,000 servers ÷ 5s interval | 200 |
Architecture Diagram
In the interview, clarify Layer 4 vs Layer 7 scope early because it dictates whether path routing and TLS termination occur at the load balancer.
Walk your interviewer through the request lifecycle across the architecture. Clients connect to a redundant load balancer cluster that actively probes backend health and balances incoming flows using configured routing algorithms. In Layer 4 mode, the load balancer forwards TCP packets by IP and port with minimal latency overhead. In Layer 7 mode, it terminates HTTP and TLS, inspects headers and URL paths, and maintains session affinity through persistent cookies. During rolling deployments, connection draining prevents new requests from reaching retiring instances while in-flight operations complete cleanly.
Component Deep Dives
The following sections analyze Layer 4 vs Layer 7 forwarding mechanics, core load balancing algorithms, active and passive health checking, Virtual Router Redundancy Protocol (VRRP) failover, and graceful connection draining.
Layer 4 (Transport Layer: TCP/UDP)
Layer 4 forwarding delivers raw packet throughput by routing on IP address and port coordinates. In NAT mode it performs address translation without inspecting application payloads, while DSR forwards packets without rewriting the Layer 3 IP headers.
- Forwarding mechanism: Evaluates routing decisions strictly using Layer 3 IP addresses and Layer 4 TCP or UDP port numbers.
- Protocol transparency: Does not inspect application payloads, HTTP headers, or request URLs.
- Packet manipulation: Implements Network Address Translation (NAT) or Direct Server Return (DSR).
- Processing efficiency: Kernel-level and hardware-accelerated packet forwarding provides sub-millisecond routing speeds.
- Target use cases: High-volume TCP streams, database connection pooling, and latency-sensitive gaming backends.
- Common implementations: AWS Network Load Balancer (NLB), F5 BIG-IP appliances, and Linux Virtual Server (IPVS).
In Layer 4 NAT mode, the load balancer receives incoming client flows, maintains flow state, performs destination and source address translation, and reverses the translation on returning responses. It does not terminate the application session in the Layer 4 sense.
Client -> LB (VIP: 10.0.0.1:80) -> LB DNATs destination to backend (10.0.1.5:8080) LB SNATs the client source to its internal address so backend responses return through the LB Backend -> LB -> LB reverses NAT -> Client Client perspective: Communicates exclusively with 10.0.0.1 while backend IPs remain internal
Direct Server Return (DSR) Optimization:
Client -> LB -> LB forwards packet to backend (rewrites Layer 2 MAC address without changing IP headers)
Backend -> Responds DIRECTLY to client (bypasses load balancer completely on the return path)
Throughput Advantage: Load balancer avoids return-path processing, reducing LB bandwidth load substantially.
In the illustrative 1 KB request vs 100 KB response mix, the LB handles roughly 100x less
response traffic than NAT mode because outbound responses are much larger than requestsLayer 7 (Application Layer: HTTP/HTTPS)
Layer 7 routing inspects HTTP headers, URL paths, and cookies to enable content-based routing, TLS termination, and granular session affinity.
- Inspection depth: Decrypts and parses HTTP request headers, query strings, cookies, and message bodies to make intelligent routing decisions.
- Advanced traffic policies: Supports URL path routing, header-based canary releases, cookie-based session persistence, response compression, and caching.
- Performance profile: Incurs higher processing overhead than Layer 4 due to full TCP termination and HTTP parsing, but provides rich routing intelligence.
- Target use cases: Web applications, microservice API gateways, and multi-tenant routing platforms.
- Common implementations: AWS Application Load Balancer (ALB), NGINX, HAProxy, and Envoy Proxy.
Content-Based Path Routing:
Rule 1: /api/v1/users/* -> User Service (servers 1 through 5) Rule 2: /api/v1/orders/* -> Order Service (servers 6 through 10) Rule 3: /static/* -> CDN or Static Asset Server Rule 4: /admin/* -> Admin Service (servers 11 through 12) Default: -> Frontend Web Service
Load Balancing Algorithms: Deep Dive
The chosen algorithm governs both resource utilization fairness and cache locality across the backend fleet.
1. Round Robin
Dispatches each consecutive request to the next server in sequence (Server 1, then Server 2, then Server 3, before cycling back to Server 1).
- Advantages: Simple to implement with deterministic, perfectly equal request distribution across identical hardware.
- Limitations: Disregards differing backend capacities and real-time connection loads, which can lead to overloaded instances.
2. Weighted Round Robin
Assigns proportional request quotas based on server capacity ratings. An instance with weight 3 receives three times the request volume of an instance with weight 1.
- Target use cases: Heterogeneous server environments combining differing CPU core counts or memory configurations.
Weights: S1 = 3, S2 = 1, S3 = 2 Routing Sequence: S1, S1, S1, S2, S3, S3, S1, S1, S1, S2, S3, S3, ...
3. Least Connections (Recommended for Long-Lived Connections)
Steers incoming requests to the backend server with the smallest count of active open connections.
- Architectural value: Mitigates hotspots when transaction processing times vary significantly across requests.
- Weighted formula: Computes effective load as
effective_load = active_connections / weight, routing to the lowest resulting value.
Server 1: 150 active connections, weight 3 -> effective load = 50 Server 2: 80 active connections, weight 1 -> effective load = 80 Server 3: 90 active connections, weight 2 -> effective load = 45 (routed here because effective load is lowest)
4. IP Hash and Consistent Hashing
Calculates a deterministic hash from the client IP address to select the target server.
- Session affinity: Requests originating from the same client always land on the same backend, providing session persistence without cookie inspection.
- Consistent hashing optimization: Adding or removing servers remaps only a fraction of keys, preserving backend cache locality.
- Target use cases: Caching layers and stateful services where memory locality directly influences latency.
5. Least Response Time
Monitors moving averages of backend response latency and active connection counts, directing traffic to the fastest responding node.
- Performance tracking: Tracks exponential moving average round-trip times across backend health and request streams.
- Target use cases: Geographically distributed or heterogeneous compute clusters where instance performance fluctuates.
6. Random and Power of Two Random Choices
Picks backend instances at random, producing a statistically uniform distribution at high traffic volumes.
- Power of Two Random Choices: Picks two servers at random and dispatches the request to the one with fewer active connections.
- Algorithmic efficiency: Reduces maximum load variance compared with naive random selection without requiring a global lock or global queue.
7. Consistent Hashing with Bounded Loads and Maglev
Enhances consistent hashing by capping the maximum requests permitted per server at (1 + ε) * average_load.
- Spillover handling: If the primary hashed backend exceeds its capacity ceiling, the load balancer sequentially routes to the next server along the hash ring.
- Production adoption: Maglev was developed by Google and is available in Envoy and other high-performance load balancing systems. Bounded-load spillover is an additional capacity policy rather than a property of basic table lookup.
Health Checks
Health checks isolate failing backend instances before client requests hit dead nodes, combining active probes with passive anomaly detection.
Active Health Checking:
Evaluation Interval: Every 5 seconds
Probing Workflow:
For each backend server in pool:
Dispatch HTTP GET /health
If response 200 OK arrives within 2 seconds: Mark probe successful
If 3 consecutive probes fail: Mark server UNHEALTHY and stop routing traffic
Detection completes within at most 15 seconds under the stated 5-second probe interval
If 2 consecutive probes succeed while unhealthy: Mark server HEALTHY and resume traffic distributionPassive Health Checking (Outlier Detection):
- Traffic analysis: Observes real-time client transactions without injecting synthetic probe traffic.
- Protocol-aware probes: For non-HTTP Layer 4 services, use TCP connect checks or protocol-specific probes instead of assuming an HTTP health endpoint.
- Outlier threshold: If an instance returns 5xx error codes on more than 50% of sampled requests over a 30-second window after a sufficient sample count, the circuit breaker can mark the node unhealthy.
- Hybrid strategy: Combines active synthetic probes for baseline availability verification with passive outlier detection for immediate incident containment.
Overload Protection, Retries, and SYN Flood Defense
Load balancers need explicit backpressure so a traffic spike does not turn proxy saturation into cascading backend failure.
- Connection and queue limits: Enforce per backend and per client limits on concurrent connections, requests, and queued work. Shed load with bounded queues and fast failures instead of allowing unbounded memory growth.
- Retry policy: Retry only safe or explicitly idempotent requests when a failure occurred before the backend could complete the operation. Use request identifiers so application-level deduplication can protect mutations that are intentionally retryable.
- Circuit breaking: Temporarily stop sending traffic to overloaded or erroring backends and use slow start when capacity recovers.
- SYN flood mitigation: Use SYN cookies or SYN proxying at the Layer 4 edge, connection rate limits, and upstream filtering when half-open connection pressure exceeds safe thresholds. Volumetric DDoS scrubbing remains an edge or WAF concern and is out of scope here.
- Admission control: Prefer rejecting excess traffic with bounded error responses rather than allowing the load balancer or backend connection pools to become memory-bound.
- Rate limiting: Apply token-bucket or leaky-bucket limits per client, API key, tenant, or connection class. Keep local enforcement on the fast path and use coordinated shared state only when a global quota is required.
High Availability via VRRP (Virtual Router Redundancy Protocol)
Virtual Router Redundancy Protocol (VRRP) maintains an active-passive load balancer pair sharing a common Virtual IP, enabling fast failover without cold-start delays. Existing NAT connections can still reset unless state is synchronized.
- Primary node (Active): Owns the Virtual IP (such as
10.0.0.1), handles production traffic, and broadcasts VRRP heartbeats to the standby node every 1 second. - Standby node (Passive): Monitors incoming heartbeats. If 3 consecutive heartbeats are missed in under 3 seconds, the standby promotes itself to active.
- Gratuitous ARP failover: The newly promoted active node broadcasts gratuitous ARP frames across the local network switch to bind the Virtual IP to its MAC address.
- Sub-second acceleration: For mission-critical environments, Bidirectional Forwarding Detection (BFD) pairs with VRRP to reduce detection windows below 500 milliseconds.
LB 1 (Active Node) <--- VRRP Heartbeat Exchange ---> LB 2 (Passive Standby)
| |
Virtual IP: 10.0.0.1 Virtual IP: Standby State
Failover Behavior:
LB 2 detects 3 consecutively missed heartbeats (within about 3 seconds)
LB 2 claims Virtual IP 10.0.0.1 by broadcasting gratuitous ARP
New connections continue using the same Virtual IP. Established NAT flows may reset unless connection state is synchronizedImplementation: Managed by the Keepalived daemon running across both load balancer nodes.
Active-Active Alternative:
- DNS round robin: Publishes multiple load balancer IP addresses, allowing traffic to distribute across both active nodes simultaneously.
- Failover latency constraint: If one node fails, DNS caching delays client traffic redirection until TTL expiry.
- Anycast routing: Preferred approach where both load balancers announce the identical IP address via BGP, achieving automated network route convergence.
Connection Draining (Graceful Shutdown)
Connection draining allows backend instances to finish processing in-flight transactions before retirement, preventing dropped requests during software deployments.
- Stop assigning new incoming connections or requests to the target backend instance.
- Allow existing active TCP streams and in-flight HTTP transactions to finish within a configurable drain timeout window, typically 30 to 300 seconds. For HTTP keep-alive connections, stop accepting new requests and signal connection closure or HTTP/2 GOAWAY as appropriate.
- For long-lived WebSocket sessions, send an application-level reconnect signal and allow the client to establish a new session on a healthy backend.
- Once the drain timeout window expires, forcefully terminate any lingering idle or stalled connections using TCP FIN or RST flags.
- Safely deregister the backend instance from the active routing pool and shut down the host process.
API Design
The load balancer control plane exposes administrative interfaces for registering backend targets, setting routing configurations, tuning health checks, and inspecting runtime cluster metrics. Administrative access should use strong authentication, role-based authorization, audit logging, and secret references rather than sending private TLS keys through configuration APIs.
Control Plane API Signatures
// Load Balancer Control Plane and Ingress Interface Signatures
export type BackendStatus = "HEALTHY" | "UNHEALTHY" | "DRAINING";
export type BalancingAlgorithm =
| "ROUND_ROBIN"
| "WEIGHTED_ROUND_ROBIN"
| "LEAST_CONNECTIONS"
| "LEAST_RESPONSE_TIME"
| "RANDOM"
| "POWER_OF_TWO_CHOICES"
| "IP_HASH"
| "CONSISTENT_HASHING"
| "MAGLEV";
export type HealthCheckProtocol = "HTTP" | "TCP" | "UDP";
export interface HealthCheckConfig {
protocol: HealthCheckProtocol;
path?: string;
intervalSec: number;
timeoutSec: number;
unhealthyThreshold: number;
healthyThreshold: number;
}
export interface BackendRegistration {
address: string; // IP:Port, e.g., "10.0.1.5:8080"
weight: number; // Capacity weight (1-100)
maxConnections: number; // Maximum concurrent connections
healthCheck: HealthCheckConfig;
}
export interface BackendRuntimeState {
backendId: string;
address: string;
status: BackendStatus;
activeConnections: number;
weight: number;
requestsTotal: number;
avgResponseTimeMs: number;
}
export interface StickySessionConfig {
enabled: boolean;
cookieName: string;
ttlSeconds: number;
}
export interface LoadBalancerConfig {
configVersion: number;
protocolMode: "L4" | "L7" | "HYBRID";
algorithm: BalancingAlgorithm;
stickySessions: StickySessionConfig;
tls: {
mode: "TERMINATE" | "PASSTHROUGH";
certificateRef?: string;
privateKeySecretRef?: string;
};
routingRules: Array<{
pathPattern: string;
backendPool: string;
}>;
}
export interface LoadBalancerControlPlane {
// Register a new backend target into an active routing pool
registerBackend(poolId: string, backend: BackendRegistration): Promise<BackendRuntimeState>;
// Gracefully drain connections and retire a backend target
deregisterBackend(poolId: string, backendId: string, drainTimeoutSec: number): Promise<void>;
// Retrieve health status and active metrics for all pool targets
listBackends(poolId: string): Promise<BackendRuntimeState[]>;
// Atomically update routing policies, algorithms, and sticky configurations
updateConfiguration(config: LoadBalancerConfig, actorId: string, requestId: string): Promise<void>;
}Backend Management HTTP Endpoints
POST /api/v1/backends
Content-Type: application/json
{
"address": "10.0.1.5:8080",
"weight": 3,
"max_connections": 1000,
"health_check": {
"protocol": "HTTP",
"path": "/health",
"interval_sec": 5,
"timeout_sec": 2,
"unhealthy_threshold": 3,
"healthy_threshold": 2
}
}
DELETE /api/v1/backends/b_10015?drain_timeout=30
GET /api/v1/backends
Response: 200 OK
Content-Type: application/json
{
"backends": [
{
"backend_id": "b_10015",
"address": "10.0.1.5:8080",
"status": "HEALTHY",
"active_connections": 150,
"weight": 3,
"requests_total": 1500000,
"avg_response_time_ms": 45
}
]
}Load Balancer Configuration HTTP Endpoints
PUT /api/v1/config
Content-Type: application/json
{
"config_version": 1842,
"protocol_mode": "L7",
"algorithm": "LEAST_CONNECTIONS",
"sticky_sessions": {
"enabled": true,
"cookie_name": "SERVERID",
"ttl_seconds": 3600
},
"tls": {
"mode": "TERMINATE",
"certificate_ref": "cert_prod_2026_01",
"private_key_secret_ref": "kms://lb/prod/tls-key"
},
"routing_rules": [
{ "path": "/api/*", "backend_pool": "api-servers" },
{ "path": "/static/*", "backend_pool": "static-servers" }
]
}Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
The load balancer maintains state across three concerns: durable control-plane configuration, transient in-memory connection and affinity state, and optional shared metadata used during failover coordination.
Configuration Control Plane and Versioned Rollout
Routing policy, backend membership, health settings, and certificates need a durable source of truth outside the packet data path. The control plane validates a new version, distributes an immutable snapshot to load balancer instances, activates it atomically, and retains the last known good version for rollback.
# Versioned control-plane configuration
store: "etcd"
config_version: 1842
last_known_good_version: 1841
rollout: "atomic snapshot to data-plane instances"
validation: "schema + backend reachability + policy checks before activation"
rollback: "revert to last_known_good_version on failed rollout"Control Plane Security and Secrets
Configuration changes are administrative operations. Authenticate operators and automation with strong identity, authorize changes by role, audit every version change, and keep private TLS material in a KMS or HSM rather than in load balancer configuration payloads. Certificate rotation uses versioned secret references and an atomic data-plane reload so private keys do not travel through the configuration API.
Connection Tracking Table (Layer 4 NAT Mode)
Session Affinity Store (Layer 7 Cookie and Redis Layout)
# Cookie-based session persistence (L7 HTTP Header)
cookie_header: "Set-Cookie: SERVERID=b_10015; Path=/; Max-Age=3600; HttpOnly; Secure"
# Optional shared affinity metadata for coordinated failover and observability
redis_cluster:
key: "sticky:client_session_hash"
value: "10.0.1.5:8080"
ttl_seconds: 3600
note: "Cookie routing can remain local to the LB, so Redis is not required on every request"Fault Tolerance
Maintaining high availability requires systematic defenses against host failure, backend crashes, connection state loss, thundering herds, and load balancer saturation.
| Concern | Solution |
|---|---|
| Load balancer hardware or host failure | Deploy an active-passive load balancer pair managed by VRRP and Keepalived, enabling automated failover in under 3 seconds. |
| Backend server process failure | Continuous active and passive health checks isolate the faulty backend, remove it from the routing table, and redistribute traffic across healthy nodes. |
| Connection table loss on failover | Under Layer 4 NAT mode, connections can drop during uncoordinated failover and require client reconnects. Direct Server Return removes the load balancer from the return path, reducing egress processing, but the inbound path still depends on the active load balancer. |
| Thundering herd on backend recovery | Enforce slow start mechanisms that gradually ramp traffic to the recovered instance, starting at 10% capacity and progressing to 100% over 30 seconds. |
| Load balancer capacity saturation | Scale horizontally across multiple load balancer Virtual IPs using DNS round robin, Anycast BGP routing, or dedicated hardware appliances for multi-terabit scale. |
Graceful Backend Crash Detection and Recovery
Timeline of Backend Failure Handling: T=0s: Server 3 experiences a hardware crash. T=0s: In-flight requests to Server 3 encounter connection reset, prompting the load balancer to retry idempotent requests on Server 1 or 2. T=5s: Active synthetic health check probe fails (first failed probe recorded). T=10s: Active synthetic health check probe fails (second failed probe recorded). T=15s: Active synthetic health check probe fails (third consecutive failure), causing Server 3 to be marked UNHEALTHY. T=15s: All subsequent incoming client traffic is dispatched exclusively to Server 1 and 2. Passive Anomaly Detection Alternative: T=0s to 2s: Over 50% of real client requests routed to Server 3 return 5xx server errors, causing outlier detection to trip the circuit breaker and mark Server 3 UNHEALTHY immediately.
Additional Considerations
Advanced architectures incorporate Global Server Load Balancing (GSLB), hardware-accelerated TLS termination, and decentralized sidecar proxies to eliminate single points of failure and scale across multiple geographical regions.
Global Server Load Balancing (GSLB)
Global Server Load Balancing steers traffic across geographically distributed data centers to reduce latency and maintain business continuity:
- GeoDNS routing: Resolves client domain queries to a regional endpoint selected from resolver location or client-subnet signals when available, subject to health checks and routing policy.
- Automated failover: If a data center experiences a total failure, health monitoring purges its IP address from DNS records, redirecting subsequent traffic to surviving regions.
- Operational use cases: Disaster recovery, latency optimization, and compliance with regional data sovereignty regulations.
TLS/SSL Termination
Terminating TLS connections at the load balancer provides centralized certificate administration and offloads cryptographic computation from backend fleets.
Client <==== HTTPS (Encrypted) ====> Load Balancer <==== HTTPS / mTLS (Preferred) ====> Backend Servers
|
+==== HTTP only on a trusted private network when policy permits ====> Backend Servers
Architectural Advantages:
1. Cryptographic Offloading: Frees backend compute instances from CPU-intensive TLS handshakes and session key negotiations.
2. Centralized Certificate Management: Certificates and renewal cycles are maintained at the load balancer rather than replicated across hundreds of backend servers.
3. Deep Layer 7 Inspection: Enables the load balancer to inspect decrypted HTTP headers, paths, and cookies for intelligent routing decisions.Modern Load Balancers: Service Mesh
In modern microservices architectures, centralized load balancers can be augmented by decentralized sidecar proxies deployed alongside each container pod. For east-west traffic, sidecars can perform discovery, balancing, retries, and circuit breaking locally:
Service A Instance -> Envoy Sidecar Proxy -> Envoy Sidecar Proxy -> Service B Instance Each microservice container pod runs a dedicated Envoy sidecar handling: 1. Client-side load balancing with latency-aware routing 2. Dynamic service discovery via control plane integration 3. Outlier detection and automated circuit breaking 4. Request retries, timeouts, and deadline propagation 5. Mutual TLS (mTLS) authentication and encryption between services 6. Distributed tracing, metrics export, and access logging
- Production implementations: Istio and Linkerd provide robust service mesh control planes.
- Data plane engine: Envoy runs as an out-of-process sidecar proxy handling Layer 4 and Layer 7 traffic routing.
- Reduces bottlenecks: Moves much of east-west load balancing, retries, and circuit breaking into the service data plane, reducing dependence on centralized internal proxy hops.
Comparison of LB Solutions
Production load balancing solutions balance throughput, Layer 7 feature richness, and operational complexity across distinct deployment models:
| Solution | Type | Layer | Scale | Use Case |
|---|---|---|---|---|
| HAProxy | Software | L4/L7 | Millions of conns | General purpose |
| NGINX | Software | L7 | 100K+ req/s | Web, reverse proxy |
| Envoy | Software | L4/L7 | Service mesh | Microservices |
| AWS ALB | Managed | L7 | Auto-scales | Cloud-native HTTP |
| AWS NLB | Managed | L4 | Millions of conns | TCP/UDP, ultra-low latency |
| F5 BIG-IP | Appliance / Software | L4/L7 | 100 Gbps+ | Enterprise traffic management |
| Maglev | Software (Google) | L3/L4 | Millions of pps | Google's internal LB |
Monitoring
Key telemetry signals required to maintain load balancer health and isolate degraded backend instances include:
- Requests per second per backend: Detects traffic imbalances and uneven algorithm distribution.
- Error rate per backend: Identifies failing servers and isolates 5xx server faults.
- Latency percentiles (p50, p95, p99) per backend: Uncovers slow nodes and tail latency degradation.
- Active connections per backend: Monitors connection pool saturation across upstream targets.
- Load balancer CPU and memory: Ensures the load balancer itself does not become an operating bottleneck.
- Health check success rate: Tracks flapping nodes and infrastructure health trends.
- Connection queue depth: Measures incoming client requests queued while waiting for backend sockets.
25-Minute Interview Cheat Sheet
Begin by establishing the architectural boundary between Layer 4 transport forwarding and Layer 7 application routing. Layer 4 forwarding prioritizes raw packet throughput and minimal latency by routing on IP and port coordinates, whereas Layer 7 inspection enables path-based routing, TLS termination, and cookie-based sticky sessions. When presenting routing algorithms, distinguish between round robin for uniform stateless fleets, least connections for long-lived connections, and consistent hashing or Maglev tables for cache locality and stateful affinity. Emphasize that health checking must pair synthetic active probes with passive 5xx error monitoring so faulty nodes are isolated before traffic black-holes. For multi-datacenter scale, explain how Anycast BGP and GeoDNS guide traffic toward an appropriate healthy regional point of presence while connection draining minimizes avoidable disruption during continuous deployments.
Interview Walkthrough
- 25-minute cut
Skip arch50 and arch75 depth unless the interview targets staff level.
- Requirements clarification and Layer 4 vs Layer 7 scope (3 min)
- Balancing algorithms and stateful session affinity (7 min)
- Active vs passive health check mechanics (5 min)
- TLS termination strategy vs passthrough (5 min)
- Connection draining and high availability failover (5 min)
- Clarify Layer 4 vs Layer 7 scope early because Layer 4 forwards packets by IP and port with minimal latency, whereas Layer 7 parses HTTP headers, URLs, and cookies to support path routing.
- Contrast core algorithms by explaining that round robin assumes homogeneous servers, least connections prevents overloading on long-lived connections, and consistent hashing preserves cache affinity.
- Detail health checks by combining active periodic probes with passive error detection, noting how grace periods during connection draining prevent dropping in-flight requests.
- Address TLS termination at the load balancer vs passthrough, noting that edge termination centralizes certificates and enables HTTP/2 multiplexing at the cost of additional cryptographic CPU load.
- Evaluate sticky session trade-offs between Layer 7 cookie injection and Layer 4 IP hashing, highlighting that IP hashing creates uneven hotspots when users sit behind corporate NAT gateways.
- Introduce Direct Server Return (DSR) as an optimization for Layer 4 architectures when egress response payload dwarfs inbound request volume.
- Highlight the pitfall of applying naive round robin across heterogeneous backend instances, which causes underpowered nodes to experience elevated p99 latencies.
- Explain overload protection with bounded queues, connection limits, safe retries, circuit breaking, and basic SYN flood defenses before treating volumetric DDoS scrubbing as an edge concern.
Related Problems and Concepts
Explore how load balancing algorithms, reverse proxies, and session management connect across core concepts and distributed system architectures:
- Load Balancing Algorithms : In-depth analysis of round robin, weighted least connections, and latency-based traffic scheduling.
- Proxy and Reverse Proxy Patterns : Forward proxies, reverse proxies, TLS termination, and edge traffic routing.
- Consistent Hashing : Ring partitioning, virtual nodes, and bounded load algorithms for distributed affinity.
- Design an API Rate Limiter : Protecting load balancers and backend services from traffic spikes and abusive clients.
- Design an API Gateway (Kong) : Layer 7 routing, authentication, request transformation, and plugin ecosystems.
- Design a Distributed Cache : Partitioning, cache locality, and load balancing across memory nodes.
- API Contract and Integration Design : Designing resilient ingress APIs, idempotent mutations, and graceful error envelopes.
Engineering Trade-offs
Interviewers frequently evaluate trade-offs around routing algorithms, kernel network bottlenecks, and high availability tiers. Be prepared to compare round robin, least connections, and consistent hashing, explaining how each behaves when managing stateful WebSocket backends.
Layer 4 vs Layer 7 Packet Path Walkthrough
Layer 4 Load Balancer (NAT Mode) Packet Flow:
1. Client sends initial packet:
Source: 198.51.100.5:4321
Destination: 10.0.0.1:80 (Virtual IP)
2. Load balancer receives packet and selects Server 2 (10.0.1.6:8080):
Rewrites destination address from 10.0.0.1:80 to 10.0.1.6:8080 (Destination NAT)
Rewrites the client source to 10.0.0.254:55000 (Source NAT) so the response returns through the LB
Stores flow mapping in connection table:
{198.51.100.5:4321, 10.0.0.1:80} -> {10.0.1.6:8080, 10.0.0.254:55000}
Forwards packet to backend network:
Source: 10.0.0.254:55000
Destination: 10.0.1.6:8080
3. Server 2 produces response:
Source: 10.0.1.6:8080
Destination: 10.0.0.254:55000
4. Load balancer intercepts return packet and consults connection table:
Reverses source and destination NAT, restoring the client-facing tuple:
Source: 10.0.0.1:80
Destination: 198.51.100.5:4321
Forwards packet to client.
Client Transparency:
Client communicates only with 10.0.0.1:80 and remains unaware of the internal backend topology.
Operational Trade-offs:
Per-packet cost: Two address rewrites per packet (inbound DNAT and outbound SNAT).
Throughput: Kernel-level routing can process millions of packets per second with appropriately tuned hardware.
Bottleneck: The load balancer must process all inbound and outbound traffic, creating a bandwidth constraint for egress-heavy workloads such as media streaming.Layer 4 Load Balancer (Direct Server Return / DSR Mode) Packet Flow:
1. Client sends initial packet:
Source: 198.51.100.5:4321
Destination: 10.0.0.1:80 (Virtual IP)
2. Load balancer receives packet and selects Server 2:
Rewrites the Layer 2 Ethernet frame MAC address to Server 2 MAC address.
Leaves the Layer 3 IP headers completely unmodified.
Packet arrives at Server 2 with Destination still set to 10.0.0.1:80.
(Server 2 must have the Virtual IP configured on a local loopback interface with ARP disabled).
3. Server 2 responds DIRECTLY to the client:
Source: 10.0.0.1:80 (Virtual IP)
Destination: 198.51.100.5:4321
Return packets travel across the network switch directly to the client gateway, bypassing the load balancer entirely.
Operational Advantages:
Asymmetric bandwidth optimization: In typical web traffic, incoming requests average 1 KB while outgoing responses exceed 100 KB.
The load balancer avoids return-path processing, so in the illustrative 1 KB request vs 100 KB response mix it handles roughly 100x less response traffic than NAT mode.
Network Constraints:
Classic ARP-based DSR typically requires the load balancer and backend servers to share a Layer 2 broadcast domain.
Routed or tunneled DSR variants can relax this constraint. DSR still depends on the load balancer for the inbound path.Layer 7 Load Balancer (Full Application Inspection) Request Flow:
1. Client initiates connection:
Client conducts TLS handshake with the load balancer.
The load balancer terminates TLS and decrypts HTTP payload.
Parses request: GET /api/v1/users/123 with Host: api.example.com
2. Load balancer evaluates application routing rules:
Path matches prefix "/api/v1/users/*" and routes to pool "user-service".
Least-connections algorithm selects Server 3 because it has the lowest active connection count.
3. Load balancer establishes an upstream TCP connection to Server 3:
Source: Load_Balancer_Internal_IP
Destination: 10.0.1.7:8080
Injects proxy headers: X-Forwarded-For: 198.51.100.5, X-Forwarded-Proto: https
Sends proxied request over a trusted private network or an HTTPS/mTLS upstream connection.
4. Server 3 processes request and replies to load balancer:
Load balancer receives upstream HTTP response, compresses payload where configured, re-encrypts over the client TLS session, and transmits response to client.
Operational Trade-offs:
Overhead: Full TCP connection termination, TLS cryptographic decryption, and HTTP parsing on every request.
Throughput: Approximately 100K requests per second per instance in this illustrative comparison, vs millions of packets per second for tuned Layer 4 forwarding.
Capabilities: Enables content-based routing, header rewriting, response caching, compression, and Web Application Firewall inspection.Connection Table Scaling
Scaling the Layer 4 NAT Connection Table: Memory Footprint Calculation: At 10M concurrent connections with 128 bytes per connection state record: 10,000,000 connections * 128 bytes = 1.28 GB RAM. Connection Table Operations: Insert: O(1) hash table insertion upon receiving initial TCP SYN packet. Lookup: O(1) hash lookup per packet indexed by 5-tuple: (src_ip, dst_ip, src_port, dst_port, protocol). Delete: O(1) removal upon TCP FIN/RST or inactivity timeout expiry. Memory Bus and Cache Pressure: While 1.28 GB easily fits into physical RAM, a 10M-entry hash table causes continuous CPU L1/L2/L3 cache misses under random traffic access. Processing 1M packets per second requires 1M random memory lookups per second, which can create memory bus pressure and CPU cache thrashing under this access pattern. High-Performance Solution: DPDK and XDP: DPDK can bypass the operating system network stack and process raw packets in user space using poll-mode drivers. XDP executes very early in the Linux networking path inside the kernel, reducing per-packet overhead without being a full user-space bypass. Pre-allocate connection table memory within huge pages (2 MB / 1 GB pages) to reduce Translation Lookaside Buffer (TLB) misses. Pin worker threads to dedicated CPU cores to reduce scheduler and context-switch overhead. Result: Can target 10M+ packets per second on appropriately tuned commodity server hardware, depending on packet size, NIC, CPU, and rule complexity. Connection Inactivity Timeouts: TCP Established state: 300 seconds is an illustrative default. TCP TIME_WAIT state: 120 seconds is an illustrative value and is operating-system dependent. UDP stream state: 30 seconds is an illustrative default. Aggressive timeouts reclaim table capacity quickly, but risk resetting connections for slow or idle clients. SNAT Port Exhaustion: NAT mode consumes source ports for each source-IP and destination tuple. At high fan-out, use multiple source IPs or equivalent port allocation strategies and monitor available SNAT ports to avoid ephemeral-port exhaustion. Connection Table High Availability Synchronization: The active load balancer replicates active connection state records to the standby node through a dedicated low-latency state replication channel. A production design can target synchronization for more than 99% of established connection states, but any unsynchronized flows may still reset during failover. Any missing connection states receive TCP RST packets, prompting clean client-side reconnection. Under Direct Server Return (DSR), no return connection table is maintained at the load balancer, simplifying failover mechanics while the inbound path still depends on the active node.
Maglev Lookup and Bounded Load Spillover
Maglev Lookup and Bounded Loads:
Limitations of Standard Ring Consistent Hashing:
Standard hash rings can produce significant imbalance when a node sits after a large gap on the ring.
For example, Server S3 might receive 40% of traffic while S1 and S2 each receive 30%.
While adding virtual nodes improves balance, it requires 100 to 200 virtual nodes per server, increasing memory and requiring O(log N) binary searches per lookup.
Google Maglev Lookup Table Construction:
Precompute a lookup table of prime size M, typically M = 65537 entries.
Each backend server generates a deterministic permutation preference list across table indices:
preference[i] = (offset + i * skip) mod M
offset = hash1(backend_name) mod M
skip = hash2(backend_name) mod (M - 1) + 1
Table Population Algorithm:
Iterate round robin through all backend servers, allowing each server to claim its next preferred unoccupied table index until all M slots are filled:
Iteration 0: Backend A claims its highest preference available slot.
Iteration 1: Backend B claims its highest preference available slot.
Iteration 2: Backend C claims its highest preference available slot.
Continue round-robin allocation until all M slots are occupied.
Packet Routing Lookup:
backend = table[hash(5-tuple) mod M]
Key Architectural Properties:
Minimal Disruption: Adding or removing 1 backend out of N instances moves only roughly 1/N of total table entries.
Near-Perfect Uniformity: Each backend receives approximately floor(M/N) or ceil(M/N) table entries for an evenly built table.
Sub-Microsecond Speed: Direct array indexing provides O(1) lookup time. A 65,537-entry table can remain cache-friendly when the stored backend reference is compact.
Bounded Loads Enforcement:
Enforces a target upper load bound: no backend server should receive more than (1 + epsilon) * average_load, where epsilon is typically 0.25 (targeting a cap of 25% over average when capacity allows).
If the hashed backend target exceeds its capacity bound:
The load balancer walks sequentially to the next backend in the preference list.
Repeats until locating an instance operating below the capacity threshold.
Outcome: Limits workload skew while retaining the connection stability of consistent hashing. If every backend is at capacity, admission control and load shedding still protect the tier.Defense in Depth: Reducing Single Points of Failure
Defense in Depth: Reducing Single Points of Failure Tier 1: Local Active-Passive Redundancy (VRRP and Keepalived) Two load balancer nodes share a common Virtual IP address (VIP). The active node processes all incoming network traffic. The passive standby node monitors the active instance using heartbeats sent every 1 second. If the active node fails, the passive node broadcasts a gratuitous ARP to claim the VIP within about 3 seconds. New connections resume through the same VIP. Established NAT flows can still reset unless connection state is synchronized. Advantages: Simple and widely used operational model. Trade-offs: Standby node remains idle during normal operations (50% hardware utilization), and failover can cause a 1 to 3 second drop in active TCP connections. Tier 2: Active-Active Clustering with DNS Round Robin Authoritative DNS returns multiple load balancer IP addresses (for example, 10.0.0.1 and 10.0.0.2). All load balancer nodes process production traffic simultaneously. If a node fails, automated DNS health probes remove the failing IP address from the record set. Advantages: 100% hardware utilization across all provisioned nodes under normal operation. Trade-offs: Recursive resolver and client DNS caching can delay failover until cached records expire, and DNS round robin does not adapt dynamically to backend load variations. Tier 3: Anycast BGP Routing Multiple load balancer clusters advertise the identical public IP address using Border Gateway Protocol (BGP). Internet routers select a path according to routing policy and network topology, which often brings traffic to a nearby healthy load balancer cluster. If an entire cluster fails, BGP route withdrawals cause upstream routers to reconverge traffic to the next nearest cluster. Advantages: Automated, network-level failover within 10 to 30 seconds with natural geographic distribution. Trade-offs: Requires BGP-capable network infrastructure, and convergence time depends on the network and policy configuration. The illustrative design assumes 10 to 30 seconds during routing shifts. Commonly used by large-scale edge networks and implemented by several infrastructure platforms. Cloud providers may abstract the same global routing ideas behind managed products. Tier 4: Distributed Sidecar Data Plane (Service Mesh) Augments centralized internal load balancing with decentralized Envoy sidecar proxies deployed alongside each application instance. Client applications dispatch requests to their local sidecar, which performs client-side load balancing, service discovery, and circuit breaking directly. Advantages: Moves much of the east-west load balancing and resiliency work into the service data plane, reducing centralized proxy bottlenecks for internal traffic. Trade-offs: Increases operational complexity and consumes additional CPU and memory resources per container pod. Commonly deployed in Kubernetes environments with service mesh control planes such as Istio or Linkerd.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.