1. What It Is
Once you have more than one app server, you need a load balancer β but the algorithm matters as much as the box. Round-robin is the default; consistent hashing and least-connections solve sticky sessions and uneven load.
What:
Load balancers act as central gateways, routing incoming traffic streams across a fleet of backend servers to prevent hotspots and ensure high availability.
Primary purpose:
Preventing single-server exhaustion, providing failover routing, and maximizing overall system QPS capacity.
Usually used for:
Stateless web server pools, API gateways, reverse proxies, and microservices internal meshes.
2. Core Mental Model
L4 balances connections; L7 balances HTTP requests β pick the layer that sees enough context to route correctly:
π Layer 4 Transport
Low overhead routing. Looks only at IP addresses and TCP port numbers, forwarding raw packets to servers without opening them.
π§ Layer 7 Smart Routing
High intelligence. Decrypts HTTP/TLS traffic to steer requests dynamically based on URL paths (`/feed`), cookies, or headers.
π₯ Active Health Checks
Continuously ping downstream nodes. Automatically eject crashed servers from the active routing pool within seconds.
In the room
Candidates draw an LB circle and stop. Say which layer (L4 vs L7), which algorithm, and whether you need sticky sessions. If WebSockets or stateful gateways are involved, consistent hashing or IP hash beats blind round-robin.
3. Why It Matters in HLD
Load balancers distribute traffic and hide backend failures β algorithm choice affects stickiness, fairness, and tail latency. Three lenses:
Needed When:
Your system scales beyond a single server instance, requiring horizontal pools of web hosts to handle QPS demand.
Avoids:
Single-point-of-failure outages, server CPU overloads, uneven user requests distribution, and network congestion.
Optimizes For:
Request latency distributions, horizontal system scale boundaries, cluster resiliency, and network resource utilizations.
4. Architecture & Data Flow
Walk the request path as interview steps. Step 1 β DNS/L4 entry: client connects to VIP; L4 LB forwards TCP to a healthy backend. Step 2 β L7 routing: HTTP LB inspects path/header and routes to service pools. Step 3 β Algorithm pick: round-robin for uniform servers, least-connections for long-lived requests, consistent hash for cache affinity. Step 4 β Health checks: evict unhealthy nodes; state interval and threshold. Step 5 β Sticky sessions: only if state cannot move to Redis.
In the room
Say L4 vs L7 and name your algorithm in one breath β "L7 round-robin for stateless REST, consistent hash for cache nodes." That signals you know the layer, not just the box.
5. Key Characteristics
Algorithm choice depends on request length, statefulness, and hot-key distribution β we compare:
- Core Load Balancing Algorithms: Matching routing logic to workloads:
| Routing Algorithm | Steering Mechanic | Primary Advantage | Severe Drawback |
|---|---|---|---|
| Round Robin | Steers requests sequentially down the server list. |
| Ignores current server CPU load or request execution complexity. |
| Weighted Round Robin | Each server gets a weight β higher-capacity nodes receive proportionally more requests. | Handles heterogeneous hardware (8-core vs 32-core instances). |
|
| Least Connections | Steers requests to the server with the fewest active TCP sockets. | Ideal for long-lived states (WebSockets, streaming channels). | Requires tracking active states, adding small lookup overheads. |
| Power of Two Choices |
| Near-optimal load spread with O(1) state β used in Envoy, Google frontends. |
|
| IP / Source Hash | Routes requests by client IP hash modulo active node count. | Simple state persistence (sticky sessions) without central caches. | Re-routing occurs if servers scale up/down (forces Consistent Hashing). |
- L4 sees TCP connections; L7 sees HTTP requests β pick the layer that has enough context to route correctly:
| Layer Profile | Protocol Focus | CPU / Latency Overhead | Security Features |
|---|---|---|---|
| Layer 4 (L4) Transport | TCP / UDP packets routing. | Ultra-low (packets are routed immediately without reading HTTP bodies). | SSL decryption happens downstream at the application servers. |
| Layer 7 (L7) Application | HTTP / HTTP/2 / WebSocket headers parsing. | Moderate (requires full request decryption to read URL paths/headers). | Enables central SSL termination, header injection, and cookie sticky routing. |
Health check types: TCP connect (port open), HTTP GET (200 on /health), and gRPC health probes. Layer-7 checks catch app-level failures TCP misses; tune interval and unhealthy threshold to avoid flapping during deploys.
6. Strategic Tradeoffs
Better distribution and failover cost complexity and potential hotspots β we name both:
| Benefit | Cost |
|---|---|
| L4 Latency Economy (routes raw network packets at wire-speed without decrypting HTTP headers, maximizing throughput) | No Content-Aware Routing (unable to route requests by URL path, query params, or client cookie types) |
L7 Smart Routing (easily direct /users to the user service fleet, terminate SSL centrally, and enforce rate limits) | High CPU Overhead (decrypting TLS keys and inspecting HTTP bodies requires significant compute power) |
7. Failure / Bottleneck Awareness
Thundering herd on recovery, sticky-session traps, and health-check flapping β we lead with these:
Problem: One load balancer in front of the fleet becomes a single point of failure β if it dies, all traffic stops.
Mitigation: Active-passive LB pairs with VRRP virtual IP failover, or DNS anycast across multiple active entry points.
Problem: Round robin spreads new connections evenly, but chat traffic concentrates on a subset of sockets β some servers overload while others sit idle.
Mitigation: Use least-connections routing for persistent sockets and drain connections slowly during rolling deploys.
8. Common HLD Usage
These scenarios map to specific LB algorithms and layers:
- AWS ALB (L7): Leveraged to route HTTP path queries dynamically (e.g. `/api/v1/checkout` routes to payment fleets; `/static/logo.png` routes directly to S3 cache buckets).
- NGINX Proxy (L7): Deployed as a high-concurrency gateway handling TLS decryption, rate limiting, and sticky session cookies for downstream monolith pools.
- HAProxy (L4/L7): Configured at L4 to forward high-concurrency database queries directly across read replicas with sub-microsecond overheads.
9. Decision Signals
Add a load balancer when multiple app instances serve the same API or WebSocket gateway:
- You are designing any horizontally scaled system architecture with more than one server node.
- You need to decouple client applications from exact physical backend server IP addresses.
- You face high SSL/TLS handshake decryption overheads that must be terminated at a central infrastructure tier.
11. Deep Dive (Optional)
DNS Anycast Edge Steering
For massive global web entries, systems bypass local L7 load balancers at the entry-level by deploying **Anycast Routing** combined with **ECMP (Equal-Cost Multi-Path)** at the BGP edge routers layer:
- Multiple load balancer gateways in different datacenters advertise the exact same IP address to global routers using BGP.
- When a client sends TCP packets, global internet routers steer them geographically to the closest physical load balancer announcing that IP.
- If datacenter A goes offline, global routers automatically re-converge BGP routes to steer traffic to datacenter B, guaranteeing near-zero entryway downtime.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.