1. What It Is
When an interviewer asks "how would this scale to a million users?" they want a staged roadmap β not a Kubernetes cluster on slide one. We add complexity only when a measured bottleneck forces our hand.
What:
The path from a single server to a distributed setup: separate tiers, cache reads, scale stateless app servers, then partition writes when one node is not enough.
Primary purpose:
Add capacity only when a measured bottleneck appears β not because the problem mentions millions of users.
Usually used for:
Scaffolding high-level design structures, justifying scalability, and proving architecture maturity.
2. Core Mental Model
We scale in stages β solve one bottleneck before adding the next layer of complexity:
π¦ Separate App & DB
First step of scaling: move the database off the application server to its own dedicated node with optimized disk I/O.
π§± Go Stateless First
Store user sessions inside Redis rather than server memory. This lets you boot up/kill application nodes dynamically behind a load balancer.
πͺ Shard the Writes
Caches and read replicas scale reads indefinitely, but scaling high-volume writes eventually requires sharding primary tables across nodes.
In the room
The biggest trap is jumping straight to microservices and sharding because the prompt says "1M users." Walk through levels 1β4, tie each step to a bottleneck (read load, session stickiness, write QPS), and mention you'd validate with back-of-the-envelope math before splitting databases.
3. Why It Matters in HLD
Scaling interviews test whether we add complexity on purpose or by reflex. We justify each layer with a measured bottleneck, using three lenses:
Needed When:
Interviewers want a staged roadmap, not a jump to microservices and sharding on day one.
Avoids:
Premature optimization, wasting operational budgets on unused clusters, and system collapses due to unshielded database hotspots.
Optimizes For:
Operational budget economy, infrastructure simplicity, developer velocity, and system resilience.
4. Architecture & Data Flow
Walk Level 3β4 as staged interview steps, not a final-state diagram. Step 1 β Separate tiers: pull the database off the app server so CPU and disk I/O do not fight. Step 2 β Cache hot reads: Redis shields the primary from repeat queries. Step 3 β Go stateless: sessions live in Redis or JWT so any app node can serve any user. Step 4 β Horizontal app pool: L7 load balancer distributes across identical stateless instances. Step 5 β CDN edge: static and cacheable API responses terminate close to users. Step 6 β Shard writes: only when back-of-envelope math proves single-primary limits are exceeded.
5. Key Characteristics
We tie each growth stage to a user band and a bottleneck β the table is our interview roadmap, not a day-one blueprint:
- Each scaling stage adds one layer of complexity β match infrastructure to measured bottlenecks, not user-count labels alone:
| Stage | User Scale | Core Focus | Architecture Details |
|---|---|---|---|
| Level 1: Single Node Monolith | 0 to 1,000 Users | Simplicity, rapid feature validation, low cost. | App and Database share a single server container (e.g. AWS EC2 instance). |
| Level 2: Multi-Tier Cache | 1,000 to 100,000 Users | Offload read query pressures from primary database. | Dedicated application server + standalone DB instance + Redis cache + DB secondary replicas. |
| Level 3: Horizontal LB Scale | 100,000 to 500,000 Users | Eliminate server bottlenecks, support failover survivability. | Stateless app server pool behind Load Balancer (L7) + Anycast DNS + CDN edge caching. |
| Level 4: Database Sharding | 500,000 to 1M+ Users | Scale database storage capacity and write throughput. | Microservices division + Kafka queue decoupling + horizontally sharded database clusters. |
In the room
Say out loud: "I would validate with back-of-envelope math before sharding." That one line separates candidates who scale on user-count labels from those who scale on measured QPS.
6. Strategic Tradeoffs
Horizontal scaling is not free. We name what we buy and what we pay:
| Benefit | Cost |
|---|---|
| Stateless Server Scaling (stateless instances let you scale application servers horizontally in seconds behind LBs) |
|
| Microservices Decoupling (independent teams deploy isolated boundaries dynamically, boosting velocity) | Distributed Operational Complexity (managing distributed transactions, tracing, network latency, and RPC errors) |
7. Failure / Bottleneck Awareness
These are the walls candidates hit when they skip stages β we call them out before the interviewer does:
Problem: Storing user session data (or local caching files) directly inside application server disk/memory forces sticky routing rules. If a server dies, active users lose their state immediately.
Mitigation: Enforce total app statelessness by extracting session storage to an external memory tier (Redis) or utilizing JWT tokens.
Problem: Read replicas scale read queries infinitely, but write queries must still route to the single primary database instance, eventually exhausting disk I/O.
Mitigation: Introduce write-decoupling buffers (e.g. Kafka or SQS queues) or partition primary databases horizontally via sharding keys.
8. Common HLD Usage
Level 4 decomposition below is illustrative β we use it to show we know what microservices look like, not to propose them on day one:
The diagram below is an illustrative Level 4 decomposition β not a prescription to microservice everything on day one. User bands in the scaling matrix are order-of-magnitude guides; always reconcile against your back-of-the-envelope QPS math before splitting services:
9. Decision Signals
Reach for progressive scaling when the interviewer wants a roadmap, not a single snapshot:
- The evaluator requests a comprehensive roadmap showing how your whiteboard design handles future user metrics leaps.
- You are designing web-based transactional systems with rapid business growth expectations.
- You must prove why horizontal scaling is mathematically superior to expensive high-tier hardware nodes (vertical scaling).
11. Deep Dive (Optional)
The Shared-Nothing Architecture
To scale horizontally to 1M+ active users, production systems embrace the **Shared-Nothing (SN) Architecture**. In this model:
- Stateless app servers operate fully independently, maintaining zero shared execution locks or file handles.
- Request context is passed inside cookies or transient header tokens (JWT).
- Any active node can process any user request. If node A crashes, the load balancer steers packets to node B without session losses.
This decouples the system's execution bounds completely, turning server capacity into an elastic utility.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.