1. What It Is
Meta-concept: system design interviews follow a repeatable structure. We clarify requirements, estimate scale, sketch high-level boxes, then deep-dive where the interviewer points.
What:
A catalog of recurring solution shapes and a delivery rhythm for 45-minute system design interviews.
Primary purpose:
Recognize which pattern applies, name it confidently, and allocate time so the interviewer sees requirements, architecture, and depth — not just one area.
Usually used for:
Product systems (feeds, chat, commerce), infra-adjacent designs (queues, storage), and any problem where the same building blocks reappear with different nouns.
2. Core Mental Model
Most interview problems reduce to one or two of these axes:
📡 Push vs Pull
Does the client need live updates? If yes, plan fan-out and connection management. If no, cache + pagination on pull is enough.
⚡ Sync vs Async
Can the user wait? Short requests stay on the API path. Anything over ~200 ms of CPU or I/O belongs on a queue with status polling or webhooks.
📈 Read vs Write Hot
Identify the hot dimension first. Read-heavy → replicas and CDN. Write-heavy → shard keys, batch writes, or aggregate counters.
In the room
Spend the first 5 minutes on scope and non-goals — not architecture. Drive the conversation: "I'll start with requirements, then back-of-envelope, then HLD." Ask which deep-dive they prefer. Silence while drawing is fine; narrate your reasoning.
3. Why It Matters in HLD
Pattern recognition saves interview time — we map the prompt to a shape before picking technologies. Three lenses:
Needed When:
You have 45 minutes and must show structured thinking — not a laundry list of technologies.
Avoids:
Spending 20 minutes on APIs with no diagram, or jumping to Kafka before clarifying scale and consistency needs.
Optimizes For:
Signal density — interviewers score pattern recognition, trade-off articulation, and communication under time pressure.
4. Architecture & Data Flow
Walk the composite diagram as interview narration. Step 1 — Sync API: client-facing REST/gRPC for interactive requests. Step 2 — Async workers: queue for long-running jobs. Step 3 — Read scale: cache + replicas + CDN on the hot path. Step 4 — Write scale: shard or pre-aggregate when math demands. Step 5 — Real-time: WebSocket/SSE + pub/sub if live updates required. Omit boxes the requirements do not need.
5. Key Characteristics
Seven recurring patterns (plus AI and platform) — we map the problem statement before naming databases:
- Seven recurring patterns — map the problem statement to one or two before choosing databases (plus Applied AI for modern stacks):
| Pattern | Core Mechanic | Primary Role | Example Problems |
|---|---|---|---|
| Real-Time Updates | WebSocket, SSE, or long-poll with pub/sub fan-out to connected clients. | Push live state (chat, bids, dashboards) without polling every resource. | #07 Chat, #50 Presence, #90 Bidding, #35 Live Likes |
| Long-Running Tasks |
| Keep API latency low for transcoding, email, reports, and batch jobs. | #28 Job Scheduler, #62 Image Pipeline, #93 Email Service |
| Contention Control | Distributed locks, compare-and-swap (CAS), or single-writer queues. | Prevent double-booking, overselling inventory, or duplicate payments. | #23 Ticketing, #68 Flash Sale, #29 Distributed Lock, #24 Payment |
| Scaling Reads | Read replicas, layered cache, and CDN edge delivery. | Absorb read-heavy traffic without overloading primary databases. | #01 URL Shortener, #06 News Feed, #09 Instagram, #17 CDN |
| Scaling Writes | Sharding, write batching, and pre-aggregation counters. | Spread write load and reduce hot-row pressure on a single node. | #04 Unique ID, #35 Live Likes, #42 Like Count, #38 Analytics |
| Large Blob Handling | Presigned multipart upload directly to object storage. | Offload multi-GB files from application servers and API gateways. | #87 S3, #25 Dropbox, #15 Video Streaming, #26 Pastebin |
| Multi-Step Processes | Saga compensations or workflow orchestrators (Temporal-style). | Coordinate checkout, booking, and onboarding across multiple services. | #22 E-Commerce, #24 Payment, #96 Hotel Booking, #103 Temporal |
| Applied AI Systems | LLM gateway, RAG retrieval, vector DB, sandboxed agents, multi-agent graphs. | Ship AI products with latency, cost, and safety constraints — not just model calls. | #108 Chat, #109 Cursor, #111 RAG, #113 Vector DB, #114 Cloud Agent, #115 Multi-Agent |
| Platform & Infra | Git hosting, CI/CD runners, K8s control plane, secrets/KMS, search clusters. | Design the tools engineers use to ship and operate software at scale. | #107 GitHub, #118 CI/CD, #119 Kubernetes, #116 Secrets, #117 Elasticsearch |
In the room
Spend the first five minutes on scope and non-goals. Say "I'll start with requirements, then back-of-envelope, then HLD" — that pacing signal matters as much as the diagram.
6. Strategic Tradeoffs
Pattern reuse accelerates design but over-application loses credibility — we compare:
| Benefit | Cost |
|---|---|
| Pattern reuse — once you recognize the shape (feed, chat, checkout), you spend less time inventing from scratch | Over-application — forcing WebSockets or sharding when a simple REST + cache design suffices loses credibility |
| Structured pacing — a time budget prevents drowning in API details before drawing architecture |
|
7. Failure / Bottleneck Awareness
WebSocket storms, lock contention, SSE mismatch, multipart orphans — we lead with these:
Problem: A celebrity goes live and millions of clients open WebSocket connections to a single region. Connection memory and fan-out CPU saturate before application logic runs.
Mitigation: Regional connection gateways, connection limits per user, SSE for one-way feeds where bidirectional channels are unnecessary, and shard fan-out by topic or room ID.
Problem: A flash sale uses a global distributed lock on inventory rows. Lock wait queues grow; P99 latency spikes and checkout times out.
Mitigation: Pre-decrement counters in Redis with CAS, partition inventory by SKU shard, or serialize purchases per SKU via a single-partition queue instead of coarse global locks.
Problem: A one-way stock ticker uses bidirectional WebSockets, wasting connection memory and complicating load balancer sticky-session config when SSE would suffice.
Mitigation: SSE for server→client only feeds; reserve WebSockets for chat and collaborative editing where the client must push frequently.
Problem: Clients start multipart uploads but never complete them. Incomplete parts accumulate storage cost and clutter lifecycle policies.
Mitigation: Short-lived presigned URLs, lifecycle rules to abort incomplete uploads after 24 hours, and server-side finalize webhooks that validate checksum before marking the object visible.
8. How to Read Our Articles (~20 min)
Use our article reading passes to build pattern fluency before mock interviews:
| Pass | What to read |
|---|---|
| First pass (~20 min) | Interview Setup → Architecture Diagram → Capacity → Deep-Dive Probes → Walkthrough |
| Mock interview prep |
|
| Staff depth | Design Evolution, Operational Reality, cross-links to core concepts |
Full curated path: Prep Plan hub. Overlapping problems (e.g. #06 vs #08) — read the essential article first; variants add product-specific constraints only.
9. Curated Learning Path
Follow the curated learning path when preparing systematically — concepts before problems:
10. Common HLD Usage (Interview Timing)
The timing table below is our default 45-minute pacing — adapt when the interviewer steers early:
| Interview Phase | Time Budget | What to Cover |
|---|---|---|
| Requirements & scope | ~5 min | Functional vs non-functional, scale assumptions, in/out of scope |
| Entities & relationships | ~2 min | Core nouns, ownership boundaries, read vs write paths |
| API surface | ~5 min | Key endpoints, idempotency keys, pagination, error contracts |
| High-level design | ~15 min | Boxes-and-arrows diagram, data flow, bottleneck callouts |
| Deep dives | ~10 min | Interviewer-chosen topics: sharding, fan-out, failure modes |
| Buffer / trade-offs | ~8 min | Explicit trade-offs, evolution path, monitoring hooks |
Treat the table as a default — senior interviewers often allocate more time to deep dives if your HLD is crisp. Always leave ~2 minutes to summarize trade-offs and next evolution steps.
11. Decision Signals
- Real-time: "Users see updates within seconds" → WebSocket/SSE + pub/sub fan-out.
- Async work: "Video processing takes minutes" → queue + workers + job status API.
- Contention: "Only one seat left" or "exactly once charge" → CAS, idempotency keys, or saga.
- Read scale: "Millions of reads, few writes" → cache-aside + read replicas + CDN.
- Write scale: "Billions of events per day" → shard by user/time, batch inserts, counter aggregation.
- Large files: "Upload 5 GB video" → presigned multipart to object storage.
- Multi-step: "Reserve, pay, ship — any step can fail" → saga or workflow engine.
- Applied AI: "Multi-agent research + codegen pipeline" → state graph (#115), LLM gateway (#108), optional RAG (#111).
13. Deep Dive (Optional)
Pattern Composition: A Live Auction Example
Real interviews rarely isolate a single pattern. An ad auction or flash sale typically composes four at once:
- Real-time bids arrive over WebSocket; a regional gateway publishes to a partitioned Kafka topic keyed by auction ID.
- Contention on the winning bid uses CAS in Redis — only increment if the new bid exceeds the current high by the minimum tick.
- Read scaling serves auction catalog pages from CDN + edge cache; bid history reads come from read replicas lagging ~100 ms behind the leader.
- Settlement after auction close triggers a saga: lock funds → record winner → notify loser wallets with compensating releases on failure.
Walking through this composition in ~3 minutes demonstrates pattern fluency without drawing every box — a strong signal at senior level.
Geographic Proximity Routing
Location-aware products — ride matching, food delivery, nearby friends, edge CDN selection — share a pattern distinct from generic read scaling:
- Index by geography: Store entities in geohash, S2, or H3 cells so "nearby drivers" becomes a bounded cell lookup, not a full table scan.
- Regional partitioning: Route users to the datacenter closest to their coordinates; cross-region queries only when the search radius spans boundaries.
- Freshness vs accuracy: Driver GPS updates every 3–5 seconds via WebSocket; matchmaking reads a slightly stale position cache — strong consistency on exact lat/lng is unnecessary, but results must refresh within one tick.
- Fallback hierarchy: Expand search radius or adjacent geohash cells when supply is thin — product logic layered on top of spatial indexing.
Name this pattern when problems mention Uber, Yelp, Tinder radius, or "find nearest X" — it combines real-time updates, spatial data structures, and regional sharding without over-sharding on day one.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.