Core Concept

System Design Interview Patterns

Seven reusable architectural patterns that appear across most HLD interviews — real-time delivery, async workers, contention control, read/write scaling, blob uploads, and multi-step workflows — plus a practical time budget for a 45-minute session.


1. What It Is

Meta-concept: system design interviews follow a repeatable structure. We clarify requirements, estimate scale, sketch high-level boxes, then deep-dive where the interviewer points.

What:

A catalog of recurring solution shapes and a delivery rhythm for 45-minute system design interviews.

Primary purpose:

Recognize which pattern applies, name it confidently, and allocate time so the interviewer sees requirements, architecture, and depth — not just one area.

Usually used for:

Product systems (feeds, chat, commerce), infra-adjacent designs (queues, storage), and any problem where the same building blocks reappear with different nouns.

2. Core Mental Model

Most interview problems reduce to one or two of these axes:

📡 Push vs Pull

Does the client need live updates? If yes, plan fan-out and connection management. If no, cache + pagination on pull is enough.

⚡ Sync vs Async

Can the user wait? Short requests stay on the API path. Anything over ~200 ms of CPU or I/O belongs on a queue with status polling or webhooks.

📈 Read vs Write Hot

Identify the hot dimension first. Read-heavy → replicas and CDN. Write-heavy → shard keys, batch writes, or aggregate counters.

In the room

Spend the first 5 minutes on scope and non-goals — not architecture. Drive the conversation: "I'll start with requirements, then back-of-envelope, then HLD." Ask which deep-dive they prefer. Silence while drawing is fine; narrate your reasoning.

3. Why It Matters in HLD

Pattern recognition saves interview time — we map the prompt to a shape before picking technologies. Three lenses:

Needed When:

You have 45 minutes and must show structured thinking — not a laundry list of technologies.

Avoids:

Spending 20 minutes on APIs with no diagram, or jumping to Kafka before clarifying scale and consistency needs.

Optimizes For:

Signal density — interviewers score pattern recognition, trade-off articulation, and communication under time pressure.

4. Architecture & Data Flow

Walk the composite diagram as interview narration. Step 1 — Sync API: client-facing REST/gRPC for interactive requests. Step 2 — Async workers: queue for long-running jobs. Step 3 — Read scale: cache + replicas + CDN on the hot path. Step 4 — Write scale: shard or pre-aggregate when math demands. Step 5 — Real-time: WebSocket/SSE + pub/sub if live updates required. Omit boxes the requirements do not need.

Loading...

5. Key Characteristics

Seven recurring patterns (plus AI and platform) — we map the problem statement before naming databases:

  • Seven recurring patterns — map the problem statement to one or two before choosing databases (plus Applied AI for modern stacks):
PatternCore MechanicPrimary RoleExample Problems
Real-Time UpdatesWebSocket, SSE, or long-poll with pub/sub fan-out to connected clients.Push live state (chat, bids, dashboards) without polling every resource.#07 Chat, #50 Presence, #90 Bidding, #35 Live Likes
Long-Running Tasks
  • Enqueue work to a durable queue
  • stateless workers process asynchronously.
Keep API latency low for transcoding, email, reports, and batch jobs.#28 Job Scheduler, #62 Image Pipeline, #93 Email Service
Contention ControlDistributed locks, compare-and-swap (CAS), or single-writer queues.Prevent double-booking, overselling inventory, or duplicate payments.#23 Ticketing, #68 Flash Sale, #29 Distributed Lock, #24 Payment
Scaling ReadsRead replicas, layered cache, and CDN edge delivery.Absorb read-heavy traffic without overloading primary databases.#01 URL Shortener, #06 News Feed, #09 Instagram, #17 CDN
Scaling WritesSharding, write batching, and pre-aggregation counters.Spread write load and reduce hot-row pressure on a single node.#04 Unique ID, #35 Live Likes, #42 Like Count, #38 Analytics
Large Blob HandlingPresigned multipart upload directly to object storage.Offload multi-GB files from application servers and API gateways.#87 S3, #25 Dropbox, #15 Video Streaming, #26 Pastebin
Multi-Step ProcessesSaga compensations or workflow orchestrators (Temporal-style).Coordinate checkout, booking, and onboarding across multiple services.#22 E-Commerce, #24 Payment, #96 Hotel Booking, #103 Temporal
Applied AI SystemsLLM gateway, RAG retrieval, vector DB, sandboxed agents, multi-agent graphs.Ship AI products with latency, cost, and safety constraints — not just model calls.#108 Chat, #109 Cursor, #111 RAG, #113 Vector DB, #114 Cloud Agent, #115 Multi-Agent
Platform & InfraGit hosting, CI/CD runners, K8s control plane, secrets/KMS, search clusters.Design the tools engineers use to ship and operate software at scale.#107 GitHub, #118 CI/CD, #119 Kubernetes, #116 Secrets, #117 Elasticsearch

In the room

Spend the first five minutes on scope and non-goals. Say "I'll start with requirements, then back-of-envelope, then HLD" — that pacing signal matters as much as the diagram.

6. Strategic Tradeoffs

Pattern reuse accelerates design but over-application loses credibility — we compare:

BenefitCost
Pattern reuse — once you recognize the shape (feed, chat, checkout), you spend less time inventing from scratchOver-application — forcing WebSockets or sharding when a simple REST + cache design suffices loses credibility
Structured pacing — a time budget prevents drowning in API details before drawing architecture
  • Rigid scripts — interviewers may jump to deep dives early
  • adapt while keeping scope explicit

7. Failure / Bottleneck Awareness

WebSocket storms, lock contention, SSE mismatch, multipart orphans — we lead with these:

🔌 WebSocket Connection Storms

Problem: A celebrity goes live and millions of clients open WebSocket connections to a single region. Connection memory and fan-out CPU saturate before application logic runs.

Mitigation: Regional connection gateways, connection limits per user, SSE for one-way feeds where bidirectional channels are unnecessary, and shard fan-out by topic or room ID.

🔒 Lock Contention Under Flash Traffic

Problem: A flash sale uses a global distributed lock on inventory rows. Lock wait queues grow; P99 latency spikes and checkout times out.

Mitigation: Pre-decrement counters in Redis with CAS, partition inventory by SKU shard, or serialize purchases per SKU via a single-partition queue instead of coarse global locks.

📡 SSE vs WebSocket Mis-Match

Problem: A one-way stock ticker uses bidirectional WebSockets, wasting connection memory and complicating load balancer sticky-session config when SSE would suffice.

Mitigation: SSE for server→client only feeds; reserve WebSockets for chat and collaborative editing where the client must push frequently.

📦 Multipart Upload Orphans

Problem: Clients start multipart uploads but never complete them. Incomplete parts accumulate storage cost and clutter lifecycle policies.

Mitigation: Short-lived presigned URLs, lifecycle rules to abort incomplete uploads after 24 hours, and server-side finalize webhooks that validate checksum before marking the object visible.

8. How to Read Our Articles (~20 min)

Use our article reading passes to build pattern fluency before mock interviews:

PassWhat to read
First pass (~20 min)Interview Setup → Architecture Diagram → Capacity → Deep-Dive Probes → Walkthrough
Mock interview prep
  • Re-read probes aloud
  • draw HLD from memory
  • compare to article diagram
Staff depthDesign Evolution, Operational Reality, cross-links to core concepts

Full curated path: Prep Plan hub. Overlapping problems (e.g. #06 vs #08) — read the essential article first; variants add product-specific constraints only.

9. Curated Learning Path

Follow the curated learning path when preparing systematically — concepts before problems:

PhaseFocus
Week 1Core concepts #01, #04, #05, #12, #31, #32
Weeks 2–4Arch 25 — 25 highest-frequency problems
Weeks 5–8Arch 50 — senior depth + infra patterns
Weeks 9–12Arch 75 — AI (#108–#115) + platform (#117–#119)

10. Common HLD Usage (Interview Timing)

The timing table below is our default 45-minute pacing — adapt when the interviewer steers early:

Interview PhaseTime BudgetWhat to Cover
Requirements & scope~5 minFunctional vs non-functional, scale assumptions, in/out of scope
Entities & relationships~2 minCore nouns, ownership boundaries, read vs write paths
API surface~5 minKey endpoints, idempotency keys, pagination, error contracts
High-level design~15 minBoxes-and-arrows diagram, data flow, bottleneck callouts
Deep dives~10 minInterviewer-chosen topics: sharding, fan-out, failure modes
Buffer / trade-offs~8 minExplicit trade-offs, evolution path, monitoring hooks

Treat the table as a default — senior interviewers often allocate more time to deep dives if your HLD is crisp. Always leave ~2 minutes to summarize trade-offs and next evolution steps.

11. Decision Signals

🎯 Reach for these patterns when:
  • Real-time: "Users see updates within seconds" → WebSocket/SSE + pub/sub fan-out.
  • Async work: "Video processing takes minutes" → queue + workers + job status API.
  • Contention: "Only one seat left" or "exactly once charge" → CAS, idempotency keys, or saga.
  • Read scale: "Millions of reads, few writes" → cache-aside + read replicas + CDN.
  • Write scale: "Billions of events per day" → shard by user/time, batch inserts, counter aggregation.
  • Large files: "Upload 5 GB video" → presigned multipart to object storage.
  • Multi-step: "Reserve, pay, ship — any step can fail" → saga or workflow engine.
  • Applied AI: "Multi-agent research + codegen pipeline" → state graph (#115), LLM gateway (#108), optional RAG (#111).

13. Deep Dive (Optional)

Pattern Composition: A Live Auction Example

Real interviews rarely isolate a single pattern. An ad auction or flash sale typically composes four at once:

  1. Real-time bids arrive over WebSocket; a regional gateway publishes to a partitioned Kafka topic keyed by auction ID.
  2. Contention on the winning bid uses CAS in Redis — only increment if the new bid exceeds the current high by the minimum tick.
  3. Read scaling serves auction catalog pages from CDN + edge cache; bid history reads come from read replicas lagging ~100 ms behind the leader.
  4. Settlement after auction close triggers a saga: lock funds → record winner → notify loser wallets with compensating releases on failure.

Walking through this composition in ~3 minutes demonstrates pattern fluency without drawing every box — a strong signal at senior level.

Geographic Proximity Routing

Location-aware products — ride matching, food delivery, nearby friends, edge CDN selection — share a pattern distinct from generic read scaling:

  1. Index by geography: Store entities in geohash, S2, or H3 cells so "nearby drivers" becomes a bounded cell lookup, not a full table scan.
  2. Regional partitioning: Route users to the datacenter closest to their coordinates; cross-region queries only when the search radius spans boundaries.
  3. Freshness vs accuracy: Driver GPS updates every 3–5 seconds via WebSocket; matchmaking reads a slightly stale position cache — strong consistency on exact lat/lng is unnecessary, but results must refresh within one tick.
  4. Fallback hierarchy: Expand search radius or adjacent geohash cells when supply is thin — product logic layered on top of spatial indexing.

Name this pattern when problems mention Uber, Yelp, Tinder radius, or "find nearest X" — it combines real-time updates, spatial data structures, and regional sharding without over-sharding on day one.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...