1. What It Is
When an interviewer asks how you'd distribute a large file or handle real-time media, you're really choosing who owns state and who pays for egress. Client-server and P2P are the two poles β most production systems land somewhere in between.
What:
In client-server, clients talk to a central host that owns state. In peer-to-peer, nodes act as equals and exchange data directly without a mandatory master.
Primary purpose:
Let you trade centralized governance and consistency against bandwidth cost and resilience when no single coordinator can carry all traffic.
Usually used for:
Large media distribution, real-time communication topologies, and centralized API gateways β often as a hybrid (tracker + P2P data plane).
2. Core Mental Model
We let the workload pick the topology β don't default to P2P just because it sounds modern:
π Single Point of Governance
Client-server fits when you need strong consistency, audit trails, and schema enforcement from one authority.
πΎ Decentralized Swarm Scale
P2P fits when egress cost dominates and peers can contribute upload bandwidth (game patches, large file swarms).
βοΈ Hybrid Coordinator
Centralize metadata (tracker, DHT bootstrap) for discovery, but move bytes peer-to-peer to avoid origin egress bills.
In the room
Pure P2P sounds elegant until you need audit logs, access control, or guaranteed delivery. If the prompt involves payments or user accounts, start with client-server and only add P2P for the data plane if egress cost is the bottleneck.
3. Why It Matters in HLD
Topology is not a aesthetic choice β it determines who owns consistency, who pays egress, and how hard failover is. We use three lenses to justify our pick:
Needed When:
Designing ultra-high egress networks (e.g. game patch downloads or video chunking) or masterless cluster states.
Avoids:
Massive cloud egress billing charges, server network interface saturations, and single-point-of-failure outages.
Optimizes For:
Bandwidth scale efficiency, failover partition tolerance, and network-level distribution speed.
4. Architecture & Data Flow
Narrate the diagram as two parallel interview paths. Client-server path: every request routes through central hosts β easy governance, predictable latency, origin pays all egress. P2P path: a discovery step (tracker or DHT) maps peers, then bytes flow on direct links β origin egress drops but control weakens. Hybrid path: centralize metadata and auth, decentralize bulk transfer. Say which path your prompt needs before drawing boxes.
5. Key Characteristics
Once we pick a topology, we compare dimensions side by side β interviewers expect us to articulate why client-server or P2P wins for this workload:
- Client-server, P2P, and hybrid topologies differ on control, consistency, and egress cost:
| Dimension | Client-Server | Peer-to-Peer | Hybrid Topology |
|---|---|---|---|
| Control Plane | Strictly Centralized | Fully Decentralized | Centralized trackers + P2P data |
| Consistency Model | Strong (Single source of truth) | Eventual (Gossip anti-entropy) | Tunable per chunk / operation |
| Network Egress Cost | Linear with traffic volume | Sub-linear (Peers upload data) |
|
In the room
If you draw P2P for a payment or booking system, expect pushback. Lead with client-server for the control plane and only add P2P to the data plane when egress cost is the stated bottleneck.
6. Strategic Tradeoffs
Neither topology is free. We state what we gain and what we give up:
| Benefit | Cost |
|---|---|
| Centralized Client-Server Auditing (trivial to enforce access controls, compliance rules, and single-source-of-truth states) | Infrastructure Egress Bills (paying for all data delivery bandwidth, leading to multi-million dollar hosting expenses) |
| Decentralized Peer Uploads (nodes absorb the bulk of bandwidth costs, scaling throughput naturally as peer count grows) | Malicious Peer Risks (peers can inject corrupted chunks, spoof trackers, or leak IP addresses) |
7. Failure / Bottleneck Awareness
P2P interviews almost always pivot to these failure modes β we bring them up proactively:
Problem: Home routers and symmetric NATs block inbound connections, so two peers often cannot open a direct socket.
Mitigation: Use STUN to discover public endpoints; fall back to TURN relay servers when direct paths fail (~10% of connections).
Problem: If downloaders never upload, the swarm runs out of seeds and new downloads stall.
Mitigation: Incentivize reciprocity β BitTorrent's tit-for-tat throttles peers that refuse to upload.
8. Common HLD Usage
Production systems rarely pick pure topologies. These examples show how real teams blend control with bandwidth economics:
| System | Topology Selection | Architectural Rationale |
|---|---|---|
| WhatsApp Messaging | Client-Server (XMPP Protocol) | Message ordering, offline queue, and delivery receipts require centralized relays. |
| Stripe Payments API | Client-Server (REST) | Ledger authority, idempotency keys, and PCI audit trails demand a single trusted origin. |
| BitTorrent Downloads | Hybrid (Tracker-assisted P2P) | Eliminates server egress costs by letting peer swarms distribute file chunks directly. |
| Steam Game Patches | Hybrid (CDN + P2P delivery) |
|
| Zoom Video Calls | Client-Server (SFU-based) |
|
9. Decision Signals
We default to client-server unless the prompt has a clear signal for decentralization. Use these checklists to decide:
- You need a single source of truth β ledgers, inventory counts, booking seats, or auth sessions where conflicting writes are catastrophic.
- Compliance, audit trails, or access control must be enforced centrally (payments, healthcare records, enterprise SaaS).
- Payloads are small and egress cost is manageable β most mobile APIs and CRUD backends never justify P2P complexity.
- Clients cannot reliably reach each other (symmetric NAT, mobile networks) and TURN relay costs would erase P2P savings.
- Strong ordering or exactly-once semantics matter more than bandwidth savings (chat, notifications, collaborative editing).
- You are designing massive bulk distribution networks (e.g., distributing 100 GB game updates to 10M concurrent users).
- You require masterless partition tolerance where any node crash must not block remaining operations.
- Origin egress is the dominant cost and peers can contribute upload bandwidth without compromising integrity checks at a central tracker.
11. Deep Dive (Optional)
Decentralized Discovery via DHT (Kademlia)
To remove the centralized tracker, P2P networks use a Distributed Hash Table (DHT) such as Kademlia:
- Both peers and file hashes are mapped to a shared 160-bit integer key space.
- Distance between keys is calculated using the XOR metric: d(x,y) = x β y.
- Peers store routing tables containing contacts to neighbors at varying XOR distances.
- To find a file chunk, peers query neighbors closer and closer to the file's hash in $O(\log N)$ hops.
This enables file lookup without a central registry, at the cost of more complex routing and weaker governance than client-server.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.