Interview Setup
Interview Prompt
Design a video conferencing system like Zoom supporting 10M concurrent meetings with 50M media streams, 175 Tbps total bandwidth, and sub-150ms end-to-end latency.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| SFU (forward streams) or MCU (mix audio/video server side)? | SFU scales with the amount of media it forwards and keeps publisher upload to one stream per sender. MCU requires server side decode and encode for each participant and is useful for legacy SIP clients. |
| P2P fallback for 1:1 calls or always through SFU? | P2P can eliminate roughly half of the media relay traffic generated by an SFU for a 1:1 call, but direct connectivity can fail behind symmetric NAT without TURN. |
| Recording in scope: cloud-side or client side? | Assuming 5M recorded meetings/day, 1.25 PB/day requires a dedicated ingest pipeline, object storage, and transcoding farm. |
| Max participants per meeting: 10, 100, or 1,000? | Drives simulcast layers, active speaker detection, and SFU fan-out design. |
Scope
In scope
- SFU vs MCU media servers
- WebRTC data channels
- Bandwidth estimation (REMB/TWCC)
- Simulcast/SVC layers
- Screen sharing
- Recording pipeline
Out of scope (state explicitly)
- Detailed frontend/UI pixel implementation
- Org structure, staffing, and hiring plan
Functional Requirements
Start by confirming meeting size, recording, and breakout room scope with your interviewer. The stated scale assumes 10M concurrent meetings, 50M concurrent participant streams, and up to 1,000 participants in a meeting. Ask whether you are designing signaling only or the full media plane including SFU topology, drawing from the real time patterns in Network Protocols: HTTP, gRPC, WebSocket & DNS.
- 1:1 and group video/audio calls (up to 1,000 participants)
- Screen sharing with annotation support
- Real-time chat during calls (text, file sharing)
- Meeting scheduling with calendar integration
- Meeting recording with cloud storage and playback
- Virtual backgrounds, noise cancellation
- Breakout rooms for large meetings
- Waiting room with host admission control
- Raise hand, reactions, polls during meetings
- Join via browser (WebRTC), desktop app, or phone (PSTN dial-in)
- End-to-end encryption for 1:1 calls
Non-Functional Requirements
Your interviewer will care most about glass-to-glass latency and the signaling versus media split. Call out sub-150 ms as the target and explain why UDP/SRTP never touches the signaling hot path.
- Ultra Low Latency: < 150ms glass-to-glass for acceptable experience. < 400ms is tolerable
- High Availability: 99.99%. Active calls should continue through partial failures.
- Scalability: 100M+ concurrent connected users, 10M+ concurrent meetings, 50M concurrent media participants, and up to 1,000 participants in a meeting
- Adaptive Quality: Gracefully degrade on poor networks
- Global: Edge servers in 50+ regions to minimize round-trip latency
- Reliability: No dropped frames under normal conditions. Reconnect within 2 seconds
- Security: E2E encryption for 1:1 calls. Use DTLS-SRTP for group media and TLS for HTTP/WebSocket signaling.
Capacity Estimations
Run this math before you size SFU clusters. Concurrent meetings and streams per meeting determine the media bandwidth budget. At peak, signaling is lightweight compared with 175 Tbps of RTP forwarding. The recording estimate assumes 5M recorded meetings per day, each lasting 30 minutes at 500 MB per hour.
| Metric | Calculation | Value |
|---|---|---|
| Concurrent meetings | Given (peak load assumption) | 10M |
| Avg participants per meeting | Given (typical workload assumption) | 5 |
| Concurrent streams | Given (peak load assumption) | 50M |
| Bandwidth per participant | Given (assumption documented in value) | 2 Mbps down + 1.5 Mbps up (720p) |
| Total bandwidth | 50M x 2 Mbps down + 50M x 1.5 Mbps up | 100 Tbps down + 75 Tbps up = 175 Tbps aggregate → Edge distribution critical |
| Recording storage / day | 5M recorded meetings/day x 30 min x 500 MB/hr | 1.25 PB/day |
| Signaling messages / sec | 10M meetings x 2 signals/sec | 20M/sec |
Architecture Diagram
In the room: separate signaling (WebSocket/TCP) from media (UDP/SRTP). Draw the two planes explicitly because they have different latency and reliability requirements.
Walk your interviewer through the signaling versus media split first. Participants connect to signaling servers over WebSocket for session orchestration such as join, SDP exchange, ICE candidate trickling, and meeting controls. They then exchange media with SFU nodes over UDP/SRTP. TURN relays are used when NAT traversal fails. Recording bots join the SFU as invisible participants and consume the required media tracks. In this design, these are two separate planes.
The interview hinges on separating signaling from media. Signaling uses reliable transport and carries 20M messages/sec, while media uses loss-tolerant UDP/SRTP and carries 175 Tbps at peak. At 10M concurrent meetings and 50M streams, clients discover their SFU assignment through signaling, while media never touches the signaling hot path.
Component Deep Dives
Overview
Video conferencing has two distinct planes. The signaling plane manages session state, while the media plane carries real time RTP forwarding. Signaling does not carry video packets, and the SFU forwards media without decoding or encoding. The sections below cover join flow, NAT traversal, breakout rooms, PSTN bridging, topology choices, and network adaptation.
Signaling Server: Session Orchestration
Signaling is entirely separate from media transport. Clients open a persistent WebSocket to the signaling cluster for join and leave events, SDP offer and answer exchange, ICE candidate trickling, and meeting controls. At 20M signals/sec the signaling tier is stateful but lightweight because it never touches RTP packets. Redis holds ephemeral meeting state such as the participant list, SFU assignment, and TURN credentials with a short TTL that is refreshed while the meeting is active.
The join flow starts with a REST POST /join call. The response contains a signaling token, an assigned SFU endpoint, and ICE server configuration. The client then opens the WebSocket, sends a join message with its SDP offer, and receives an SDP answer through signaling. ICE candidates are exchanged in both directions until the media path is established over UDP/SRTP or a TURN relay.
NAT Traversal: ICE, STUN, and TURN
After signaling exchanges SDP, the client and SFU establish the media path through ICE. STUN provides server reflexive candidates so the endpoints can discover reachable public addresses. ICE connectivity checks select a viable candidate pair. When direct connectivity fails, especially behind restrictive or symmetric NAT, the client falls back to a TURN relay. TURN allocation state is kept with short lived credentials so relay access is scoped to the meeting.
The scaling concern is relay bandwidth. If NAT traversal fails for 30% of participants, TURN relays 15M streams at 3.5 Mbps each, which is about 52.5 Tbps through relay infrastructure, rounded to roughly 52 Tbps at the stated planning precision. Deploy TURN at edge PoPs, prefer SFU direct paths whenever possible, and scale relay capacity horizontally.
Breakout Rooms
Breakout rooms are logical submeetings spawned from a parent webinar. The host assigns participants to room buckets. Signaling updates each client's meeting_id and triggers an ICE restart on a new SFU shard. Parent meeting state is preserved in Redis so participants can be recalled atomically. Recording and chat are scoped to the active room. Returning to the main session is another signaling-driven SFU reassignment, not a full page reload.
Host: POST /api/meetings/{id}/breakout
{ "rooms": 4, "assignment": "manual" | "random", "duration_minutes": 15 }
Signaling broadcasts to each participant:
{ "type": "breakout_assign", "room_id": "br-2", "sfu_endpoint": "sfu-us-east-7...", "token": "..." }
On recall:
{ "type": "breakout_close", "return_sfu": "sfu-us-east-3...", "main_meeting_id": "m_123" }
State in Redis:
breakout:{parent_meeting_id} → Hash { room_id → [user_ids], timer_expires_at }PSTN Dial-In (SIP Gateway)
Phone participants cannot run WebRTC in a browser. They connect through a SIP/PSTN media gateway that bridges the PSTN leg to the SFU as a synthetic participant. The gateway decodes G.711 audio from the phone network, transcodes it to Opus, and injects it into the SFU media path. Dial-in numbers map to meeting PINs, and DTMF tones authenticate the caller. This is where MCU-style processing can appear in an SFU-first architecture. Only the phone leg needs transcoding rather than the entire meeting.
PSTN join flow: 1. User dials +1-800-xxx-xxxx, enters meeting PIN via DTMF 2. SIP gateway validates PIN → looks up meeting_id in PostgreSQL 3. Gateway registers as synthetic participant "Phone User" on SFU 4. Audio: PSTN (G.711 64 Kbps) ↔ gateway ↔ Opus ↔ SFU ↔ other participants Capacity: ~5K concurrent PSTN legs per gateway cluster. Scale horizontally by region. Latency budget: PSTN adds ~100-200ms one-way vs native WebRTC
Why SFU Over MCU and Mesh?
Topology choice drives both cost and latency. Mesh works for 1:1, but peer connections grow as N². MCU reduces each client to one upload and one download stream, but the server must decode and re-encode the participants. SFU forwards packets without transcoding, which allows per-subscriber quality selection.
| Topology | Pros | Cons | Use When |
|---|---|---|---|
| Mesh | No server cost, lowest latency for 1:1 | N² streams, unscalable past 4-5 users | 1:1 calls |
| MCU | Client uploads 1, downloads 1 stream | Massive server CPU (decode+encode), no per-client quality | Legacy telephony |
| SFU ✅ | Low server CPU (no decode), per-client quality selection, simulcast | More downstream bandwidth than MCU | 2+ participant calls |
SFU is attractive for group calls because the server forwards packets without decoding or encoding. Simulcast lets each receiver get an appropriate quality layer, and the SFU can forward selected streams rather than every available stream. In comparable workloads, this can use roughly 10x less media processing CPU than an MCU, although the exact ratio depends on codecs and meeting size.
Simulcast & SVC
Receivers on 3G and gigabit fiber cannot share one bitrate. Simulcast sends three independent encodes (180p / 720p / 1080p), and the SFU picks the right layer per subscriber. SVC achieves a similar outcome with a single layered bitstream. It can use bandwidth more efficiently, but it is harder to implement consistently across codec vendors.
Simulcast: Sender encodes 3 independent streams (High: 1080p @ 2.5 Mbps, Medium: 720p @ 1 Mbps, Low: 180p @ 150 Kbps). The SFU selects a layer for each receiver. An active speaker can receive High from the speaker and Low from others. Gallery view can receive Medium from all, and mobile on 3G can receive Low.
SVC: Single layered stream (base + enhancement layers). SFU drops upper layers for constrained receivers. Advantage: single encode, seamless quality transitions. VP9 and AV1 can support SVC. Product deployments vary by codec and platform.
SFU Architecture & Cascading
Each meeting pins to a primary SFU with a hot standby that already holds participant state. For global meetings, regional SFU nodes cascade streams so each video track crosses a region boundary exactly once. In the representative scenario below, this results in 30 Mbps of inter-region relay versus 870 Mbps if every participant connected directly.
Each meeting has a primary SFU and a hot standby. The standby receives the participant list and ICE candidates from signaling. On primary failure, signaling detects the heartbeat miss in < 3 sec, signals clients to reconnect to the backup, and clients perform an ICE restart in about 1 to 2 sec.
Multi-region cascading: A meeting with participants in US-East, EU-West, and APAC uses SFU-US-East, SFU-EU-West, and SFU-APAC. Each region fans out locally, so each stream crosses a region boundary once through an SFU to SFU relay. Bandwidth is 30 Mbps inter-region versus 870 Mbps if all participants connected directly, which is a 29x saving.
Network Quality Adaptation
Video quality must degrade gracefully before the call drops entirely. Clients monitor packet loss and RTT and then step down resolution when network conditions deteriorate. FEC adds redundant packets for recovery without retransmission. NACK requests retransmission from the SFU jitter buffer only when RTT stays under 100 ms.
Client monitors packet loss, RTT, and available bandwidth. Loss < 2% maps to 1080p 30fps. Loss 2-5% maps to 720p. Loss 5-10% maps to 360p 15fps. Loss > 10% maps to audio-only. FEC sends redundant packets (10% overhead) for recovery without retransmission. NACK requests retransmission from the SFU jitter buffer and is useful only if RTT < 100ms.
Large Meeting Optimization (1,000 participants)
A thousand WebRTC connections per client is impractical. Webinar mode keeps a small set of active speakers on the interactive SFU path, while passive viewers consume a CDN HLS stream with a 5 to 10 second delay. Viewers can be promoted to the interactive speaker role without changing the overall architecture.
Tiered architecture: Tier 1 Speakers (5-10) use full SFU send and receive video. Tier 2 Viewers receive the webinar stream through CDN HLS with a 5-10s delay and can be promoted to an interactive speaker role when needed. In the representative 1,000 participant case, keeping about 10 speakers on WebRTC leaves about 990 viewers on CDN delivery.
API Design
REST APIs
type MeetingId = string;
type UserId = string;
type Jwt = string;
type SfuEndpoint = string;
interface CreateMeetingRequest {
title?: string;
scheduledStart?: string;
maxParticipants?: number;
}
interface CreateMeetingResponse {
meetingId: MeetingId;
}
interface JoinMeetingRequest {
audio: boolean;
video: boolean;
}
interface JoinMeetingResponse {
meetingId: MeetingId;
sfuEndpoint: SfuEndpoint;
signalingToken: Jwt;
iceServers: string[];
}
interface RecordingRequest {
enabled: boolean;
}
interface BreakoutRequest {
rooms: number;
assignment: "manual" | "random";
durationMinutes: number;
}POST /api/meetings
GET /api/meetings/{meeting_id}
POST /api/meetings/{meeting_id}/join
POST /api/meetings/{meeting_id}/record
POST /api/meetings/{meeting_id}/breakoutWebSocket Signaling Protocol
type MeetingId = string;
type UserId = string;
type Jwt = string;
type SfuEndpoint = string;
type MediaOptions = {
audio: boolean;
video: boolean;
};
type JoinMeetingMessage = {
type: "join";
meetingId: MeetingId;
signalingToken: Jwt;
media: MediaOptions;
};
type SdpMessage = {
type: "offer" | "answer";
sdp: string;
target?: "sfu";
from?: "sfu";
};
type IceCandidateMessage = {
type: "ice_candidate";
candidate: string;
sdpMid: string;
};
type MeetingControl =
| { type: "mute"; targetUserId: UserId; media: "audio" | "video" }
| { type: "raise_hand" }
| { type: "screen_share_start"; streamId: string };
type ParticipantEvent =
| { type: "participant_joined"; user: { id: UserId; name: string } }
| { type: "active_speaker"; userId: UserId };
type SfuAssignmentEvent = {
type: "sfu_assignment";
meetingId: MeetingId;
sfuEndpoint: SfuEndpoint;
signalingToken: Jwt;
}Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff 440 Login Timeout: WebSocket connection session expired, client reconnect is required
Data Model
PostgreSQL
CREATE TABLE meetings (
meeting_id UUID PRIMARY KEY,
host_id UUID NOT NULL,
title TEXT,
scheduled_start TIMESTAMPTZ,
actual_start TIMESTAMPTZ,
actual_end TIMESTAMPTZ,
status TEXT DEFAULT 'scheduled',
password_hash TEXT,
settings JSONB,
max_participants INT DEFAULT 100 CHECK (max_participants <= 1000),
created_at TIMESTAMPTZ DEFAULT NOW()
);
CREATE TABLE meeting_participants (
meeting_id UUID REFERENCES meetings(meeting_id),
user_id UUID,
display_name TEXT,
join_time TIMESTAMPTZ,
leave_time TIMESTAMPTZ,
role TEXT DEFAULT 'participant',
PRIMARY KEY (meeting_id, user_id, join_time)
);Store the meeting password as a hash rather than plaintext. The default participant limit is 100, while the product requirement permits configuration up to 1,000.
Redis: Active Meeting State
HSET meeting:active:{meeting_id} host "u_123" sfu_endpoint "sfu-us-east-3.zoom.us:443" participant_count 15
ZADD meeting:participants:{meeting_id} {join_timestamp} {user_id}
SET turn:alloc:{user_id}:{meeting_id} "turn-server-17:3478" EX 3600Event Bus Design (Kafka)
Call-quality telemetry and chat archives ride a separate asynchronous bus. Kafka is never on the media hot path. SFU nodes and client SDKs emit quality metrics without blocking RTP forwarding, which lets operations teams alert on regional packet-loss spikes without adding media latency.
Topic: call-quality-events
Partitions: 64
Partition key: meeting_id
Retention: 30 days
Topic: chat-events
Partition key: meeting_id
Producer: SFU + client SDK (async, non-blocking)
Event: { meeting_id, user_id, metric: "packet_loss"|"jitter"|"bitrate", value, timestamp }
Consumer groups:
1. analytics: ClickHouse call quality per region, network diagnostics
2. alerting: packet_loss > 5% sustained triggers notification to host
3. chat-archive: persist chat history to PostgreSQL
Media path is WebRTC/SRTP: Kafka is analytics only, never on media hot path
DLQ: call-quality-events-dlqFault Tolerance
SFU Failure Recovery
Each meeting has a primary SFU and a hot standby. On primary failure, signaling detects the heartbeat miss in < 3 sec, signals all clients to reconnect to the backup, and clients perform an ICE restart in about 1 to 2 sec. In the warm SFU optimization, the standby already has ICE candidates cached, so the system can skip full ICE negotiation and reconnect in < 1 sec.
End-to-End Encryption (E2E)
Challenge: the SFU needs RTP headers for routing but should not decrypt media. Solution: Insertable Streams API (SFrame). The client encrypts the encoded media payload before RTP packetization, while the RTP headers remain available for routing. An authenticated key exchange can run through the signaling channel. In this design, E2EE is scoped to 1:1 calls. It can also be extended to group calls, but recording then requires a trusted endpoint with access to decrypted media, and server side processing features must move to the client or another trusted endpoint.
Recording Architecture
A recording bot joins the SFU as an invisible participant and receives the required streams. It performs real time compositing for an active speaker or gallery layout and writes chunks to local SSD. When the meeting ends, the system uploads the chunks to S3, triggers transcoding through the dedicated Video Transcoding Pipeline, and generates searchable transcripts through speech-to-text.
Noise Cancellation & Virtual Background
Run ML models locally on the client, such as RNNoise for audio and MediaPipe for segmentation. Processing on the client keeps these operations off the server and preserves the stated privacy goal. GPU acceleration can use WebGPU or WebGL.
Additional Considerations
Why WebRTC Specifically?
WebRTC provides NAT traversal through ICE, STUN, and TURN. It also provides encrypted media with SRTP and DTLS, adaptive bitrate with transport feedback, codec negotiation through SDP for H.264, VP8, VP9, and AV1, browser native support without plugins, and subsecond latency. An alternative is a custom UDP protocol. That can reduce protocol overhead, but it does not work directly in browsers without additional native support.
Meeting Quality Monitoring
Every 5 seconds, each client reports packet loss, jitter, RTT, resolution, CPU usage, encode time, and available bandwidth. The backend aggregates these metrics in ClickHouse for real time dashboards, alerts when > 20% of participants have poor quality, regional analysis, and historical 95th percentile quality by region.
Interview Walkthrough
- 25-minute cut
Skip arch50/arch75 depth unless staff.
- Signaling (WebSocket) separate from media (UDP/SRTP) (8 min)
- SFU forwards streams, while MCU is used only for legacy clients (9 min)
- TURN relay only when NAT traversal fails (8 min)
- Separate signaling, which uses reliable WebSocket over TCP, from media transport, which uses loss-tolerant UDP/SRTP. Make this split immediately.
- Walk through signaling join as a REST token exchange followed by WebSocket SDP and ICE exchange, then the SFU media path. Use Load Balancing Algorithms to explain SFU assignment and regional capacity balancing. Signaling at 20M messages/sec stays off the media hot path.
- Cover breakout rooms by showing signaling reassigning participants to submeeting SFU shards. Parent state in Redis enables atomic recall.
- Explain PSTN dial-in via SIP gateway, where the phone leg transcodes G.711 to Opus and joins the SFU as a synthetic participant.
- Recommend SFU over mesh and MCU for 2+ participants: server forwards packets without decode/encode, enabling simulcast and per-client quality selection.
- Explain simulcast: the sender encodes 3 layers, 1080p, 720p, and 180p. The SFU forwards the appropriate layer per receiver based on bandwidth estimation.
- Regional SFU cascading for global meetings, ensuring each stream crosses a region boundary once via SFU-to-SFU relay, not N² direct connections.
- Large meetings (1000+): tier active speakers on WebRTC, route passive viewers through CDN (HLS) to cap connection count.
- Network adaptation uses packet loss thresholds to trigger resolution downgrade. FEC and NACK provide recovery, and audio-only mode starts above 10% loss.
- Common pitfall: choosing MCU because clients upload only one stream, even though the server must decode and re-encode N streams, making CPU the bottleneck at scale.
Engineering Trade-offs
P2P (Mesh) vs SFU vs MCU: Media Routing Architecture
Video conferencing balances media quality, client bandwidth, and server compute. The central architectural choice is selective forwarding with an SFU versus transcoding with an MCU.
P2P (Mesh) has no SFU relay cost, and direct peer media can provide end to end transport encryption. Client upload = N-1 streams. It often becomes impractical beyond 4-5 participants because bandwidth grows quickly, so use it for 1:1 calls.
SFU ⭐ keeps client upload at 1 stream. The server forwards packets without transcoding, and near realtime latency can stay < 150ms. Client download still scales with N. At 100 participants, one receiver could receive up to 99 streams. Use it for 2-1000 participants, with webinar tiering for very large meetings.
MCU gives each client 1 download stream, which works well on weak devices. Server compute requires N decoders + N encoders per meeting for personalized outputs (EXPENSIVE). It can add 500-1000ms of latency. Use it for PSTN integration and weak-client meetings.
Hybrid topology example: use SFU for interactive participants and introduce MCU only when server side composition or legacy interoperability is required. Large webinars can combine MCU for panelists with SFU and CDN delivery for attendees. The exact topology threshold is a workload and product decision.
UDP vs TCP for Media Transport
TCP retransmits lost packets. Under loss, head of line blocking can delay subsequent packets and produce 100-300ms audio or video freezes.
UDP with custom loss handling: FEC or packet skipping can keep the stream moving. A lost video packet can corrupt one macroblock for 1 frame, which is often barely noticeable. Lost audio (20ms) can be masked by PLC, and the Opus codec has built-in FEC.
Rule: Use UDP with FEC and PLC for media. Use reliable WebSocket over TCP, or another reliable QUIC based signaling transport, for signaling.
Simulcast: Serving All Clients at Their Optimal Quality
Each sender encodes 3 resolutions simultaneously (180p, 720p, 1080p). The listed layer profile is 1080p @ 2.5 Mbps, 720p @ 1 Mbps, and 180p @ 150 Kbps, which is 3.65 Mbps of aggregate encoded video before protocol overhead when all three layers are sent. The SFU forwards the appropriate layer per receiver based on bandwidth estimation through RTCP feedback. This is materially higher than the separate 1.5 Mbps single stream target used in the capacity estimate, so the two figures should not be conflated. The benefit is that each receiver gets an appropriate quality level without server side transcoding. A 360p profile can be used as an additional congestion fallback outside the three layer baseline. SVC is an alternative with a single layered bitstream and one upload stream, but the codec path is more complex. VP9 and AV1 can support SVC. A simulcast profile can also be chosen for H.264 compatibility.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.