Interview Setup
Interview Prompt
Design a podcast platform like Spotify/Apple Podcasts: creators upload episodes, listeners stream or download, RSS feeds syndicate to aggregators, with subscriptions and playback sync.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Stream vs download ratio? | 500M downloads/day x 50 MB = 25 PB/day CDN, which dominates total cost. |
| RSS compatibility required? |
|
| Dynamic ad insertion? |
|
| Offline listening? | Signed download URLs with expiry vs open RSS enclosures. |
Scope
In scope
- Episode upload + processing
- Multi-format transcode
- CDN streaming/download
- RSS feeds
- Subscriptions + home feed
- Playback progress sync
Out of scope (state explicitly)
- Music catalog licensing
- Podcast discovery machine learning
- Live podcast streaming
Functional Requirements
Start by confirming RSS syndication scope with your interviewer. External aggregators poll your public feeds, so stable enclosure URLs and full-file download bandwidth dominate infrastructure cost. Audio delivery integrates closely with the Video Transcoding Pipeline and the CDN and Edge Delivery architecture.
- Upload podcasts: Creators upload audio episodes with metadata (title, description, show notes, chapters)
- Streaming playback: Stream episodes with seeking, speed control (0.5x to 3x), skip silence
- Downloading: Offline download for listening without internet
- RSS feed: Generate and serve RSS/Atom feeds for distribution to Apple Podcasts, Spotify, etc.
- Show management: Create shows (series), manage episodes, schedule future releases
- Discovery: Browse by category, charts (top podcasts), search, recommendations
- Subscriptions: Users subscribe to shows; new episodes appear in their feed
- Playback state: Sync playback position across devices (resume where you left off)
- Analytics: Download counts, listener demographics, retention graphs per episode
- Monetization: Dynamic ad insertion (pre-roll, mid-roll, post-roll), premium subscriptions
Non-Functional Requirements
Starting audio playback in under 2 seconds globally is the primary non-functional target, necessitating CDN-cached audio files and pre-transcoded bitrates. Strict RSS specification compliance is equally critical because a single malformed feed XML breaks distribution across Apple Podcasts, Spotify, and external podcast apps simultaneously.
- Playback Start: Audio begins playing within 2 seconds globally
- Availability: 99.99%: listeners can always play episodes
- Durability: Audio files never lost
- Scalability: 5M+ podcast shows, 100M+ episodes, 100M+ active listeners/month
- Global: Low-latency playback worldwide via CDN
- Bandwidth Efficiency: Adaptive bitrate; compressed audio formats (Opus, AAC)
- Analytics Accuracy: Download/listen counts accurate within ±1%
- RSS Compliance: Valid RSS 2.0 with iTunes podcast extensions
Capacity Estimations
RSS feed polling by aggregators creates a massive read spike independent of active listener counts: 5M shows polled by 10 platforms every 15 minutes generates tens of millions of feed requests per hour. Factor this polling load into CDN and cache sizing alongside audio egress.
| Metric | Calculation | Value |
|---|---|---|
| Total shows | Given (assumption documented in value) | 5M |
| Total episodes | Given (assumption documented in value) | 100M |
| New episodes / day | Given (assumption documented in value) | 100K |
| Avg episode duration | Given (typical workload assumption) | 45 minutes |
| Avg episode size | Given (typical workload assumption) | 50 MB (128 kbps MP3) |
| Upload storage / day | 100K x 50 MB | 5 TB |
| Total storage | Given (assumption documented in value) | 5 PB |
| Daily active listeners | Given (assumption documented in value) | 30M |
| Concurrent streams | Given (peak load assumption) | 5M |
| Stream bandwidth | 5M x 128 kbps | 640 Gbps |
| Downloads / day | Given (assumption documented in value) | 500M (including RSS aggregators) |
| Download bandwidth | 500M x 50 MB | 25 PB / day |
Architecture Diagram
In the room: emphasize that RSS download syndication (25 PB/day) dominates streaming bandwidth because aggregators pull full files rather than incremental byte-range streams.
Structure the architecture around three consumers of the episode catalog: the native mobile app for streaming and playback state synchronization, RSS aggregators that continuously poll feeds, and the analytics pipeline tracking listen events. Upload and transcoding run asynchronously, ensuring nothing blocks the creator's publish response.
Creators upload episodes that pass through an asynchronous audio pipeline, while listeners stream from edge CDNs and RSS syndicators fetch from the published episode catalog.
The system leverages CloudFront CDN for high-availability RSS feed delivery and globally cached audio, utilizing PostgreSQL for primary transactional tables, Redis for volatile playback tracking and caching, and S3 for processing files and artwork assets.
Component Deep Dives
1. Audio Processing Pipeline
We start with the audio processing pipeline because every downstream surface, including RSS syndication, native app playback, and dynamic ad insertion, depends on normalized, multi-bitrate files landing in object storage.
Creators upload raw audio (WAV, FLAC, MP3). The pipeline normalizes loudness to industry standards, trims silence, transcodes into multiple formats and bitrates, and generates waveforms, chapters, and optional transcripts before anything is published.
Step 1: Validate + Probe (FFprobe: extract duration, format, code, reject if >12 hrs or >2GB)
Step 2: Normalize Audio
- Normalize loudness: target -16 LUFS (loudness standard for podcasts)
ffmpeg -i input.wav -af "loudnorm=I=-16:TP=-1.5:LRA=11" normalized.wav
- Silence trimming: remove > 3 seconds of silence at start/end
Step 3: Transcode to Multiple Formats/Bitrates
- MP3 128 kbps: Universal compatibility (RSS feed reference)
- AAC 128 kbps: iOS/Android native high quality
- Opus 48 kbps: Incredible compression for modern clients (50% smaller than MP3!)
Step 4: Chapter Markers (Embed in ID3/M4A tags: { title, start_time, end_time })
Step 5: Generate Waveform (RMS amplitude per 100ms for custom seekbars, ~50KB JSON)
Step 6: Speech-to-Text Transcription (Whisper API, cost-optimized: only run for shows with >100 subs)2. RSS Feed Service: The Core Distribution Mechanism
Podcasts are distributed via RSS. Every external aggregator (Spotify, Apple, Overcast) polls RSS feeds constantly.
Feed serving at scale:
5M shows x polled every 15-30 minutes by 10+ aggregators = ~30M feed requests/hour
Strategy:
1. Pre-generate RSS XML for each show → store in S3
2. Serve via CDN with 15-minute TTL
3. On new episode publish: regenerate feed XML → invalidate CDN cache
4. Conditional requests: ETag/If-Modified-Since → 304 Not Modified (saves 90% bandwidth)
Stable RSS URL format: https://feeds.example.com/shows/{show_id}/rss3. Playback Sync: Resume Across Devices
Syncs playback position dynamically, allowing a seamless transition from phone commute to desktop browser.
Sync mechanism:
Client reports position every 30 seconds:
POST /api/v1/playback/progress { episode_id, position_seconds, speed }
On opening:
GET /api/v1/playback/progress/{episode_id} → resumes from stored point
Storage: Redis
Key: playback:{user_id}:{episode_id}
Value: Hash { position, speed, duration, updated_at }
TTL: 90 days (auto-cleanup old progress)
Scale: 30M DAU x update every 30 sec = ~100K writes/sec (easily handled by 10 Redis Cluster shards)4. Dynamic Ad Insertion (DAI): The Revenue Engine
Rather than relying on static baked-in ads, Server-Side Ad Insertion (SSAI) stitches targeted commercial spots into audio streams dynamically at request time.
SSAI Splicing Flow:
1. Creator marks ad breaks: { "breaks": [{"position": 0, "type": "pre-roll"}, {"position": 1200, "type": "mid-roll"}] }
2. On request, Ad Decision Service evaluates demographics, frequency caps, and targets ads.
3. Audio Stitching Service:
- Splicing HLS segments dynamically on edge.
- Segmented playlist: [segment_pre, ad_1, segment_mid, ad_2]
- Allows pre-encoded segments cached on CDN separately. No real-time heavy CPU re-encoding.
Ad Impression Tracking:
Client fires event when passing ad bounds: { ad_id, event: "impression|start|50%|complete" } → Kafka → ClickHouseEvent Bus Design (Kafka)
Topic: episode-uploaded
Partitions: 64
Partition key: show_id (serialize processing per podcast show)
Retention: 7 days
Replication factor: 3, min.insync.replicas: 2
Producer: Creator Service on S3 upload complete
Event: { episode_id, show_id, s3_key, format, chapters[] }
Consumer groups:
1. transcode-pipeline: normalize (-16 LUFS), transcode to MP3, AAC, and Opus, and generate waveform JSON
2. rss-regenerator: rebuild RSS XML, purge CDN cache, and ping aggregator webhooks
3. subscriber-notifier: push new episode notifications to subscribed listeners
4. stt-pipeline: Whisper transcription and Elasticsearch full-text indexing
Topic: playback-events (play, seek, complete) for ClickHouse analytics and IAB metrics
Topic: ad-events (impression, 50%, complete) for monetization reporting
Topic: download-events for compliance filters in dynamic ad insertion
Sync path: upload ACK and episode status set to processing, with asynchronous publishing
DLQ: episode-uploaded-dlq after 3 retriesAPI Design
Upload Episode
POST /api/v1/shows/{show_id}/episodes
{
"title": "Episode 42: System Design",
"description": "In this episode...",
"audio_file_key": "uploads/ep-42-raw.wav",
"publish_at": "2025-03-14T08:00:00Z",
"season": 3,
"episode_number": 42,
"explicit": false,
"chapters": [
{"title": "Introduction", "start": 0},
{"title": "Main Topic", "start": 180},
{"title": "Interview", "start": 1200}
],
"ad_breaks": [
{"position": 0, "type": "pre_roll", "max_duration": 30},
{"position": 1200, "type": "mid_roll", "max_duration": 60}
]
}Stream Episode
GET /api/v1/episodes/{episode_id}/stream?format=aac&quality=128k
Response: 302 Redirect
Location: https://cdn.example.com/audio/ep-uuid/aac_128k.m4a
Or with ad insertion:
Location: https://cdn.example.com/dai/ep-uuid/playlist.m3u8Get User's Subscription Feed
GET /api/v1/feed?limit=20&cursor={last}
Response: 200 OK
{
"episodes": [
{
"episode_id": "ep-uuid",
"show": {"id": "show-uuid", "title": "Tech Talk", "art": "..."},
"title": "Episode 42: System Design",
"duration_seconds": 2700,
"published_at": "2025-03-14T08:00:00Z",
"progress": {"position": 1847, "percent": 68},
"stream_url": "https://cdn.example.com/audio/ep-uuid/aac_128k.m4a",
"download_url": "https://cdn.example.com/audio/ep-uuid/mp3_128k.mp3",
"file_size_bytes": 32400000
}
]
}Update Playback Progress
POST /api/v1/playback/progress
{
"episode_id": "ep-uuid",
"position_seconds": 1847,
"speed": 1.5,
"duration_seconds": 2700
}Get RSS XML (Aggregator facing)
<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:podcast="https://podcastindex.org/namespace/1.0">
<channel>
<title>My Podcast Show</title>
<link>https://example.com/shows/my-podcast</link>
<itunes:author>John Doe</itunes:author>
<itunes:category text="Technology"/>
<itunes:image href="https://cdn.example.com/art/show-123.jpg"/>
<item>
<title>Episode 42: System Design</title>
<enclosure url="https://cdn.example.com/audio/ep-42.mp3" length="57000000" type="audio/mpeg"/>
<pubDate>Fri, 14 Mar 2025 08:00:00 GMT</pubDate>
<itunes:duration>3600</itunes:duration>
<description>In this episode we discuss...</description>
<podcast:chapters url="https://cdn.example.com/chapters/ep-42.json"/>
</item>
</channel>
</rss>Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpointData Model
PostgreSQL: Core Relational Data
CREATE TABLE shows (
show_id UUID PRIMARY KEY,
creator_id UUID NOT NULL,
title VARCHAR(255) NOT NULL,
description TEXT,
category VARCHAR(50),
subcategory VARCHAR(50),
language CHAR(5),
artwork_url TEXT,
website_url TEXT,
rss_feed_url TEXT NOT NULL, -- public feed URL (stable, permanent)
explicit BOOLEAN DEFAULT FALSE,
subscriber_count INT DEFAULT 0,
total_episodes INT DEFAULT 0,
status ENUM('active', 'paused', 'archived') DEFAULT 'active',
created_at TIMESTAMP,
updated_at TIMESTAMP,
INDEX idx_category (category, subscriber_count DESC),
INDEX idx_creator (creator_id)
);
CREATE TABLE episodes (
episode_id UUID PRIMARY KEY,
show_id UUID NOT NULL,
title VARCHAR(255) NOT NULL,
description TEXT,
show_notes TEXT,
season SMALLINT,
episode_number SMALLINT,
duration_seconds INT,
audio_url_mp3 TEXT, -- CDN URL for MP3
audio_url_aac TEXT, -- CDN URL for AAC
audio_url_opus TEXT, -- CDN URL for Opus
original_s3_key TEXT,
file_size_bytes INT,
chapters JSONB,
ad_breaks JSONB,
transcript_url TEXT,
waveform_url TEXT,
explicit BOOLEAN DEFAULT FALSE,
status ENUM('draft','processing','scheduled','published','archived'),
published_at TIMESTAMPTZ,
created_at TIMESTAMP,
INDEX idx_show (show_id, published_at DESC),
INDEX idx_published (status, published_at DESC)
);
CREATE TABLE subscriptions (
user_id UUID NOT NULL,
show_id UUID NOT NULL,
subscribed_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
notifications BOOLEAN DEFAULT TRUE,
PRIMARY KEY (user_id, show_id),
INDEX idx_show (show_id) -- "who subscribes to this show"
);Redis Key Schemas
# Playback progress
playback:{user_id}:{episode_id} → Hash { position, speed, updated_at } (TTL: 90 days)
# User's episode queue
queue:{user_id} → List of episode_ids (ordered)
# RSS feed cache
rss:{show_id} → String (RSS XML blob) (TTL: 15 minutes)
# Podcast charts (Sorted Sets)
charts:top:{category} → Sorted Set { show_id: score }
charts:trending → Sorted Set { show_id: growth_score }
# Episode download counter (incremented in Redis, flushed daily to ClickHouse)
downloads:{episode_id}:{date} → INT (INCR) (TTL: 2 days)S3 Storage Layout
Bucket: podcast-originals (cross-region replicated, permanent)
/{show_id}/{episode_id}/original.wav
Bucket: podcast-processed (CDN-served)
/{show_id}/{episode_id}/mp3_128k.mp3
/{show_id}/{episode_id}/aac_128k.m4a
/{show_id}/{episode_id}/opus_48k.ogg
/{show_id}/{episode_id}/waveform.json
/{show_id}/{episode_id}/transcript.json
Bucket: podcast-artwork
/{show_id}/artwork_3000x3000.jpg
/{show_id}/artwork_600x600.jpgKafka Message Bus Topics
Topic: episode-published (triggers RSS regeneration and push notifications) Topic: playback-events (play, seek, complete events feeding ClickHouse analytics) Topic: download-events (download started events used for IAB compliance filters) Topic: ad-events (ad impressions feeding monetization statistics)
ClickHouse: Analytics DB
CREATE TABLE episode_plays (
episode_id UUID,
show_id UUID,
user_id UUID,
event_type Enum8('play'=0,'pause'=1,'seek'=2,'complete'=3,'download'=4),
position_seconds UInt32,
duration_seconds UInt32,
speed Float32,
platform Enum8('ios'=0,'android'=1,'web'=2,'rss'=3),
country FixedString(2),
city String,
event_date Date MATERIALIZED toDate(timestamp),
timestamp DateTime
) ENGINE = MergeTree()
PARTITION BY toYYYYMM(timestamp)
ORDER BY (show_id, episode_id, timestamp);Fault Tolerance
| Concern | Solution |
|---|---|
| Audio file corruption |
|
| CDN failure |
|
| RSS feed stale |
|
| Playback sync loss | Client buffers progress locally, then retries sync when online |
| Ad insertion failure | Serve episode without ads (degrade gracefully, which is better than no audio) |
| Processing pipeline failure |
|
| Download counter loss | Redis AOF plus batch flush to ClickHouse every hour, with ClickHouse as the source of truth |
Specific: RSS Polling Storm (Thundering Herd)
Aggregators sync feeds simultaneously on the hour, triggering a massive thundering herd request spike of 55K req/sec.
- CDN Edge Caching: Feeds are cached globally on CDN edges with a 15-minute TTL. Only 5% of requests hit the origin.
- Conditional Requests: Aggregators support ETag and If-None-Match headers. 90% of requests return
304 Not Modified, saving massive bandwidth. - WebSub PubSubHubbub Push: Pushes new episode announcements to aggregators in real-time webhook endpoints instead of regular polling, eliminating 99% of requests.
Specific: Download Counting Accuracy (IAB Standard)
Advertisers pay per 1000 downloads (CPM), making overcounting (fraud) or undercounting (lost revenue) highly sensitive issues.
IAB Podcast Measurement Guidelines:
1. Deduplication: In Flink stream, generate key = SHA256(ip + user_agent + episode_id).
Window of 24 hours. Ignore matches within this window.
2. Bot filtering: Filter out automated crawlers matching the IAB bot list.
3. Byte-range filtering:
- Ignore byte 0-1000 requests (metadata fetching only).
- Ignore downloads where total bytes served < 50% of episode size.Additional Considerations
Interview Walkthrough
- 25-minute cut
Skip arch50 and arch75 depth unless interviewing for a staff-level role.
- Upload, transcode MP3/AAC/Opus, and publish to CDN and RSS enclosures (5 min)
- RSS feed design: stable URLs, ETag caching, and aggregator polling storms (6 min)
- Playback progress synchronization across listener devices (5 min)
- Download bandwidth calculations: 500M downloads at 50 MB yielding 25 PB/day (5 min)
- Dynamic ad insertion architecture as an advanced staff topic (4 min)
- Lead with the dual delivery model: RSS feeds for aggregator syndication alongside direct streaming for native apps, both backed by the same storage origin.
- Explain CDN edge caching with 15-minute TTL and ETag with If-None-Match headers so that 90% of aggregator polls return 304 Not Modified without hitting the origin.
- Cover WebSub push to eliminate the hourly RSS polling thundering herd that spikes origin servers to 55K requests per second.
- Walk through multi-codec storage: MP3 in RSS for universal compatibility, and Opus or AAC for native app playback to save 50% on CDN bandwidth.
- Describe IAB-compliant download counting: deduplication windows, bot filtering, and byte-range thresholds processed inside a real-time Flink stream.
- Mention silence-skipping implemented as client-side precomputed bounds JSON, avoiding audio re-encoding while preserving chapter markers and ad cues.
- Avoid the common pitfall of counting every byte-range request as a distinct download, because metadata probing inflates CPM figures and compromises advertiser trust.
Engineering Trade-offs
1. Audio Codec Choice: MP3 vs AAC vs Opus
Podcast platforms continuously balance codec compatibility against network bandwidth, contrasting universal MP3 support for RSS syndication with Opus compression for streaming efficiency.
- MP3 128k: Universal compatibility across legacy automotive systems, older browsers, and desktop podcast clients. Mandatory for public RSS enclosure feeds, though it exhibits lower compression efficiency than modern codecs.
- AAC 128k: Native hardware support across iOS and Android mobile platforms, achieving 30% greater bandwidth efficiency than MP3 at equivalent fidelity.
- Opus 48k: Industry-leading speech compression. Opus at 48 kbps matches the perceptual quality of 128 kbps MP3 while reducing CDN egress costs by more than 50%. While ideal for proprietary mobile apps, it is not yet universally supported by third-party RSS aggregators.
Strategy: Encode and store all three formats. Reference MP3 inside the public RSS XML feed, and configure native mobile players to negotiate optimal support: Opus first, falling back to AAC, and finally MP3.
2. Silence Detection and Skip
Automatically skipping silent gaps to keep podcasts fast and engaging.
Detection: Analyze RMS amplitude in 50ms windows. If RMS < threshold for > 500ms, mark as silence. Adaptive threshold: Measure noise floor from first 2 seconds. Set threshold = noise_floor + 6 dB. 1. Client-Side Skip (Chosen): - Precompute silence bounds JSON: [(start1, end1), (start2, end2)]. - Client player seeks past bounds. - Extremely flexible, zero extra storage or re-encoding costs. 2. Server-Side Skip: - Trim silences and save a condensed MP3. - Saves CDN bandwidth, but breaks chapter markers and dynamic ad boundaries.
3. Podcast Discovery: Charts and Recommendations
For comprehensive ranking system designs, multi-stage candidate generation, and vector retrieval, explore the Video Recommendation Engine.
Charts Ranking Formula:
score = w1 x new_subscribers_7d + w2 x downloads_7d + w3 x listener_retention + w4 x growth_velocity
- High velocity weighting allows rising trending podcasts to break into charts.
- Computed daily via Spark and cached in Redis Sorted Sets: charts:top:{category}.
Recommendations:
- Collaborative Filtering: User-show subscription matrices factored for matching shows.
- Content-based: Generate semantic embeddings from title, description, and STT transcripts.
- Serve recommendations via cosine similarity rankings, cached in Redis.Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.