System Design Problem

Design a Podcast Delivery Platform

Commonly Asked By:SpotifyAppleGoogle

Interview Setup

Interview Prompt

Design a podcast platform like Spotify/Apple Podcasts: creators upload episodes, listeners stream or download, RSS feeds syndicate to aggregators, with subscriptions and playback sync.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Stream vs download ratio?500M downloads/day x 50 MB = 25 PB/day CDN, which dominates total cost.
RSS compatibility required?
  • Public rss_feed_url must stay stable
  • enclosure URLs point to CDN.
Dynamic ad insertion?
  • DAI adds per-user m3u8 manifest
  • separate from static MP3.
Offline listening?Signed download URLs with expiry vs open RSS enclosures.

Scope

In scope

  • Episode upload + processing
  • Multi-format transcode
  • CDN streaming/download
  • RSS feeds
  • Subscriptions + home feed
  • Playback progress sync

Out of scope (state explicitly)

  • Music catalog licensing
  • Podcast discovery machine learning
  • Live podcast streaming

Functional Requirements

Start by confirming RSS syndication scope with your interviewer. External aggregators poll your public feeds, so stable enclosure URLs and full-file download bandwidth dominate infrastructure cost. Audio delivery integrates closely with the Video Transcoding Pipeline and the CDN and Edge Delivery architecture.

  • Upload podcasts: Creators upload audio episodes with metadata (title, description, show notes, chapters)
  • Streaming playback: Stream episodes with seeking, speed control (0.5x to 3x), skip silence
  • Downloading: Offline download for listening without internet
  • RSS feed: Generate and serve RSS/Atom feeds for distribution to Apple Podcasts, Spotify, etc.
  • Show management: Create shows (series), manage episodes, schedule future releases
  • Discovery: Browse by category, charts (top podcasts), search, recommendations
  • Subscriptions: Users subscribe to shows; new episodes appear in their feed
  • Playback state: Sync playback position across devices (resume where you left off)
  • Analytics: Download counts, listener demographics, retention graphs per episode
  • Monetization: Dynamic ad insertion (pre-roll, mid-roll, post-roll), premium subscriptions

Non-Functional Requirements

Starting audio playback in under 2 seconds globally is the primary non-functional target, necessitating CDN-cached audio files and pre-transcoded bitrates. Strict RSS specification compliance is equally critical because a single malformed feed XML breaks distribution across Apple Podcasts, Spotify, and external podcast apps simultaneously.

  • Playback Start: Audio begins playing within 2 seconds globally
  • Availability: 99.99%: listeners can always play episodes
  • Durability: Audio files never lost
  • Scalability: 5M+ podcast shows, 100M+ episodes, 100M+ active listeners/month
  • Global: Low-latency playback worldwide via CDN
  • Bandwidth Efficiency: Adaptive bitrate; compressed audio formats (Opus, AAC)
  • Analytics Accuracy: Download/listen counts accurate within ±1%
  • RSS Compliance: Valid RSS 2.0 with iTunes podcast extensions

Capacity Estimations

RSS feed polling by aggregators creates a massive read spike independent of active listener counts: 5M shows polled by 10 platforms every 15 minutes generates tens of millions of feed requests per hour. Factor this polling load into CDN and cache sizing alongside audio egress.

MetricCalculationValue
Total showsGiven (assumption documented in value)5M
Total episodesGiven (assumption documented in value)100M
New episodes / dayGiven (assumption documented in value)100K
Avg episode durationGiven (typical workload assumption)45 minutes
Avg episode sizeGiven (typical workload assumption)50 MB (128 kbps MP3)
Upload storage / day100K x 50 MB5 TB
Total storageGiven (assumption documented in value)5 PB
Daily active listenersGiven (assumption documented in value)30M
Concurrent streamsGiven (peak load assumption)5M
Stream bandwidth5M x 128 kbps640 Gbps
Downloads / dayGiven (assumption documented in value)500M (including RSS aggregators)
Download bandwidth500M x 50 MB25 PB / day

Architecture Diagram

In the room: emphasize that RSS download syndication (25 PB/day) dominates streaming bandwidth because aggregators pull full files rather than incremental byte-range streams.

Structure the architecture around three consumers of the episode catalog: the native mobile app for streaming and playback state synchronization, RSS aggregators that continuously poll feeds, and the analytics pipeline tracking listen events. Upload and transcoding run asynchronously, ensuring nothing blocks the creator's publish response.

Creators upload episodes that pass through an asynchronous audio pipeline, while listeners stream from edge CDNs and RSS syndicators fetch from the published episode catalog.

Loading...

The system leverages CloudFront CDN for high-availability RSS feed delivery and globally cached audio, utilizing PostgreSQL for primary transactional tables, Redis for volatile playback tracking and caching, and S3 for processing files and artwork assets.

Component Deep Dives

1. Audio Processing Pipeline

We start with the audio processing pipeline because every downstream surface, including RSS syndication, native app playback, and dynamic ad insertion, depends on normalized, multi-bitrate files landing in object storage.

Creators upload raw audio (WAV, FLAC, MP3). The pipeline normalizes loudness to industry standards, trims silence, transcodes into multiple formats and bitrates, and generates waveforms, chapters, and optional transcripts before anything is published.

Loading...
Step 1: Validate + Probe (FFprobe: extract duration, format, code, reject if >12 hrs or >2GB)
Step 2: Normalize Audio
  - Normalize loudness: target -16 LUFS (loudness standard for podcasts)
    ffmpeg -i input.wav -af "loudnorm=I=-16:TP=-1.5:LRA=11" normalized.wav
  - Silence trimming: remove > 3 seconds of silence at start/end
Step 3: Transcode to Multiple Formats/Bitrates
  - MP3 128 kbps: Universal compatibility (RSS feed reference)
  - AAC 128 kbps: iOS/Android native high quality
  - Opus 48 kbps: Incredible compression for modern clients (50% smaller than MP3!)
Step 4: Chapter Markers (Embed in ID3/M4A tags: { title, start_time, end_time })
Step 5: Generate Waveform (RMS amplitude per 100ms for custom seekbars, ~50KB JSON)
Step 6: Speech-to-Text Transcription (Whisper API, cost-optimized: only run for shows with >100 subs)

2. RSS Feed Service: The Core Distribution Mechanism

Podcasts are distributed via RSS. Every external aggregator (Spotify, Apple, Overcast) polls RSS feeds constantly.

Loading...
Feed serving at scale:
  5M shows x polled every 15-30 minutes by 10+ aggregators = ~30M feed requests/hour
  
Strategy:
  1. Pre-generate RSS XML for each show → store in S3
  2. Serve via CDN with 15-minute TTL
  3. On new episode publish: regenerate feed XML → invalidate CDN cache
  4. Conditional requests: ETag/If-Modified-Since → 304 Not Modified (saves 90% bandwidth)
  
Stable RSS URL format: https://feeds.example.com/shows/{show_id}/rss

3. Playback Sync: Resume Across Devices

Syncs playback position dynamically, allowing a seamless transition from phone commute to desktop browser.

Sync mechanism:
  Client reports position every 30 seconds:
    POST /api/v1/playback/progress  { episode_id, position_seconds, speed }
  On opening:
    GET /api/v1/playback/progress/{episode_id}  → resumes from stored point

Storage: Redis
  Key: playback:{user_id}:{episode_id}
  Value: Hash { position, speed, duration, updated_at }
  TTL: 90 days (auto-cleanup old progress)

Scale: 30M DAU x update every 30 sec = ~100K writes/sec (easily handled by 10 Redis Cluster shards)

4. Dynamic Ad Insertion (DAI): The Revenue Engine

Rather than relying on static baked-in ads, Server-Side Ad Insertion (SSAI) stitches targeted commercial spots into audio streams dynamically at request time.

SSAI Splicing Flow:
  1. Creator marks ad breaks: { "breaks": [{"position": 0, "type": "pre-roll"}, {"position": 1200, "type": "mid-roll"}] }
  2. On request, Ad Decision Service evaluates demographics, frequency caps, and targets ads.
  3. Audio Stitching Service:
     - Splicing HLS segments dynamically on edge.
     - Segmented playlist: [segment_pre, ad_1, segment_mid, ad_2]
     - Allows pre-encoded segments cached on CDN separately. No real-time heavy CPU re-encoding.

Ad Impression Tracking:
  Client fires event when passing ad bounds: { ad_id, event: "impression|start|50%|complete" } → Kafka → ClickHouse

Event Bus Design (Kafka)

Topic: episode-uploaded
  Partitions: 64
  Partition key: show_id (serialize processing per podcast show)
  Retention: 7 days
  Replication factor: 3, min.insync.replicas: 2

Producer: Creator Service on S3 upload complete
  Event: { episode_id, show_id, s3_key, format, chapters[] }

Consumer groups:
  1. transcode-pipeline: normalize (-16 LUFS), transcode to MP3, AAC, and Opus, and generate waveform JSON
  2. rss-regenerator: rebuild RSS XML, purge CDN cache, and ping aggregator webhooks
  3. subscriber-notifier: push new episode notifications to subscribed listeners
  4. stt-pipeline: Whisper transcription and Elasticsearch full-text indexing

Topic: playback-events (play, seek, complete) for ClickHouse analytics and IAB metrics
Topic: ad-events (impression, 50%, complete) for monetization reporting
Topic: download-events for compliance filters in dynamic ad insertion

Sync path: upload ACK and episode status set to processing, with asynchronous publishing
DLQ: episode-uploaded-dlq after 3 retries

API Design

Upload Episode

HTTP
POST /api/v1/shows/{show_id}/episodes
{
  "title": "Episode 42: System Design",
  "description": "In this episode...",
  "audio_file_key": "uploads/ep-42-raw.wav",
  "publish_at": "2025-03-14T08:00:00Z",
  "season": 3,
  "episode_number": 42,
  "explicit": false,
  "chapters": [
    {"title": "Introduction", "start": 0},
    {"title": "Main Topic", "start": 180},
    {"title": "Interview", "start": 1200}
  ],
  "ad_breaks": [
    {"position": 0, "type": "pre_roll", "max_duration": 30},
    {"position": 1200, "type": "mid_roll", "max_duration": 60}
  ]
}

Stream Episode

HTTP
GET /api/v1/episodes/{episode_id}/stream?format=aac&quality=128k
Response: 302 Redirect
Location: https://cdn.example.com/audio/ep-uuid/aac_128k.m4a

Or with ad insertion:
Location: https://cdn.example.com/dai/ep-uuid/playlist.m3u8

Get User's Subscription Feed

HTTP
GET /api/v1/feed?limit=20&cursor={last}
Response: 200 OK
{
  "episodes": [
    {
      "episode_id": "ep-uuid",
      "show": {"id": "show-uuid", "title": "Tech Talk", "art": "..."},
      "title": "Episode 42: System Design",
      "duration_seconds": 2700,
      "published_at": "2025-03-14T08:00:00Z",
      "progress": {"position": 1847, "percent": 68},
      "stream_url": "https://cdn.example.com/audio/ep-uuid/aac_128k.m4a",
      "download_url": "https://cdn.example.com/audio/ep-uuid/mp3_128k.mp3",
      "file_size_bytes": 32400000
    }
  ]
}

Update Playback Progress

HTTP
POST /api/v1/playback/progress
{
  "episode_id": "ep-uuid",
  "position_seconds": 1847,
  "speed": 1.5,
  "duration_seconds": 2700
}

Get RSS XML (Aggregator facing)

XML
<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:podcast="https://podcastindex.org/namespace/1.0">
  <channel>
    <title>My Podcast Show</title>
    <link>https://example.com/shows/my-podcast</link>
    <itunes:author>John Doe</itunes:author>
    <itunes:category text="Technology"/>
    <itunes:image href="https://cdn.example.com/art/show-123.jpg"/>
    
    <item>
      <title>Episode 42: System Design</title>
      <enclosure url="https://cdn.example.com/audio/ep-42.mp3" length="57000000" type="audio/mpeg"/>
      <pubDate>Fri, 14 Mar 2025 08:00:00 GMT</pubDate>
      <itunes:duration>3600</itunes:duration>
      <description>In this episode we discuss...</description>
      <podcast:chapters url="https://cdn.example.com/chapters/ep-42.json"/>
    </item>
  </channel>
</rss>

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpoint

Data Model

PostgreSQL: Core Relational Data

SQL
CREATE TABLE shows (
    show_id         UUID PRIMARY KEY,
    creator_id      UUID NOT NULL,
    title           VARCHAR(255) NOT NULL,
    description     TEXT,
    category        VARCHAR(50),
    subcategory     VARCHAR(50),
    language        CHAR(5),
    artwork_url     TEXT,
    website_url     TEXT,
    rss_feed_url    TEXT NOT NULL,          -- public feed URL (stable, permanent)
    explicit        BOOLEAN DEFAULT FALSE,
    subscriber_count INT DEFAULT 0,
    total_episodes  INT DEFAULT 0,
    status          ENUM('active', 'paused', 'archived') DEFAULT 'active',
    created_at      TIMESTAMP,
    updated_at      TIMESTAMP,
    INDEX idx_category (category, subscriber_count DESC),
    INDEX idx_creator (creator_id)
);

CREATE TABLE episodes (
    episode_id      UUID PRIMARY KEY,
    show_id         UUID NOT NULL,
    title           VARCHAR(255) NOT NULL,
    description     TEXT,
    show_notes      TEXT,
    season          SMALLINT,
    episode_number  SMALLINT,
    duration_seconds INT,
    audio_url_mp3   TEXT,                  -- CDN URL for MP3
    audio_url_aac   TEXT,                  -- CDN URL for AAC
    audio_url_opus  TEXT,                  -- CDN URL for Opus
    original_s3_key TEXT,
    file_size_bytes INT,
    chapters        JSONB,
    ad_breaks       JSONB,
    transcript_url  TEXT,
    waveform_url    TEXT,
    explicit        BOOLEAN DEFAULT FALSE,
    status          ENUM('draft','processing','scheduled','published','archived'),
    published_at    TIMESTAMPTZ,
    created_at      TIMESTAMP,
    INDEX idx_show (show_id, published_at DESC),
    INDEX idx_published (status, published_at DESC)
);

CREATE TABLE subscriptions (
    user_id         UUID NOT NULL,
    show_id         UUID NOT NULL,
    subscribed_at   TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
    notifications   BOOLEAN DEFAULT TRUE,
    PRIMARY KEY (user_id, show_id),
    INDEX idx_show (show_id)               -- "who subscribes to this show"
);

Redis Key Schemas

# Playback progress
playback:{user_id}:{episode_id}  → Hash { position, speed, updated_at } (TTL: 90 days)

# User's episode queue
queue:{user_id}  → List of episode_ids (ordered)

# RSS feed cache
rss:{show_id}  → String (RSS XML blob) (TTL: 15 minutes)

# Podcast charts (Sorted Sets)
charts:top:{category}    → Sorted Set { show_id: score }
charts:trending          → Sorted Set { show_id: growth_score }

# Episode download counter (incremented in Redis, flushed daily to ClickHouse)
downloads:{episode_id}:{date}  → INT (INCR) (TTL: 2 days)

S3 Storage Layout

Bucket: podcast-originals (cross-region replicated, permanent)
  /{show_id}/{episode_id}/original.wav

Bucket: podcast-processed (CDN-served)
  /{show_id}/{episode_id}/mp3_128k.mp3
  /{show_id}/{episode_id}/aac_128k.m4a
  /{show_id}/{episode_id}/opus_48k.ogg
  /{show_id}/{episode_id}/waveform.json
  /{show_id}/{episode_id}/transcript.json

Bucket: podcast-artwork
  /{show_id}/artwork_3000x3000.jpg
  /{show_id}/artwork_600x600.jpg

Kafka Message Bus Topics

Topic: episode-published     (triggers RSS regeneration and push notifications)
Topic: playback-events       (play, seek, complete events feeding ClickHouse analytics)
Topic: download-events       (download started events used for IAB compliance filters)
Topic: ad-events             (ad impressions feeding monetization statistics)

ClickHouse: Analytics DB

SQL
CREATE TABLE episode_plays (
    episode_id      UUID,
    show_id         UUID,
    user_id         UUID,
    event_type      Enum8('play'=0,'pause'=1,'seek'=2,'complete'=3,'download'=4),
    position_seconds UInt32,
    duration_seconds UInt32,
    speed           Float32,
    platform        Enum8('ios'=0,'android'=1,'web'=2,'rss'=3),
    country         FixedString(2),
    city            String,
    event_date      Date MATERIALIZED toDate(timestamp),
    timestamp       DateTime
) ENGINE = MergeTree()
PARTITION BY toYYYYMM(timestamp)
ORDER BY (show_id, episode_id, timestamp);

Fault Tolerance

ConcernSolution
Audio file corruption
  • Checksum verification after upload
  • re-upload from creator if corrupt
CDN failure
  • Multi-CDN (CloudFront + Akamai)
  • DNS failover in < 30 seconds
RSS feed stale
  • Max TTL 15 min
  • manual cache purge on publish
  • ETag for conditional requests
Playback sync lossClient buffers progress locally, then retries sync when online
Ad insertion failureServe episode without ads (degrade gracefully, which is better than no audio)
Processing pipeline failure
  • Retry 3 times from Kafka
  • route persistent failures to DLQ and alert creator
Download counter lossRedis AOF plus batch flush to ClickHouse every hour, with ClickHouse as the source of truth

Specific: RSS Polling Storm (Thundering Herd)

Aggregators sync feeds simultaneously on the hour, triggering a massive thundering herd request spike of 55K req/sec.

  • CDN Edge Caching: Feeds are cached globally on CDN edges with a 15-minute TTL. Only 5% of requests hit the origin.
  • Conditional Requests: Aggregators support ETag and If-None-Match headers. 90% of requests return 304 Not Modified, saving massive bandwidth.
  • WebSub PubSubHubbub Push: Pushes new episode announcements to aggregators in real-time webhook endpoints instead of regular polling, eliminating 99% of requests.

Specific: Download Counting Accuracy (IAB Standard)

Advertisers pay per 1000 downloads (CPM), making overcounting (fraud) or undercounting (lost revenue) highly sensitive issues.

IAB Podcast Measurement Guidelines:
  1. Deduplication: In Flink stream, generate key = SHA256(ip + user_agent + episode_id).
     Window of 24 hours. Ignore matches within this window.
  2. Bot filtering: Filter out automated crawlers matching the IAB bot list.
  3. Byte-range filtering:
     - Ignore byte 0-1000 requests (metadata fetching only).
     - Ignore downloads where total bytes served < 50% of episode size.

Additional Considerations

Interview Walkthrough

  • 25-minute cut

    Skip arch50 and arch75 depth unless interviewing for a staff-level role.

    • Upload, transcode MP3/AAC/Opus, and publish to CDN and RSS enclosures (5 min)
    • RSS feed design: stable URLs, ETag caching, and aggregator polling storms (6 min)
    • Playback progress synchronization across listener devices (5 min)
    • Download bandwidth calculations: 500M downloads at 50 MB yielding 25 PB/day (5 min)
    • Dynamic ad insertion architecture as an advanced staff topic (4 min)
  • Lead with the dual delivery model: RSS feeds for aggregator syndication alongside direct streaming for native apps, both backed by the same storage origin.
  • Explain CDN edge caching with 15-minute TTL and ETag with If-None-Match headers so that 90% of aggregator polls return 304 Not Modified without hitting the origin.
  • Cover WebSub push to eliminate the hourly RSS polling thundering herd that spikes origin servers to 55K requests per second.
  • Walk through multi-codec storage: MP3 in RSS for universal compatibility, and Opus or AAC for native app playback to save 50% on CDN bandwidth.
  • Describe IAB-compliant download counting: deduplication windows, bot filtering, and byte-range thresholds processed inside a real-time Flink stream.
  • Mention silence-skipping implemented as client-side precomputed bounds JSON, avoiding audio re-encoding while preserving chapter markers and ad cues.
  • Avoid the common pitfall of counting every byte-range request as a distinct download, because metadata probing inflates CPM figures and compromises advertiser trust.

Engineering Trade-offs

1. Audio Codec Choice: MP3 vs AAC vs Opus

Podcast platforms continuously balance codec compatibility against network bandwidth, contrasting universal MP3 support for RSS syndication with Opus compression for streaming efficiency.

  • MP3 128k: Universal compatibility across legacy automotive systems, older browsers, and desktop podcast clients. Mandatory for public RSS enclosure feeds, though it exhibits lower compression efficiency than modern codecs.
  • AAC 128k: Native hardware support across iOS and Android mobile platforms, achieving 30% greater bandwidth efficiency than MP3 at equivalent fidelity.
  • Opus 48k: Industry-leading speech compression. Opus at 48 kbps matches the perceptual quality of 128 kbps MP3 while reducing CDN egress costs by more than 50%. While ideal for proprietary mobile apps, it is not yet universally supported by third-party RSS aggregators.

Strategy: Encode and store all three formats. Reference MP3 inside the public RSS XML feed, and configure native mobile players to negotiate optimal support: Opus first, falling back to AAC, and finally MP3.

2. Silence Detection and Skip

Automatically skipping silent gaps to keep podcasts fast and engaging.

Detection: Analyze RMS amplitude in 50ms windows. If RMS < threshold for > 500ms, mark as silence.
Adaptive threshold: Measure noise floor from first 2 seconds. Set threshold = noise_floor + 6 dB.

1. Client-Side Skip (Chosen):
   - Precompute silence bounds JSON: [(start1, end1), (start2, end2)].
   - Client player seeks past bounds.
   - Extremely flexible, zero extra storage or re-encoding costs.

2. Server-Side Skip:
   - Trim silences and save a condensed MP3.
   - Saves CDN bandwidth, but breaks chapter markers and dynamic ad boundaries.

3. Podcast Discovery: Charts and Recommendations

For comprehensive ranking system designs, multi-stage candidate generation, and vector retrieval, explore the Video Recommendation Engine.

Charts Ranking Formula:
  score = w1 x new_subscribers_7d + w2 x downloads_7d + w3 x listener_retention + w4 x growth_velocity

  - High velocity weighting allows rising trending podcasts to break into charts.
  - Computed daily via Spark and cached in Redis Sorted Sets: charts:top:{category}.

Recommendations:
  - Collaborative Filtering: User-show subscription matrices factored for matching shows.
  - Content-based: Generate semantic embeddings from title, description, and STT transcripts.
  - Serve recommendations via cosine similarity rankings, cached in Redis.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...