System Design Problem

Design a Content Moderation System

Commonly Asked By:MetaByteDanceGoogleTwitter

Interview Setup

Interview Prompt

Design a content moderation system processing 500M text posts, 100M images, and 10M videos per day, with ML classifiers routing ~1% (5M items) to 15K human reviewers.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Pre-publish blocking or post-publish async review?
  • Text can block in 50ms
  • video at 120/sec needs async with takedown, requiring a different architecture.
Fail-open or fail-close when moderation service is down?
  • Fail-close blocks new posts (safe)
  • fail-open for trusted users avoids platform freeze.
Regional policy differences (EU hate speech vs US First Amendment)?
  • Policy engine must version rules per jurisdiction
  • hot-reload every 60s.
Appeal flow in scope? What's the SLA for human review?
  • 5M human review cases/day needs priority queue
  • appeals add 20% volume.

Scope

In scope

  • Multi-modal detection (text, image, video)
  • ML classifier + human review queue
  • Appeal flow
  • False positive handling
  • Latency vs accuracy trade-off
  • Capacity estimation with shown math

Out of scope (state explicitly)

  • Detailed frontend/UI pixel implementation
  • Org structure, staffing, and hiring plan

Functional Requirements

Start by confirming which content modalities are in scope, including text, image, video, and audio. Ask whether pre-publish blocking or post-publish review applies to each, and whether regional policy differences matter, connecting to our broader Fraud Detection System patterns for abusive actors.

  • Multi-modal moderation: Moderate text, images, video, and audio content
  • Real-time scoring: Score content for violations before or immediately after publishing
  • Policy engine: Configurable rules per content type, region, and community standards
  • Violation categories: Hate speech, nudity/NSFW, violence/gore, spam, misinformation, harassment, copyright, CSAM
  • Action framework: Auto-remove (high confidence), auto-flag for review (medium), allow (low risk)
  • Human review queue: Prioritized queue for flagged content with analyst tooling, backed by the asynchronous streaming patterns from Message Queues Fundamentals
  • Appeals: Users can appeal moderation decisions
  • User reporting: Users report content; reports feed into moderation pipeline
  • Audit trail: Every moderation decision logged with reason, model version, reviewer
  • Feedback loop: Reviewer decisions retrain ML models

Non-Functional Requirements

Your interviewer will care most about recall on harmful content and cost-efficient tiered filtering. CSAM and hate speech demand high recall; only ~1% of volume should reach human reviewers.

  • Low Latency: Pre-publish moderation in < 500 ms (text); < 5 seconds (image); < 30 seconds (video)
  • High Recall: Catch > 99.5% of CSAM, > 95% of hate speech
  • Acceptable Precision: False positive rate < 5%
  • Scale: 500M+ posts/day, 100M+ images/day, 10M+ videos/day
  • Availability: 99.99%
  • Regional Compliance: Different rules per country
  • Reviewer Wellbeing: Limit harmful content exposure; rotate, counsel

Capacity Estimations

Run this math before you size GPU pools. Posts, images, and videos per day tell you pipeline throughput; the human review queue at ~1% of volume drives analyst staffing.

MetricCalculationValue
Text posts / dayGiven500M
Images / dayGiven100M
Videos / dayGiven10M
Text moderation / secDerived from daily volume ÷ 86400 (+ peak factor)~6K
Image moderation / secDerived from daily volume ÷ 86400 (+ peak factor)~1.2K
Video moderation / secDerived from daily volume ÷ 86400 (+ peak factor)~120
Human review cases / day~1% of 610M total5M (~1% of total)
Human reviewersGiven~15K

Architecture Diagram

In the room: pre-publish vs post-publish differs by modality; text can block in 200ms while video needs async frame sampling.

Walk your interviewer through the tiered cascade first. Content enters modality-specific pipelines as soon as a user publishes: text runs synchronously on the publish path (< 100 ms), while images and video publish first and review asynchronously. Each pipeline applies a tiered filter: cheap rules first, GPU classifiers second, and human reviewers last, ensuring cost scales with uncertainty rather than raw volume.

A shared policy engine applies regional thresholds before any action fires. At 610M items/day with only ~1% routed to 15K human reviewers, the diagram below shows how auto-remove, auto-flag, and allow decisions branch from a single scoring API without blocking the entire platform on ML latency.

Loading...

Component Deep Dives

Next we walk through each component on the diagram. Moderation is not a single monolithic model; it is a cascade of increasingly thorough checks gated by confidence scores. Text blocks publish; images and video go live first and get reviewed within SLA. The pipelines below share a policy engine and audit trail but route to separate GPU pools and human queues because each modality has a different latency budget and failure mode.

Text Moderation Pipeline

Text is the highest-volume path (~6K/sec) and the only one that blocks publish. Layers 1 and 2 handle 90% of traffic in under 50 ms; Layer 3 LLM analysis runs only on borderline scores (0.3 to 0.7) where context matters and a false positive would silence legitimate speech.

Layer 1: Keyword/Regex Filter (< 1ms)
  Bloom filter + Aho-Corasick for known bad terms
  Catches: obvious slurs, known spam phrases

Layer 2: ML Text Classifier (< 50ms)
  Fine-tuned BERT/DistilBERT
  Multi-label: P(hate_speech), P(harassment), P(spam), P(violence)
  Batch of 32 texts on GPU in ~50ms = 1.5ms per text

Layer 3: LLM-based Analysis (borderline cases only, < 2s)
  Only if Layer 2 score is 0.3-0.7
  Expensive: ~$0.01 per text -> only 10% of traffic

Layer 4: Context Enrichment
  User's history, conversation context, community norms

Image Moderation Pipeline

Images demand a hash-first strategy: PhotoDNA matches against known CSAM databases in under 10 ms and trigger immediate removal plus NCMEC reporting without consuming any ML latency budget. Only after the hash check passes does the CV classifier score nudity, violence, and hate symbols; OCR catches text embedded in memes that bypass the text pipeline entirely.

Step 1: Hash matching - PhotoDNA / pHash (< 10ms)
  Compare against known illegal content database (CSAM, terrorism)
  Match found -> IMMEDIATE removal + report to NCMEC

Step 2: ML Classification (< 200ms)
  EfficientNet/ResNet-50 fine-tuned on moderation data
  P(nudity), P(violence), P(hate_symbol), P(drugs), P(gore)
  If P(minor) > 0.5 AND P(nudity) > 0.5 -> escalate to CSAM team

Step 3: OCR + Text Moderation (for text in images)

Step 4: Object Detection (YOLO/Faster-RCNN)
  Weapons, flags/symbols associated with extremism

Video Moderation Pipeline

Full-frame analysis of a 10-minute video would require 18,000 image inferences, which is far too expensive at 120 videos/sec. The pipeline samples key frames every 2 seconds and runs the image classifier in parallel, similar to the ingestion and chunking techniques in a Video Transcoding Pipeline, with a prioritized scan that expands to all frames only when early samples flag risk.

Approach: Sample + Classify (not every frame)

Step 1: Extract key frames (every 2 seconds) + audio track
  10-min video -> 300 frames + audio

Step 2: Run image moderation on each key frame (parallel)
  First pass: sample 10 frames -> quick assessment
  If clean -> low-priority full scan later
  If flagged -> immediately scan all 300 frames
  Result: 90% of videos need only 10 frames analyzed

Step 3: Audio moderation
  Speech-to-text -> text pipeline; gunshots/screams detection; copyright matching

Policy Engine: Regional Compliance

ML scores are raw probabilities; the policy engine maps them to actions using jurisdiction-specific thresholds. The same hate-speech score might auto-remove in Germany but only flag for review in the US; geo-restricted removal lets content stay visible where local law permits it.

Same content may be legal in one country and illegal in another.

Implementation:
  Policy rules configurable per: country, content type, user age, community type
  On moderation: apply model scores against regional thresholds
  Geo-restricted removal: content removed in Germany but visible in US

Human Review Queue: Prioritization

Five million flagged items per day cannot be FIFO, since a viral CSAM clip and a borderline spam post would compete for the same 15K reviewers. Priority scoring weights severity, reach, and time sensitivity so the queue surfaces harm that spreads fastest, not content that arrived first.

5M flagged items/day. 15K reviewers. Not all items equal priority.

Priority scoring:
  priority = severity_weight x reach x time_sensitivity

  severity: CSAM=1000, Violence=100, Hate=50, Nudity=30, Spam=10
  reach: Viral 10x, Regular 1x, Private 0.5x
  time: Going viral 10x, Steady 1x, Old 0.5x

Queue: Redis sorted set ZADD review_queue {priority} {content_id}
Reviewer dequeues: ZPOPMAX review_queue

API Design

Score Content (Pre-Publish)

HTTP
POST /api/v1/moderation/score
{
  "content_id": "post-uuid",
  "content_type": "text",
  "text": "This is the post content...",
  "user_id": "user-uuid",
  "region": "US",
  "community_id": "comm-uuid"
}
Response:
{
  "decision": "allow",
  "scores": { "hate_speech": 0.05, "spam": 0.02, "violence": 0.01 },
  "flags": [], "review_required": false
}

Report Content

HTTP
POST /api/v1/moderation/report
{ "content_id": "post-uuid", "reporter_id": "user-uuid",
  "reason": "hate_speech", "details": "Contains racial slurs" }
-> { "report_id": "rpt-uuid", "status": "received" }

Review Decision & Appeal

HTTP
POST /api/v1/moderation/review/{content_id}/decide
{ "reviewer_id": "...", "decision": "remove",
  "violation_category": "hate_speech", "notes": "..." }
-> { "decision_id": "...", "action_taken": "removed" }

POST /api/v1/moderation/appeal
{ "content_id": "post-uuid", "user_id": "user-uuid",
  "reason": "This is a news quote, not hate speech" }
-> { "appeal_id": "app-uuid", "status": "under_review" }

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpoint

Data Model

PostgreSQL: Moderation Decisions & Policies

SQL
CREATE TABLE moderation_decisions (
    decision_id UUID PRIMARY KEY, content_id UUID NOT NULL,
    content_type ENUM('text','image','video','audio') NOT NULL,
    user_id UUID NOT NULL, auto_scores JSONB,
    auto_decision ENUM('allow','review','remove'),
    final_decision ENUM('allow','remove','restrict','warning'),
    violation_category VARCHAR(50), reviewer_id UUID,
    policy_version VARCHAR(20), model_version VARCHAR(20),
    region VARCHAR(10), appealed BOOLEAN DEFAULT FALSE,
    created_at TIMESTAMPTZ DEFAULT NOW(), decided_at TIMESTAMPTZ
);

CREATE TABLE moderation_policies (
    policy_id UUID PRIMARY KEY, region VARCHAR(10) NOT NULL,
    content_type VARCHAR(20) NOT NULL,
    violation_category VARCHAR(50) NOT NULL,
    auto_remove_threshold DECIMAL(4,3) DEFAULT 0.90,
    review_threshold DECIMAL(4,3) DEFAULT 0.30,
    enabled BOOLEAN DEFAULT TRUE, effective_from TIMESTAMPTZ,
    UNIQUE (region, content_type, violation_category, effective_from)
);

Redis: Real-Time State

review_queue:{category}  -> Sorted Set { content_id: priority_score }
user_trust:{user_id}     -> FLOAT (0.0 = untrusted, 1.0 = highly trusted)
strikes:{user_id}        -> INT
mod_score:{content_hash} -> JSON (scores) TTL: 3600

Event Bus Design (Kafka)

Topic: moderation-requests
  Partitions: 128
  Partition key: content_type (text, image, video, audio pipelines)
  Retention: 7 days

Topic: moderation-decisions
  Partition key: content_id

Consumer groups (moderation-requests):
  1. text-pipeline: Bloom filter -> BERT -> LLM borderline scoring
  2. image-pipeline: PhotoDNA (mandatory) -> EfficientNet -> OCR -> text pipeline
  3. video-pipeline: key frame extract -> image pipeline per frame + Whisper STT

Consumer groups (moderation-decisions):
  1. user-notification: removal reason + appeal link
  2. model-feedback: reviewer decisions -> weekly ML retrain
  3. strike-service: progressive enforcement (warn -> ban)
  4. analytics: ClickHouse precision/recall, review time, model drift

Sync path: content hidden pending review; auto-allow < 0.3 score immediately
DLQ: moderation-requests-dlq; CSAM hits bypass queue -> immediate auto-remove

Fault Tolerance

ConcernSolution
Moderation service down
  • Fail-close for new accounts
  • fail-open for trusted users
ML model degradation
  • Monitor precision/recall daily
  • auto-rollback if drops > 5%
Review queue backlog
  • Auto-adjust thresholds
  • hire surge reviewers
False positive spikeCircuit breaker: switch all decisions to 'review'
Hash database unavailable
  • Continue with ML-only
  • maintain local cache + replicas
Regional policy update
  • Version policies with effective dates
  • hot-reload every 60s

Handling Viral Harmful Content

Scenario: Terrorist attack live-streamed. Video going viral in real-time.

Response playbook:
  1. CSAM/terrorism hash match -> immediate auto-removal
  2. Hash all known copies (perceptual hash survives re-encoding)
  3. Block ALL uploads that match hash within 10 seconds
  4. ML model: flag visually similar content
  5. Keyword filters: block titles/descriptions referencing the event
  6. Rate limit new account uploads
  7. Human review: "war room" mode

Prevention: shared industry hash databases (GIFCT, PhotoDNA/NCMEC)

Additional Considerations

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • Multi-modal pipeline with different latency per type (5 min)
    • Tiered filtering: Bloom + regex → ML classifier (6 min)
    • Image path: PhotoDNA hash match for CSAM first (5 min)
    • Three-way action: auto-remove, review queue, allow (5 min)
    • Human review queue prioritized by severity x reach (4 min)
  • Frame as a multi-modal pipeline with different latency budgets: text (< 500 ms), image (< 5 s), video (sampled key frames, not every frame).
  • Describe tiered filtering: Bloom filter + regex to ML classifier (BERT) to LLM for borderline cases only, ensuring cost scales with uncertainty rather than volume.
  • Image path: PhotoDNA hash match for CSAM first (immediate removal + NCMEC report), then CV classification for nudity/violence/gore.
  • Three-way action framework: auto-remove (high confidence), auto-flag for human review (medium), allow (low risk), with thresholds tailored per violation category.
  • Prioritize the human review queue with severity x reach x time_sensitivity in a Redis sorted set, because 5M flagged items/day cannot be processed FIFO.
  • Hybrid pre-publish vs post-publish routing: strict for new/low-trust users, async for established accounts with high trust scores.
  • Common pitfall: using the same ML threshold for CSAM and spam; CSAM demands maximum recall at any precision cost, whereas spam tolerates false positives.

Engineering Trade-offs

Pre-Publish vs Post-Publish Moderation

Moderation trades user experience against safety, balancing pre-publish blocking against post-publish takedown latency.

Pre-publish: moderate BEFORE content is visible
  ✓ Harmful content never seen
  ✗ Adds latency (500ms text, 5s image, 30s video)
  Use for: high-risk content types

Post-publish: publish immediately, moderate async
  ✓ Zero latency
  ✗ Harmful content visible for seconds-to-minutes
  Use for: low-risk users with trust score > 0.8

Hybrid (recommended):
  New users/low trust: pre-publish (stricter)
  Established users/high trust: post-publish (faster)

Model Accuracy vs Coverage (Precision-Recall)

Threshold = 0.95: Precision 99%, Recall 70%
Threshold = 0.70: Precision 85%, Recall 95%

Different thresholds per category:
  CSAM: threshold 0.50 (maximize recall at all costs)
  Spam: threshold 0.80 (false positives acceptable)
  Hate speech: threshold 0.90 (context matters)
  Nudity: threshold 0.85

Principle: the higher the harm of a false negative, the lower the threshold.

Cost of Moderation at Scale

ML inference costs (per day):
  Text: 500M x $0.0001 = $50K
  Images: 100M x $0.001 = $100K
  Videos: 10M x $0.01 = $100K
  LLM (borderline): 50M x $0.01 = $500K
  Total ML: ~$750K/day = ~$274M/year

Human review: 15K reviewers x $15/hr x 8hr = $1.8M/day

Optimized cost: ~$150-200M/year (vs $900M+ without optimization)
  Trust-based routing, efficient models, tiered approach, hash matching

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...