Interview Setup
Interview Prompt
Design a content moderation system processing 500M text posts, 100M images, and 10M videos per day, with ML classifiers routing ~1% (5M items) to 15K human reviewers.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Pre-publish blocking or post-publish async review? |
|
| Fail-open or fail-close when moderation service is down? |
|
| Regional policy differences (EU hate speech vs US First Amendment)? |
|
| Appeal flow in scope? What's the SLA for human review? |
|
Scope
In scope
- Multi-modal detection (text, image, video)
- ML classifier + human review queue
- Appeal flow
- False positive handling
- Latency vs accuracy trade-off
- Capacity estimation with shown math
Out of scope (state explicitly)
- Detailed frontend/UI pixel implementation
- Org structure, staffing, and hiring plan
Functional Requirements
Start by confirming which content modalities are in scope, including text, image, video, and audio. Ask whether pre-publish blocking or post-publish review applies to each, and whether regional policy differences matter, connecting to our broader Fraud Detection System patterns for abusive actors.
- Multi-modal moderation: Moderate text, images, video, and audio content
- Real-time scoring: Score content for violations before or immediately after publishing
- Policy engine: Configurable rules per content type, region, and community standards
- Violation categories: Hate speech, nudity/NSFW, violence/gore, spam, misinformation, harassment, copyright, CSAM
- Action framework: Auto-remove (high confidence), auto-flag for review (medium), allow (low risk)
- Human review queue: Prioritized queue for flagged content with analyst tooling, backed by the asynchronous streaming patterns from Message Queues Fundamentals
- Appeals: Users can appeal moderation decisions
- User reporting: Users report content; reports feed into moderation pipeline
- Audit trail: Every moderation decision logged with reason, model version, reviewer
- Feedback loop: Reviewer decisions retrain ML models
Non-Functional Requirements
Your interviewer will care most about recall on harmful content and cost-efficient tiered filtering. CSAM and hate speech demand high recall; only ~1% of volume should reach human reviewers.
- Low Latency: Pre-publish moderation in < 500 ms (text); < 5 seconds (image); < 30 seconds (video)
- High Recall: Catch > 99.5% of CSAM, > 95% of hate speech
- Acceptable Precision: False positive rate < 5%
- Scale: 500M+ posts/day, 100M+ images/day, 10M+ videos/day
- Availability: 99.99%
- Regional Compliance: Different rules per country
- Reviewer Wellbeing: Limit harmful content exposure; rotate, counsel
Capacity Estimations
Run this math before you size GPU pools. Posts, images, and videos per day tell you pipeline throughput; the human review queue at ~1% of volume drives analyst staffing.
| Metric | Calculation | Value |
|---|---|---|
| Text posts / day | Given | 500M |
| Images / day | Given | 100M |
| Videos / day | Given | 10M |
| Text moderation / sec | Derived from daily volume ÷ 86400 (+ peak factor) | ~6K |
| Image moderation / sec | Derived from daily volume ÷ 86400 (+ peak factor) | ~1.2K |
| Video moderation / sec | Derived from daily volume ÷ 86400 (+ peak factor) | ~120 |
| Human review cases / day | ~1% of 610M total | 5M (~1% of total) |
| Human reviewers | Given | ~15K |
Architecture Diagram
In the room: pre-publish vs post-publish differs by modality; text can block in 200ms while video needs async frame sampling.
Walk your interviewer through the tiered cascade first. Content enters modality-specific pipelines as soon as a user publishes: text runs synchronously on the publish path (< 100 ms), while images and video publish first and review asynchronously. Each pipeline applies a tiered filter: cheap rules first, GPU classifiers second, and human reviewers last, ensuring cost scales with uncertainty rather than raw volume.
A shared policy engine applies regional thresholds before any action fires. At 610M items/day with only ~1% routed to 15K human reviewers, the diagram below shows how auto-remove, auto-flag, and allow decisions branch from a single scoring API without blocking the entire platform on ML latency.
Component Deep Dives
Next we walk through each component on the diagram. Moderation is not a single monolithic model; it is a cascade of increasingly thorough checks gated by confidence scores. Text blocks publish; images and video go live first and get reviewed within SLA. The pipelines below share a policy engine and audit trail but route to separate GPU pools and human queues because each modality has a different latency budget and failure mode.
Text Moderation Pipeline
Text is the highest-volume path (~6K/sec) and the only one that blocks publish. Layers 1 and 2 handle 90% of traffic in under 50 ms; Layer 3 LLM analysis runs only on borderline scores (0.3 to 0.7) where context matters and a false positive would silence legitimate speech.
Layer 1: Keyword/Regex Filter (< 1ms) Bloom filter + Aho-Corasick for known bad terms Catches: obvious slurs, known spam phrases Layer 2: ML Text Classifier (< 50ms) Fine-tuned BERT/DistilBERT Multi-label: P(hate_speech), P(harassment), P(spam), P(violence) Batch of 32 texts on GPU in ~50ms = 1.5ms per text Layer 3: LLM-based Analysis (borderline cases only, < 2s) Only if Layer 2 score is 0.3-0.7 Expensive: ~$0.01 per text -> only 10% of traffic Layer 4: Context Enrichment User's history, conversation context, community norms
Image Moderation Pipeline
Images demand a hash-first strategy: PhotoDNA matches against known CSAM databases in under 10 ms and trigger immediate removal plus NCMEC reporting without consuming any ML latency budget. Only after the hash check passes does the CV classifier score nudity, violence, and hate symbols; OCR catches text embedded in memes that bypass the text pipeline entirely.
Step 1: Hash matching - PhotoDNA / pHash (< 10ms) Compare against known illegal content database (CSAM, terrorism) Match found -> IMMEDIATE removal + report to NCMEC Step 2: ML Classification (< 200ms) EfficientNet/ResNet-50 fine-tuned on moderation data P(nudity), P(violence), P(hate_symbol), P(drugs), P(gore) If P(minor) > 0.5 AND P(nudity) > 0.5 -> escalate to CSAM team Step 3: OCR + Text Moderation (for text in images) Step 4: Object Detection (YOLO/Faster-RCNN) Weapons, flags/symbols associated with extremism
Video Moderation Pipeline
Full-frame analysis of a 10-minute video would require 18,000 image inferences, which is far too expensive at 120 videos/sec. The pipeline samples key frames every 2 seconds and runs the image classifier in parallel, similar to the ingestion and chunking techniques in a Video Transcoding Pipeline, with a prioritized scan that expands to all frames only when early samples flag risk.
Approach: Sample + Classify (not every frame) Step 1: Extract key frames (every 2 seconds) + audio track 10-min video -> 300 frames + audio Step 2: Run image moderation on each key frame (parallel) First pass: sample 10 frames -> quick assessment If clean -> low-priority full scan later If flagged -> immediately scan all 300 frames Result: 90% of videos need only 10 frames analyzed Step 3: Audio moderation Speech-to-text -> text pipeline; gunshots/screams detection; copyright matching
Policy Engine: Regional Compliance
ML scores are raw probabilities; the policy engine maps them to actions using jurisdiction-specific thresholds. The same hate-speech score might auto-remove in Germany but only flag for review in the US; geo-restricted removal lets content stay visible where local law permits it.
Same content may be legal in one country and illegal in another. Implementation: Policy rules configurable per: country, content type, user age, community type On moderation: apply model scores against regional thresholds Geo-restricted removal: content removed in Germany but visible in US
Human Review Queue: Prioritization
Five million flagged items per day cannot be FIFO, since a viral CSAM clip and a borderline spam post would compete for the same 15K reviewers. Priority scoring weights severity, reach, and time sensitivity so the queue surfaces harm that spreads fastest, not content that arrived first.
5M flagged items/day. 15K reviewers. Not all items equal priority.
Priority scoring:
priority = severity_weight x reach x time_sensitivity
severity: CSAM=1000, Violence=100, Hate=50, Nudity=30, Spam=10
reach: Viral 10x, Regular 1x, Private 0.5x
time: Going viral 10x, Steady 1x, Old 0.5x
Queue: Redis sorted set ZADD review_queue {priority} {content_id}
Reviewer dequeues: ZPOPMAX review_queueAPI Design
Score Content (Pre-Publish)
POST /api/v1/moderation/score
{
"content_id": "post-uuid",
"content_type": "text",
"text": "This is the post content...",
"user_id": "user-uuid",
"region": "US",
"community_id": "comm-uuid"
}
Response:
{
"decision": "allow",
"scores": { "hate_speech": 0.05, "spam": 0.02, "violence": 0.01 },
"flags": [], "review_required": false
}Report Content
POST /api/v1/moderation/report
{ "content_id": "post-uuid", "reporter_id": "user-uuid",
"reason": "hate_speech", "details": "Contains racial slurs" }
-> { "report_id": "rpt-uuid", "status": "received" }Review Decision & Appeal
POST /api/v1/moderation/review/{content_id}/decide
{ "reviewer_id": "...", "decision": "remove",
"violation_category": "hate_speech", "notes": "..." }
-> { "decision_id": "...", "action_taken": "removed" }
POST /api/v1/moderation/appeal
{ "content_id": "post-uuid", "user_id": "user-uuid",
"reason": "This is a news quote, not hate speech" }
-> { "appeal_id": "app-uuid", "status": "under_review" }Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpointData Model
PostgreSQL: Moderation Decisions & Policies
CREATE TABLE moderation_decisions (
decision_id UUID PRIMARY KEY, content_id UUID NOT NULL,
content_type ENUM('text','image','video','audio') NOT NULL,
user_id UUID NOT NULL, auto_scores JSONB,
auto_decision ENUM('allow','review','remove'),
final_decision ENUM('allow','remove','restrict','warning'),
violation_category VARCHAR(50), reviewer_id UUID,
policy_version VARCHAR(20), model_version VARCHAR(20),
region VARCHAR(10), appealed BOOLEAN DEFAULT FALSE,
created_at TIMESTAMPTZ DEFAULT NOW(), decided_at TIMESTAMPTZ
);
CREATE TABLE moderation_policies (
policy_id UUID PRIMARY KEY, region VARCHAR(10) NOT NULL,
content_type VARCHAR(20) NOT NULL,
violation_category VARCHAR(50) NOT NULL,
auto_remove_threshold DECIMAL(4,3) DEFAULT 0.90,
review_threshold DECIMAL(4,3) DEFAULT 0.30,
enabled BOOLEAN DEFAULT TRUE, effective_from TIMESTAMPTZ,
UNIQUE (region, content_type, violation_category, effective_from)
);Redis: Real-Time State
review_queue:{category} -> Sorted Set { content_id: priority_score }
user_trust:{user_id} -> FLOAT (0.0 = untrusted, 1.0 = highly trusted)
strikes:{user_id} -> INT
mod_score:{content_hash} -> JSON (scores) TTL: 3600Event Bus Design (Kafka)
Topic: moderation-requests Partitions: 128 Partition key: content_type (text, image, video, audio pipelines) Retention: 7 days Topic: moderation-decisions Partition key: content_id Consumer groups (moderation-requests): 1. text-pipeline: Bloom filter -> BERT -> LLM borderline scoring 2. image-pipeline: PhotoDNA (mandatory) -> EfficientNet -> OCR -> text pipeline 3. video-pipeline: key frame extract -> image pipeline per frame + Whisper STT Consumer groups (moderation-decisions): 1. user-notification: removal reason + appeal link 2. model-feedback: reviewer decisions -> weekly ML retrain 3. strike-service: progressive enforcement (warn -> ban) 4. analytics: ClickHouse precision/recall, review time, model drift Sync path: content hidden pending review; auto-allow < 0.3 score immediately DLQ: moderation-requests-dlq; CSAM hits bypass queue -> immediate auto-remove
Fault Tolerance
| Concern | Solution |
|---|---|
| Moderation service down |
|
| ML model degradation |
|
| Review queue backlog |
|
| False positive spike | Circuit breaker: switch all decisions to 'review' |
| Hash database unavailable |
|
| Regional policy update |
|
Handling Viral Harmful Content
Scenario: Terrorist attack live-streamed. Video going viral in real-time. Response playbook: 1. CSAM/terrorism hash match -> immediate auto-removal 2. Hash all known copies (perceptual hash survives re-encoding) 3. Block ALL uploads that match hash within 10 seconds 4. ML model: flag visually similar content 5. Keyword filters: block titles/descriptions referencing the event 6. Rate limit new account uploads 7. Human review: "war room" mode Prevention: shared industry hash databases (GIFCT, PhotoDNA/NCMEC)
Additional Considerations
Interview Walkthrough
- 25-minute cut
Skip arch50/arch75 depth unless staff.
- Multi-modal pipeline with different latency per type (5 min)
- Tiered filtering: Bloom + regex → ML classifier (6 min)
- Image path: PhotoDNA hash match for CSAM first (5 min)
- Three-way action: auto-remove, review queue, allow (5 min)
- Human review queue prioritized by severity x reach (4 min)
- Frame as a multi-modal pipeline with different latency budgets: text (< 500 ms), image (< 5 s), video (sampled key frames, not every frame).
- Describe tiered filtering: Bloom filter + regex to ML classifier (BERT) to LLM for borderline cases only, ensuring cost scales with uncertainty rather than volume.
- Image path: PhotoDNA hash match for CSAM first (immediate removal + NCMEC report), then CV classification for nudity/violence/gore.
- Three-way action framework: auto-remove (high confidence), auto-flag for human review (medium), allow (low risk), with thresholds tailored per violation category.
- Prioritize the human review queue with
severity x reach x time_sensitivityin a Redis sorted set, because 5M flagged items/day cannot be processed FIFO. - Hybrid pre-publish vs post-publish routing: strict for new/low-trust users, async for established accounts with high trust scores.
- Common pitfall: using the same ML threshold for CSAM and spam; CSAM demands maximum recall at any precision cost, whereas spam tolerates false positives.
Engineering Trade-offs
Pre-Publish vs Post-Publish Moderation
Moderation trades user experience against safety, balancing pre-publish blocking against post-publish takedown latency.
Pre-publish: moderate BEFORE content is visible ✓ Harmful content never seen ✗ Adds latency (500ms text, 5s image, 30s video) Use for: high-risk content types Post-publish: publish immediately, moderate async ✓ Zero latency ✗ Harmful content visible for seconds-to-minutes Use for: low-risk users with trust score > 0.8 Hybrid (recommended): New users/low trust: pre-publish (stricter) Established users/high trust: post-publish (faster)
Model Accuracy vs Coverage (Precision-Recall)
Threshold = 0.95: Precision 99%, Recall 70% Threshold = 0.70: Precision 85%, Recall 95% Different thresholds per category: CSAM: threshold 0.50 (maximize recall at all costs) Spam: threshold 0.80 (false positives acceptable) Hate speech: threshold 0.90 (context matters) Nudity: threshold 0.85 Principle: the higher the harm of a false negative, the lower the threshold.
Cost of Moderation at Scale
ML inference costs (per day): Text: 500M x $0.0001 = $50K Images: 100M x $0.001 = $100K Videos: 10M x $0.01 = $100K LLM (borderline): 50M x $0.01 = $500K Total ML: ~$750K/day = ~$274M/year Human review: 15K reviewers x $15/hr x 8hr = $1.8M/day Optimized cost: ~$150-200M/year (vs $900M+ without optimization) Trust-based routing, efficient models, tiered approach, hash matching
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.