Interview Setup
Interview Prompt
Design Instagram, a photo-sharing app with a personalized feed, ephemeral Stories, and content discovery.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| For the home feed, do we fan-out on write, read, or a hybrid? How do celebrity accounts (illustrative threshold: ~50K+ followers) change the answer? | This is the central feed trade-off. A celebrity post that fans out on write becomes millions of Redis writes, which means you need the hybrid model before anything else scales. |
| What is the Stories lifecycle and do expired stories need hard deletion or is lazy expiry OK? | Stories have 500M uploads/day with a 24-hour logical lifetime. Cassandra row TTL (86400s) and expires_at validation govern application visibility, while S3 lifecycle rules asynchronously delete media files. |
| How do we store photos, including original plus resized variants? What is the acceptable upload-to-visible latency? | At 100M uploads/day (3 MB originals, four resized versions), you're looking at 600 TB/day (~220 PB/year raw annual ingest). The answer determines S3 layout, CDN warming, and the two-phase publish gate. |
| What scale are we designing for in terms of DAU, feed reads, and storage growth? | Anchors the math: 500M DAU, 5B feed reads/day (~58K/sec), 100M photo uploads/day. Without these numbers every component choice is guesswork. |
Scope
In scope
Out of scope (state explicitly)
- Full ML ranking model training pipeline
- Direct messaging / chat (acknowledged as a core product requirement, but out of scope for this architecture, see Real-Time Chat)
- Ad insertion and monetization
Functional Requirements
Clarify core functional boundaries across photo uploads, personalized feeds, ephemeral stories, and user relationship graphs. Confirm whether real-time chat and interest discovery reside in scope.
- Upload photos and videos with captions, tags, geographic locations, and image filters.
- News feed: Personalized feed displaying posts published by followed creators.
- Stories: Ephemeral multimedia content that expires automatically after 24 hours.
- Follow and unfollow: Social graph relationship management.
- Like, comment, and save: Real-time interactions and post bookmarks.
- Explore page: Content discovery recommendations based on viewer interest profiles.
- Direct messaging (DMs): Acknowledged as a core product requirement, but out of scope for this system design exercise (see Real-Time Chat).
- User profiles: User timeline grid, biographical details, follower counts, and following counts.
- Hashtags and location search: Topic-based and geographic content discovery.
- Notifications: Push and in-app alerts for likes, comments, follows, and user mentions.
Non-Functional Requirements
Photo upload latency and feed read speed matter most. Interviewers will frequently ask whether fan-out starts before transcoding finishes, since CDN and object storage choices follow directly from media payload sizes.
- High Availability: 99.99% uptime target across core feed and upload endpoints.
- Low Latency: Feed loads in under 200 ms (bounded candidate retrieval and ML ranking), while image thumbnails render in under 500 ms at the CDN edge.
- Scalability: Architecture supports 2B+ registered accounts and 500M+ daily active users.
- Durability: Uploaded media must achieve 11 nines durability with zero unrecoverable asset loss across ~220 PB/year raw annual ingest.
- Read-Heavy Profile: 100:1 read to write ratio between feed browsing and post creation.
- Eventual Consistency: Eventual consistency is acceptable for feed ordering, like counts, and follower totals.
- Global Reach: Multi-region CDN edge POPs deliver low-latency asset caching worldwide.
Capacity Estimations
Storage math for media ingestion and feed fan-out memory requirements define the foundational boundaries of the system architecture.
| Metric | Calculation | Value |
|---|---|---|
| DAU | Given (product assumption) | 500M |
| Photo uploads / day | 500M DAU x 0.2 | 100M |
| Avg photo size (original) | Typical camera payload assumption | 3 MB |
| Original photo storage / day | 100M x 3 MB | 300 TB/day |
| Original photo storage / year | 300 TB/day x 365 | ~110 PB/year |
| Resized versions per photo | 4 variants (thumbnail, small, medium, large) | 3 MB total resizes |
| Total media storage / day | 300 TB originals + 300 TB resizes | 600 TB/day |
| Total media ingest / year | 600 TB/day x 365 (raw annual ingest) | ~220 PB/year |
| Feed reads / day | 500M DAU x 10 (read-heavy, ~58K reads/sec) | 5B |
| Stories uploads / day | 500M DAU x 1 (24-hour ephemeral lifecycle) | 500M |
Architecture Diagram
In the interview, clearly separate media upload from feed fan-out, and never push a post to followers until all media variants are ready.
Walk two distinct pipelines in parallel: the media upload pipeline, where the client transfers binaries to S3 using a pre-signed URL followed by asynchronous transcoding triggered by a post-created event, and the feed delivery pipeline, where the completion of media processing atomically updates the database to status = published and emits a post-published event. Downstream fan-out workers consume post-published to populate Redis candidate timelines for CDN-served thumbnails. These two pipelines converge only after publication, ensuring fan-out never triggers on a post still in processing.
Normal accounts below the illustrative ~50K follower threshold receive fan-out on write into each follower's Redis sorted-set timeline, scored by creation timestamp for bounded candidate retrieval. High-follower creators skip push fan-out because followers merge celebrity posts dynamically at read time, following the hybrid pattern explored in News Feed System and Twitter Timeline. Stories follow an independent path with a 24-hour logical lifetime governed by Cassandra TTL and expires_at validation, asynchronous S3 lifecycle deletion, isolated CDN caching, and no fan-out to the main feed.
Kafka decouples ingestion and delivers events to independent, parallel consumer groups: Media Processing (consuming post-created), Fan-Out (consuming post-published), Search Indexing, Notifications, and Analytics. None of these consumers depend on Fan-Out finishing first.
Component Deep Dives
Media handling and feed fan-out represent decoupled architectural domains. The design establishes a strict media readiness gate before timeline distribution begins, then walks through hybrid feed aggregation, ephemeral story caching, and discovery pipelines.
Post Service
The Post Service manages photo and video intake, persists initial metadata, orchestrates downstream processing, and authorizes deletion workflows without proxying large media payloads through application servers.
- Handles photo and video uploads along with post creation and deletion lifecycle events.
- Upload Flow:
- The client requests a pre-signed S3 upload URL from Post Service.
- The client uploads media binaries directly to S3, preventing large payloads from congesting application servers.
- The client sends post metadata including captions, tags, and location to Post Service.
- Post Service persists metadata to Cassandra with
status = processing. - Post Service publishes a
post-createdevent to Kafka for the Media Processing Pipeline.
- Post Deletion Flow:
- When a user requests post deletion, Post Service verifies account ownership and authorization.
- Post Service marks the post record as deleted or tombstones it in Cassandra to prevent read hydration.
- Post Service emits a
post-deletedevent to Kafka. - Background workers consume
post-deletedto asynchronously remove the post ID from follower Redis timelines and delete the post document from Elasticsearch. - An edge invalidation worker triggers an asynchronous CDN cache purge for the post's media URL patterns. Because CDN edge purge is asynchronous and takes time to propagate globally, logical deletion correctness relies on read-time tombstone and authorization checks that immediately suppress the post and block media hydration during cache propagation.
Media Processing Pipeline
The Media Processing Pipeline executes asynchronous transcoding, variant generation, and safety filters before content is eligible for distribution.
- Triggered by: The
post-createdKafka event emitted after initial metadata persistence. - Processing steps:
- Resize: Generate 4 standardized versions (150x150 thumbnail, 320x320 small, 640x640 medium, and 1080x1080 large).
- Apply filters: If the user selected an image filter, apply the transformation server-side or validate client-side pre-processing.
- Generate blurhash: Create a compact string hash representation for progressive placeholder rendering.
- Extract EXIF: Read GPS coordinates and camera metadata, and strip sensitive privacy data.
- Content moderation: Run automated machine learning models for NSFW and violence detection.
- Video processing: Transcode videos into multiple adaptive bitrate streams (360p, 720p, 1080p) using HLS or DASH and extract video thumbnails.
- Publishing transition: When all variants and moderation checks pass, atomically update the post record to
status = publishedin Cassandra, warm CDN edge caches, and emit apost-publishedevent to Kafka via transactional outbox or CDC.
- Technology: AWS Lambda for parallel image transformations, FFmpeg for video encoding, and dedicated GPU clusters for computer vision moderation.
- Output: Multiple resized assets stored back in S3 with corresponding CDN URLs generated.
Fan-Out Service
The Fan-Out Service consumes published post events and distributes chronological timeline pointers into Redis caches according to creator audience size.
- Consumes from Kafka
post_eventstopic specifically forpost-publishedevents (neverpost-created). Fan-out is triggered exclusively by this event after media processing is complete. - Normal users (< 50K followers): Using an illustrative ~50K threshold (tuned dynamically in production based on worker capacity and queue lag), the worker fetches the author's follower list from the MySQL social graph and executes
ZADDforpost_idinto each follower'sfeed:{user_id}Redis sorted set (scored bycreated_attimestamp), maintaining a rolling cap of 500 entries. - Celebrities (≥ 50K followers): The worker skips push fan-out entirely. Posts remain in the author's
user_postsCassandra partition and are pulled dynamically at read time. - Deletion handling: On receiving a
post-deletedevent, the worker executesZREMto evictpost_idfrom follower Redis timelines asynchronously. - Consumer idempotency: Redis
ZADDandZREMoperations are naturally idempotent under Kafka at-least-once delivery, ensuring duplicate event replays cause no timeline corruption.
Feed Service (Hybrid Fan-Out)
Feed Service coordinates bounded candidate retrieval, celebrity timeline pulls, read-time authorization filtering, and dynamic ML ranking before hydrating the final timeline response.
- Adopts the hybrid approach established in News Feed System and Twitter Timeline.
- Read-Time Assembly Pipeline:
- Bounded Candidate Retrieval: Fetches pre-computed candidate post IDs from the viewer's Redis sorted set
feed:{user_id}(scored by creation timestamp). Candidate retrieval is bounded to at most 500 entries, so latency remains effectively constant with respect to total feed history. - Celebrity Pull: Queries the
user_poststable for the latest ~50 posts from each followed celebrity (accounts at or above the illustrative ~50K threshold). - Visibility and Authorization Filtering: Evaluates candidate posts against active social graph permissions, dropping stale pointers from recent unfollows, account blocks, post deletions, and private accounts. Stale Redis pointers are harmless because they never bypass this authorization gate.
- ML Ranking at Read Time: Passes the combined candidate pool (~500 posts) to the ML Ranking Service, which scores and orders posts dynamically based on viewer affinity and engagement velocity.
- Hydration and Response: Hydrates the top 50 ranked post IDs with full post metadata, author profiles, and CDN media URLs, returning a cursor-paginated response to the client.
- Bounded Candidate Retrieval: Fetches pre-computed candidate post IDs from the viewer's Redis sorted set
- Feed ranking signals: The machine learning model re-ranks candidate posts using multiple signals:
- Relationship strength, measuring interaction frequency and bi-directional user-author interaction history between viewer and author.
- Post recency, applying exponential decay to older content.
- Engagement velocity, tracking like and comment acceleration within the initial 30-minute window.
- Content format affinity, evaluating preferences across static photos, video clips, and carousels.
- Historical viewer behavior and topical engagement clusters.
Story Service
Story Service manages ephemeral multimedia sharing with automatic 24-hour product expiration enforced through database TTLs, application-level timestamp validation, and asynchronous object lifecycle policies.
- Ephemeral content lifecycle: Stories maintain a strict 24-hour product visibility window. Cassandra enforces primary row expiry with
default_time_to_live = 86400, while story records carry an explicitexpires_attimestamp. - Read-time validation: Story Service and feed requests evaluate
expires_aton every read, immediately rejecting expired content regardless of storage or cache propagation states. - Storage management: Media assets reside in dedicated S3 prefixes where S3 lifecycle policies make objects eligible for asynchronous deletion after 24 hours.
- CDN delivery and authorization: CDN media URLs require non-expired signed access, ensuring edge caching never serves logically expired content after the 24-hour window.
- Pre-fetching: Story thumbnails and initial frames for followed creators are pre-fetched by the client application on launch to eliminate playback buffering.
- Story tray cache: Redis sorted sets (
story_tray:{user_id}) store pre-computed lists of followed creator IDs with active stories ordered by latest publication timestamp. The tray is asynchronously updated when stories are published or expire, and is periodically rebuilt as a recovery mechanism. Story Service acts as the authoritative enforcement point by validatingexpires_aton every read, while mobile clients may also defensively hide expired items.
Explore Service
Explore Service surfaces discovery content from accounts the user does not currently follow by combining graph clustering, topical embeddings, and engagement velocity.
- Discovery objective: Introduce relevant content outside the follower graph to drive discovery and creator growth.
- Candidate generation:
- Collaborative filtering identifies posts liked by users with similar historical consumption profiles.
- Content-based filtering pairs posts that share hashtags, topics, or location clusters with user affinity vectors.
- Engagement velocity surfaces rapidly trending posts exhibiting anomalous like and comment acceleration.
- Architecture implementation:
- Offline pipelines run Apache Spark jobs to cluster user embeddings and pre-generate candidate pools per interest category.
- Online ranking models score hundreds of candidate posts against real-time user context within a strict 100-millisecond latency budget.
- Ranked Explore grids are cached in Redis with short expiration windows to balance personalization freshness with computation cost.
Social Service (Follow, Like, Comment)
Social Service handles core engagement interactions by coordinating atomic cache operations with asynchronous durability pipelines.
- Follow operations: Persists edge relationships to MySQL as the source of truth for the social graph, refreshes Redis follower sets, and publishes a
follow-eventto Kafka for downstream timeline invalidation. - Like interactions: Executes an atomic check-and-set in Redis using
SADD liked:{post_id} {user_id}, increments the counter withINCR like_count:{post_id}, and emits a Kafka event for asynchronous Cassandra persistence and notification triggers. - Comments: Writes comment payloads directly to Cassandra using
post_idas the partition key andcomment_idas the clustering key for chronological ordering, while publishing an event to Kafka for the Notification Service. - Deduplication: The Redis set for
liked:{post_id}naturally prevents duplicate likes, ensuring that retried network requests remain idempotent. - Anti-spam protection: Applies rate limiting capped at 20 comments per minute per user alongside machine learning spam classifiers that analyze comment text prior to ingestion.
Notification Service
Channel adapters isolate provider-specific retry logic and fan out push notifications across external mobile gateways and in-app activity feeds. For deep architectural mechanics, review the Notification System.
- Consumes notification events from Kafka topics including
post-published,like-events,follow-events,comment-events, andmention-events. - Delivery channels: Dispatches mobile push notifications through Apple Push Notification service (APNs) and Firebase Cloud Messaging (FCM), while updating an internal inbox for the in-app Activity tab.
- Consumer idempotency: Deduplicates push notifications using Redis
SETNXonevent_idor time-windowedpost_id:user_idkeys to prevent duplicate pushes under Kafka at-least-once delivery. - Aggregation and batching: Collapses multiple engagement events on the same post over a short duration into a single summary notification, such as stating that multiple users liked your photo rather than dispatching dozens of individual pushes.
- Activity feed storage: Stores notification timelines in Cassandra partitioned by
user_idand ordered by event timestamp to power the user activity view. - Rate limiting: Enforces a ceiling of at most one push notification per post within a 5-minute window for identical event types to prevent notification fatigue.
Search Service (Elasticsearch)
Search Service powers typeahead queries and exploratory discovery across user accounts, hashtags, and geographic locations using distributed inverted indexes.
- Indexed entities: Indexes user profiles (username, display name, and bio), hashtags (post count and real-time velocity metrics), and geographic locations (place name and coordinates).
- Query features: Supports edge n-gram prefix autocomplete for usernames, fuzzy spell correction, and trending hashtag lookups.
- Consumer idempotency: Elasticsearch consumers ingest
post-publishedevents to index hashtags and locations usingpost_idas the document ID for idempotent upserts, and delete the corresponding document upon receivingpost-deletedevents. - Ranking logic: User searches prioritize verified badges, follower volume, and mutual connections. Hashtag searches prioritize total post count combined with recent trending velocity.
- Data synchronization: Changes to profile metadata flow through Kafka change data capture (CDC) pipelines, allowing Elasticsearch consumer workers to update index documents asynchronously.
- Architectural constraint: Post captions are deliberately excluded from full-text search indexing because discovery on Instagram relies primarily on hashtags, Explore recommendations, and location tags rather than open text retrieval.
ML Ranking Service
The ML Ranking Service runs on the read serving path to score merged candidate posts retrieved from Redis timeline caches and celebrity pulls within low latency bounds.
- Scores approximately 500 candidate post IDs sourced from the user's Redis feed cache (scored by creation timestamp) combined with real-time celebrity pulls.
- Read-time scoring: Batch-scores candidate posts with ONNX runtime on CPU within a strict 100-millisecond latency budget before passing the top 50 to hydration.
- Model features: Evaluates relationship strength, post recency, engagement velocity, content format (photo, video, or carousel), and historical topic preferences.
- Service consumers: Serves both Feed Service for personalized home feeds and Explore Service for discovery candidate scoring.
- Fallback resilience: Automatically falls back to a simple reverse-chronological sort if the ranking cluster experiences downtime or latency spikes.
Analytics Pipeline
Analytics runs on a separate service level objective from the hot serving path, aggregating telemetry and engagement events without placing load on transactional databases.
- Consumes all Kafka event streams into Apache Flink for real-time windowed aggregations before sinking data into ClickHouse.
- Idempotency: Flink windowing deduplicates incoming events by
event_idbefore database insertion. - Computed metrics: Tracks post engagement rates, story completion percentages, follower growth curves, and creator analytics dashboards.
- Downstream consumers: Supplies feature pipelines for Explore ranking models, enriches content moderation heuristics, and powers advertising attribution.
Event Bus Design (Kafka)
The event bus decouples producers from consumers, buffers traffic spikes, and coordinates independent parallel consumer groups across asynchronous processing stages.
topics:
post_events: "post-created (triggers media pipeline), post-published (triggers fan-out, search indexing, and notifications), or post-deleted (triggers cache and index eviction)"
story_events: "story-created or story-expired"
like_events: "like or unlike interactions on posts"
follow_events: "follow, unfollow, or block events for relationship graph updates"
comment_events: "new comment or comment deletion on posts"
notification_events: "batched mobile push alerts and in-app activity triggers"
post_events_configuration:
partitions: 128
partition_key: "author_id (guarantees strict per-creator ordering across Kafka partitions)"
retention_period: "7 days"
replication_factor: 3
min_insync_replicas: 2
producer_settings:
idempotence: true # enable.idempotence=true — prevents duplicate Kafka records caused by producer retries
payload_schema:
event_id: "UUID"
event_type: "post-created | post-published | post-deleted"
post_id: "int64"
author_id: "UUID"
status: "processing | published | deleted"
media_urls: "string[]"
created_at: "timestamp"
consumer_groups:
media_pipeline:
consumes: "post-created"
action: "Resize variants, generate blurhash, and run moderation, then on success, atomically mark published in DB and emit post-published"
fan_out_worker:
consumes: "post-published"
action: "ZADD post_id to follower Redis feeds for accounts below the illustrative ~50K celebrity threshold"
search_indexer:
consumes: "post-published and post-deleted"
action: "Upsert or delete hashtags and locations in Elasticsearch inverted indexes"
notification_worker:
consumes: "post-published, like_events, and comment_events"
action: "Dispatches batched mobile push notifications and activity feed updates"
analytics_pipeline:
consumes: "All post_events, like_events, and story_events"
action: "Streams events to Apache Flink for windowed aggregations and sinks to ClickHouse"
consumer_idempotency:
delivery_guarantee: "At-least-once delivery, requiring all consumers to implement idempotent execution"
fan_out_workers: "Redis ZADD is inherently idempotent on duplicate event replay"
search_indexer: "Elasticsearch writes use post_id as document ID for idempotent upserts"
notification_worker: "Deduplicates alerts using Redis SETNX on event_id or post_id:user_id time windows"
analytics_pipeline: "Flink windowing deduplicates incoming events by event_id before database insertion"
pipeline_paths:
sync_upload_path: "Client uploads to S3 -> Post Service persists post(status=processing) -> emit post-created -> return 201 Created"
async_publishing_path: "Media pipeline finishes -> atomically update status=published in DB -> emit post-published event via transactional outbox/CDC -> fan-out, search, and notification consumers process in parallel"
deletion_path: "Authorize delete -> canonical tombstone/delete in DB -> emit post-deleted event -> async workers remove feed pointers, delete search index docs, and trigger async CDN purge while read-time checks suppress content during propagation"
dead_letter_queue: "Route to post-events-dlq after 3 retries, and alert when consumer lag exceeds 60s"API Design
RESTful API contracts and client interface definitions for media upload authorization, feed ingestion, ephemeral stories, and real-time social engagement.
Client API Type Definitions
TypeScript domain interfaces defining upload parameters, feed payloads, story trays, and client service methods:
export type MediaFilter = "normal" | "clarendon" | "gingham" | "juno" | "lark";
export type MediaType = "image" | "video";
export interface PresignedUploadRequest {
mediaType: MediaType;
contentType: string;
byteSize: number;
}
export interface PresignedUploadResponse {
uploadUrl: string;
mediaId: string;
expiresInSeconds: number;
}
export interface GeoLocation {
latitude: number;
longitude: number;
name: string;
}
export interface CreatePostRequest {
mediaIds: string[];
caption: string;
location?: GeoLocation;
taggedUserIds?: string[];
filter?: MediaFilter;
}
export interface PostResponse {
postId: string;
authorId: string;
status: "processing" | "published";
createdAt: string;
}
export interface FeedPost {
postId: string;
author: {
userId: string;
username: string;
avatarUrl: string;
isVerified: boolean;
};
caption: string;
mediaUrls: {
thumbnail: string;
small: string;
medium: string;
large: string;
};
blurhash: string;
likeCount: number;
commentCount: number;
hasLiked: boolean;
createdAt: string;
}
export interface GetFeedResponse {
posts: FeedPost[];
nextCursor: string | null;
hasMore: boolean;
}
export interface StoryItem {
storyId: string;
mediaUrl: string;
mediaType: MediaType;
createdAt: string;
expiresAt: string;
}
export interface StoryTrayItem {
user: {
userId: string;
username: string;
avatarUrl: string;
};
stories: StoryItem[];
allSeen: boolean;
}
export interface InstagramClient {
getUploadUrl(request: PresignedUploadRequest): Promise<PresignedUploadResponse>;
createPost(request: CreatePostRequest): Promise<PostResponse>;
getFeed(limit?: number, cursor?: string): Promise<GetFeedResponse>;
getStoriesTray(): Promise<{ storyTrays: StoryTrayItem[] }>;
createStory(mediaId: string, stickers?: string[]): Promise<{ storyId: string; expiresAt: string }>;
likePost(postId: string): Promise<{ success: boolean; likeCount: number }>;
addComment(postId: string, text: string): Promise<{ commentId: string; createdAt: string }>;
}Upload Photo
Dispatches finalized post metadata, referenced media IDs, and filter options after S3 binary upload completes:
POST /api/v1/posts HTTP/1.1
Host: api.instagram.com
Authorization: Bearer <user_token>
Content-Type: application/json
{
"media_ids": ["media-uuid-1"],
"caption": "Beautiful sunset! #nature",
"location": { "lat": 37.7749, "lng": -122.4194, "name": "San Francisco" },
"tagged_users": ["user-456"],
"filter": "clarendon"
}
HTTP/1.1 201 Created
Content-Type: application/json
{
"post_id": "719284910284719203",
"status": "processing",
"created_at": "2026-03-15T18:42:00Z"
}Get Pre-signed Upload URL
Authorizes direct binary transfer from the client to object storage, bypassing application servers:
GET /api/v1/media/upload-url?type=image&content_type=image%2Fjpeg HTTP/1.1
Host: api.instagram.com
Authorization: Bearer <user_token>
HTTP/1.1 200 OK
Content-Type: application/json
{
"upload_url": "https://s3.amazonaws.com/instagram-media/media-uuid-1?AWSAccessKeyId=AKIAIOSFODNN7EXAMPLE&Signature=vjbyPxybdZaNmGa%2ByT272YEAiv4%3D&Expires=1742064120",
"media_id": "media-uuid-1",
"expires_in": 900
}Get Feed
Retrieves a paginated list of ranked timeline posts using an opaque cursor:
GET /api/v1/feed?cursor=719284910284719200&limit=10 HTTP/1.1
Host: api.instagram.com
Authorization: Bearer <user_token>
HTTP/1.1 200 OK
Content-Type: application/json
{
"posts": [
{
"post_id": "719284910284719200",
"author": {
"user_id": "usr_948271",
"username": "photographer_jane",
"avatar_url": "https://cdn.instagram.com/avatars/usr_948271.jpg",
"is_verified": true
},
"caption": "Golden gate reflections at dusk.",
"media_urls": {
"thumbnail": "https://cdn.instagram.com/media/usr_948271/2026/03/m1/thumb_150.jpg",
"small": "https://cdn.instagram.com/media/usr_948271/2026/03/m1/small_320.jpg",
"medium": "https://cdn.instagram.com/media/usr_948271/2026/03/m1/medium_640.jpg",
"large": "https://cdn.instagram.com/media/usr_948271/2026/03/m1/large_1080.jpg"
},
"blurhash": "LEHLk~WB2yk8pyo0adR*.7kCMdnj",
"like_count": 1420,
"comment_count": 89,
"has_liked": false,
"created_at": "2026-03-15T18:30:00Z"
}
],
"next_cursor": "719284910284719185",
"has_more": true
}Get Stories
Fetches active story trays for followed accounts within the active 24-hour expiration window:
GET /api/v1/stories/feed HTTP/1.1
Host: api.instagram.com
Authorization: Bearer <user_token>
HTTP/1.1 200 OK
Content-Type: application/json
{
"story_trays": [
{
"user": {
"user_id": "usr_102938",
"username": "travel_dan",
"avatar_url": "https://cdn.instagram.com/avatars/usr_102938.jpg"
},
"stories": [
{
"story_id": "st_9482019283",
"media_url": "https://cdn.instagram.com/stories/st_9482019283.jpg",
"created_at": "2026-03-15T14:10:00Z",
"expires_at": "2026-03-16T14:10:00Z"
}
]
}
]
}Post Story
Publishes an ephemeral multimedia asset with interactive stickers and audio attachments:
POST /api/v1/stories HTTP/1.1
Host: api.instagram.com
Authorization: Bearer <user_token>
Content-Type: application/json
{
"media_id": "media-uuid-story-1",
"stickers": ["location_tag_sf", "poll_question_1"],
"music_id": "track_98124"
}
HTTP/1.1 201 Created
Content-Type: application/json
{
"story_id": "st_9482019283",
"created_at": "2026-03-15T18:45:00Z",
"expires_at": "2026-03-16T18:45:00Z"
}Like and Comment
Executes atomic engagement updates against published post records:
POST /api/v1/posts/719284910284719200/like HTTP/1.1
Host: api.instagram.com
Authorization: Bearer <user_token>
HTTP/1.1 200 OK
Content-Type: application/json
{ "success": true, "like_count": 1421 }
POST /api/v1/posts/719284910284719200/comments HTTP/1.1
Host: api.instagram.com
Authorization: Bearer <user_token>
Content-Type: application/json
{ "text": "Amazing photo!" }
HTTP/1.1 201 Created
Content-Type: application/json
{ "comment_id": "cm_82910482", "created_at": "2026-03-15T18:46:12Z" }Common Error Responses
Standardized HTTP error envelopes across media validation, rate limiting, and authorization gates:
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
Photos reside in object storage, while post metadata, timelines, and social graph relationships reside across dedicated NoSQL and relational data stores.
Cassandra: Posts
Primary table for post metadata, status flags, media variant references, and engagement counter columns:
CREATE TABLE posts (
post_id BIGINT, -- Snowflake ID
user_id UUID,
caption TEXT,
media_urls LIST<TEXT>, -- CDN URLs for different sizes
location TEXT,
hashtags SET<TEXT>,
tagged_users SET<UUID>,
status VARCHAR, -- 'processing', 'published', 'deleted'
like_count COUNTER,
comment_count COUNTER,
created_at TIMESTAMP,
PRIMARY KEY (post_id)
);
-- User's own posts (profile grid and celebrity pull source)
CREATE TABLE user_posts (
user_id UUID,
post_id BIGINT,
media_thumb TEXT,
created_at TIMESTAMP,
PRIMARY KEY (user_id, post_id)
) WITH CLUSTERING ORDER BY (post_id DESC);Cassandra: Stories (with TTL)
Ephemeral story table utilizing automatic row expiration through native Cassandra TTL configuration alongside application-level expires_at validation:
CREATE TABLE stories (
user_id UUID,
story_id BIGINT,
media_url TEXT,
media_type VARCHAR, -- image, video
created_at TIMESTAMP,
expires_at TIMESTAMP, -- application-level expiry check
PRIMARY KEY (user_id, story_id)
) WITH CLUSTERING ORDER BY (story_id DESC)
AND default_time_to_live = 86400;Redis: Feed Cache
In-memory sorted sets maintaining bounded chronological candidate references per user (scored by creation timestamp), with ML ranking scores computed dynamically at read time:
# Redis Feed Cache Schema & CLI Operations
# Key: feed:{user_id}
# Type: Sorted Set (ZSET)
# Member: post_id
# Score: created_at timestamp
# Purpose: Bounded candidate retrieval with ML ranking computed dynamically at read time
# Max: 500 entries
# Add post pointer to user timeline (scored by creation timestamp)
ZADD feed:usr_8921 1742064120 719284910284719203
# Retrieve bounded chronological candidate post IDs for read-time ranking
ZREVRANGE feed:usr_8921 0 499
# Maintain sliding window cap of 500 posts
ZREMRANGEBYRANK feed:usr_8921 0 -501
# Asynchronous eviction on post deletion or unfollow
ZREM feed:usr_8921 719284910284719203Redis: Story Tray
Pre-computed sorted set indexing followed accounts with active stories ordered by recency, where read queries filter against expires_at:
# Redis Story Tray Schema & CLI Operations
# Key: story_tray:{user_id}
# Type: Sorted Set (ZSET)
# Member: poster_user_id
# Score: latest_story_timestamp
# Lifecycle: Asynchronously updated on story publish or expiry, and periodically rebuilt as a recovery mechanism
# Add or update poster in follower story tray when a story is published
ZADD story_tray:usr_8921 1742064120 usr_102938
EXPIRE story_tray:usr_8921 3600
# Read active story creators ordered by latest timestamp (Story Service authoritatively validates expires_at)
ZREVRANGE story_tray:usr_8921 0 -1 WITHSCORESMySQL: Users and Social Graph
ACID-compliant relational tables managing user profile attributes and bidirectional follow relationships:
CREATE TABLE users (
user_id UUID PRIMARY KEY,
username VARCHAR(30) UNIQUE,
display_name VARCHAR(64),
bio TEXT,
avatar_url TEXT,
post_count INT DEFAULT 0,
follower_count INT DEFAULT 0,
following_count INT DEFAULT 0,
is_private BOOLEAN DEFAULT FALSE,
created_at TIMESTAMP
);
CREATE TABLE follows (
follower_id UUID,
followee_id UUID,
status ENUM('active', 'pending'),
created_at TIMESTAMP,
PRIMARY KEY (follower_id, followee_id),
INDEX idx_followee (followee_id)
);S3: Media Storage Structure
Hierarchical prefix organization for original photo assets, processed resolutions, and progressive blurhash tokens:
bucket: "instagram-media"
prefix_structure: "/{user_id}/{year}/{month}/{media_id}/"
stored_objects:
- name: "original.jpg"
description: "Original master upload archived to Glacier after 180 days"
- name: "thumb_150.jpg"
description: "150x150 thumbnail for profile grid and search previews"
- name: "small_320.jpg"
description: "320x320 resolution for compact mobile displays"
- name: "medium_640.jpg"
description: "640x640 resolution for standard mobile feed display"
- name: "large_1080.jpg"
description: "1080x1080 resolution for high-density screens and full-screen view"
- name: "blurhash.txt"
description: "Compact string hash placeholder for instant blur preview"Fault Tolerance
Resiliency strategies for handling media loss, transcode failures, cache evictions, and upstream network interruptions.
General Failure Scenarios
Architectural mechanisms mitigating infrastructure degradation across storage, cache, and queue boundaries:
| Concern | Solution |
|---|---|
| Media loss | S3 with 11 nines durability, cross-region replication |
| Upload failure | Client retries with idempotent media_id, and uses S3 multipart upload for large files |
| Feed cache loss | Rebuild user timeline from recent posts of followed accounts in user_posts table, degrading to reverse-chronological order without ML ranking while the cache warms |
| Story expiration accuracy | Cassandra TTL (86400s) and application-level expires_at validation govern logical expiry, while S3 lifecycle asynchronously purges media files and CDN caching requires valid tokens |
| Post deletion propagation | Authorize deletion, tombstone or delete the canonical record in the database, and emit a post-deleted event. Background workers asynchronously evict Redis feed pointers, delete Elasticsearch documents, and trigger asynchronous CDN purge, while read-time tombstone and authorization checks immediately suppress deleted posts so correctness does not depend on instant cache invalidation. |
| Media processing failure | Dead-letter queue for failed processing jobs, retried up to 3 times before asking user to re-upload |
| Kafka consumer restarts | Consumers are idempotent because Kafka delivery is at-least-once (Redis ZADD/ZREM replay safe, ES upsert by ID, notification deduplication via event_id) |
Specific: Handling Image Upload Failures
Step-by-step recovery workflow for client and server errors during media ingestion:
- The client uploads to S3 using resumable multipart upload protocols.
- If the network drops mid-stream, the client retries starting from the last verified byte chunk rather than restarting the entire payload.
- If post metadata persists successfully but media processing experiences a worker crash, the post remains in the
processingstate. - An asynchronous dead-letter queue worker retries image processing jobs up to 3 times with exponential backoff.
- After 3 consecutive failures, mark the processing job as failed and notify the user to re-upload the asset.
Additional Considerations
Deep-dive operational topics covering progressive image rendering, automated safety filters, geo-discovery, privacy controls, video streaming, and race condition prevention.
Progressive Image Loading
Progressive loading provides instant visual feedback while minimizing cellular bandwidth usage on mobile clients.
- First, the client renders a blurhash placeholder string that decodes into an instant blurred preview with zero network latency.
- Next, the client fetches the lightweight 150px thumbnail to give the user a low-resolution recognizable view.
- Then, the client downloads the final resolution (320px, 640px, or 1080px) matched to the device display density and viewport dimensions.
- The web and mobile clients utilize
srcsetandsizesattributes to let the browser request the optimal image asset.
Content Moderation Pipeline
The content moderation pipeline runs automated computer vision and NLP models on every uploaded asset before fan-out begins. For a dedicated deep dive on moderation architectures, review the Content Moderation System.
stages:
stage_1_intake:
trigger: "Media binary uploaded to S3 staging prefix"
event: "Post Service persists post(status=processing) and emits post-created event to Kafka"
stage_2_automated_ml:
computer_vision: "NSFW, violence, and gore detection models score image frames"
nlp_analysis: "Optical character recognition (OCR) and caption analysis evaluate hate speech"
stage_3_decision_tree:
auto_reject:
condition: "Violation score > 0.85"
outcome: "Quarantine asset, update status to rejected in DB, and notify author without publishing"
human_review:
condition: "Violation score between 0.40 and 0.85"
outcome: "Route asset to internal human moderation queue with 15-minute SLA"
auto_publish:
condition: "Violation score < 0.40"
outcome: "Promote media to public CDN, atomically update post status to published in DB, and emit post-published event to Kafka"Hashtag and Location Pages
Dedicated aggregation pages index posts by semantic tags and geographic metadata to facilitate organic content discovery.
- Hashtag aggregation: Queries Elasticsearch for post records containing the requested hashtag token, filtered by public visibility.
- Location queries: Executes geospatial boundary queries using Elasticsearch
geo_pointfields or PostGIS indexes to locate posts created within specific geographical coordinates. - Sorting tabs: Supports both recent views sorted strictly by timestamp and top views weighted by engagement acceleration.
Private Accounts
Privacy boundaries restrict content distribution across feed caches, search indexes, and exploration surfaces.
- Follow requests require explicit creator approval, creating relationship records with
status = pending. - Posts published by private accounts are viewable only by approved followers with active relationship status.
- Feed fan-out workers distribute timeline post pointers exclusively to approved followers in the social graph.
- Feed authorization guarantee: Even if a denormalized post pointer lingers in a user's Redis feed cache following a privacy change or unfollow, the read-time visibility filter validates the active relationship against the social graph before ranking and hydration, ensuring unapproved viewers never see private posts.
- Explore ranking models, public hashtag clusters, and trending indexers explicitly filter out private account media.
Instagram Reels (Video Feed)
Reels operates as an independent short-form video feed optimized for low-latency playback and continuous scroll engagement.
- Operates as an immersive, full-screen vertical video feed powered by a dedicated recommendation pipeline.
- Transcodes ingested video into multi-bitrate HLS and DASH streams to support adaptive streaming across varying cellular networks.
- The recommendation engine evaluates completion rates, rewatch signals, and multimodal video embeddings to generate personalized video streams.
- Mobile clients pre-fetch the next 3 video segments in the feed to achieve instant, zero-buffer scrolling transitions.
Feed Ranking Model: ML Deep Dive
The feed ranking system shifts from chronological ordering to predicted user engagement through multi-task learning models that balance immediate interactions with long-term retention.
model_objective: "Multi-task neural network predicting P(like), P(comment), P(save), P(share), and P(dwell_time > 3s)"
scoring_function: "Score = w1*P(like) + w2*P(comment) + w3*P(save) + w4*P(share) + w5*P(dwell)"
weight_priority: "Saves and shares carry highest weights because they indicate stronger positive intent than likes"
feature_engineering:
user_author_affinity:
interaction_score: "Exponential decay of user likes and comments on author content"
profile_visit_count: "Frequency of explicit profile navigations"
interaction_history: "Bi-directional user-author interaction history elevates relevant author posts to top of feed"
post_signals:
recency_decay: "Exponential decay based on post age in minutes"
content_format: "Photo, video, or carousel (video receives approximately 1.3x implicit boost)"
engagement_velocity: "Ratio of likes in first 30 minutes over impressions in first 30 minutes"
metadata_completeness: "Presence of location tags, caption quality, and hashtag count"
context_signals:
temporal: "User local time of day and day of week"
session_depth: "First daily session prioritizes peak content, while fifth session explores deeper inventory"
network_connection: "Deprioritize heavy video assets on 3G or constrained cellular connections"
serving_pipeline:
candidate_generation: "Retrieve bounded candidates from Redis feed cache (scored by created_at) and pull recent celebrity posts"
visibility_filtering: "Filter out stale pointers from unfollows, blocks, deleted posts, and private accounts"
inference_stage: "Batch score ~500 merged candidates with ONNX runtime on CPU within 100ms latency budget"
diversity_injection: "Enforce constraint of no more than 2 consecutive posts from the same author"
hydration_and_output: "Hydrate top 50 ranked posts with user and media metadata, returning cursor-paginated payload"Race Condition: Post Visible Before Media Processed
A two-phase publishing gate prevents followers from loading feed entries while image variants and moderation pipelines are still completing.
race_condition_scenario:
problem: "User creates post -> metadata saved to DB -> premature fan-out triggers -> followers see post in timeline"
failure_mode: "Followers load feed before media variants are resized or moderated, causing broken image icons"
two_phase_publishing_solution:
phase_1_staging:
actions: "Upload media to the S3 staging prefix, while Post Service stores the post with status = 'processing' and emits a post-created event"
client_behavior: "Displays upload spinner while waiting for processing confirmation (typically 3 to 8 seconds)"
phase_2_release:
prerequisite: "All 4 image variants generated, CDN edge warmed, and moderation check passed"
state_transition: "Atomically update post record in database to status = 'published'"
event_emission: "Publishing pipeline emits post-published event to Kafka via transactional outbox or CDC"
fan_out_dispatch: "Fan-Out Service consumes post-published event to populate follower Redis feeds (< ~50K followers)"
guarantee: "Followers never receive timeline references to posts in 'processing' status because fan-out only consumes post-published"Related Problems and Concepts
Feed generation patterns overlap with News Feed System and Twitter Timeline for hybrid fan-out mechanics and timeline cache management. Direct messaging patterns are detailed in Real-Time Chat. Media ingestion and asynchronous transformation pipelines connect directly to Image Processing Pipeline. Ephemeral content storage patterns connect to Ephemeral Stories, while push delivery integrates with Notification System. Deepen your foundational systems knowledge in CDN and Edge Delivery, Storage Types (Block, File, Object), Caching Patterns and Invalidation, and Back-of-the-Envelope Estimation.
Interview Walkthrough
Structuring the 45-minute Instagram system design interview across a disciplined sequence of requirements, core architecture, edge cases, and operational scaling.
- 25-Minute Interview Strategy
Focus on the media ingestion pipeline and hybrid fan-out mechanics before detailing Stories TTL, Explore discovery ranking, and CDN caching tiering.
- Phase 1: Requirements, Scope, and Capacity Math (4 min)
- Phase 2: Upload Flow & Asynchronous Media Processing (6 min)
- Phase 3: Hybrid Feed Fan-Out & Read-Time Assembly (6 min)
- Phase 4: Stories Ephemeral Lifecycle & Explore Discovery (5 min)
- Phase 5: Storage Tiering, CDN Caching, and Resiliency (4 min)
- Phase 1: Requirements, Scope, and Capacity Math (4 min): Explicitly establish 500M DAU, 100M photo uploads/day, and 5B feed reads/day (~58K reads/sec). Clarify that direct messaging is out of scope for the session. Calculate 100M x 3 MB = 300 TB/day of original photos (~110 PB/year originals), scaling with 4 resized variants to 600 TB/day total ingest (~220 PB/year raw annual ingest).
- Phase 2: Upload Flow & Asynchronous Media Processing (6 min): Detail the pre-signed S3 upload path that prevents application servers from proxying large payloads. Walk through initial post creation with
status = processingand the emission of apost-createdKafka event. Explain variant generation, blurhash calculation, and automated moderation. Highlight the two-phase publishing gate: only after all variants pass moderation does the system atomically update the database tostatus = publishedand emit apost-publishedevent, ensuring followers never encounter broken image links. - Phase 3: Hybrid Feed Fan-Out & Read-Time Assembly (6 min): Walk through the hybrid fan-out model using an illustrative ~50K follower threshold (tuned dynamically based on worker lag). For normal users, workers consume
post-publishedto push post IDs to Redis sorted set timelines scored by creation timestamp (capped at 500 entries). For celebrities, bypass push fan-out and rely on read-time pulls fromuser_posts. Detail the 5-step read pipeline: bounded candidate retrieval from Redis, celebrity pull-merge, read-time visibility filtering (dropping stale pointers from unfollows, blocks, deleted posts, or private accounts), ML ranking of ~500 candidates within a 100ms latency budget, and final hydration. - Phase 4: Stories Ephemeral Lifecycle & Explore Discovery (5 min): Detail the 24-hour story lifecycle: logical expiration via Cassandra
default_time_to_live = 86400and authoritative Story Serviceexpires_atvalidation, asynchronous S3 lifecycle deletion, and Redis story tray caching (asynchronously updated and periodically rebuilt). Explain that Stories never fan out to the main feed. Introduce the Explore discovery pipeline: offline Spark candidate clustering, real-time ML ranking based on engagement velocity and computer vision embeddings, and diversity filtering. - Phase 5: Storage Tiering, CDN Caching, and Resiliency (4 min): Present storage tiering math (hot CDN edge for 80% of serves, S3 Standard for 30 to 180 days, S3 Glacier for originals older than 180 days saving 70%). Walk through the complete post deletion lifecycle (authorization, canonical tombstone in DB,
post-deletedevent, async Redis ZREM, Elasticsearch document deletion, and asynchronous CDN purge, emphasizing that read-time tombstone checks ensure deletion correctness does not depend on instant edge cache invalidation). Explain feed cache recovery fromuser_posts, dead-letter queue retries for failed media processing jobs, and consumer idempotency under Kafka at-least-once delivery.
Engineering Trade-offs
Core architectural trade-offs between fan-out models, media storage tiering strategies, and discovery ranking algorithms.
Fan-Out on Write vs Fan-Out on Read for Feed Generation
The central feed generation trade-off balances write amplification against read-time aggregation latency by partitioning accounts around an audience threshold.
fan_out_on_write_push_model:
mechanism: "When a creator posts, immediately write post_id to every follower's Redis feed cache (scored by created_at)"
advantages:
- "Candidate retrieval is bounded to at most 500 entries, so latency remains effectively constant with respect to total feed history"
- "Consistent timeline order across devices and rapid in-memory candidate reads"
disadvantages:
- "Severe write amplification: an account with 100M followers causes 100M Redis writes per post"
- "High-follower accounts create thundering herd spikes on worker queues"
- "Wasted compute and memory for followers who remain inactive"
fan_out_on_read_pull_model:
mechanism: "When a viewer opens their feed, dynamically query the latest posts from all followed accounts"
advantages:
- "Zero write amplification at publication time"
- "Always returns fresh content without maintaining persistent caches"
disadvantages:
- "Read amplification: 500 followed creators require 500 database lookups per feed load"
- "Unacceptable p99 latency caused by parallel scatter-gather network overhead"
instagram_hybrid_approach:
regular_users:
condition: "Audience < 50,000 followers (illustrative threshold tuned dynamically by fan-out cost and lag)"
strategy: "Fan-out on write: async workers write post_id to Redis feed:{follower_id} sorted set (max 500 entries, scored by created_at)"
overhead: "Manageable write cost for typical social graphs (200 to 1,000 followers)"
celebrity_accounts:
condition: "Audience >= 50,000 followers (illustrative threshold)"
strategy: "Fan-out on read: bypass push fan-out so posts remain in the creator's user_posts table and are pulled at read time"
read_time_assembly:
bounded_candidate_retrieval: "Fetch bounded candidates from Redis feed:{user_id} (regular accounts, scored by created_at)"
celebrity_timeline: "Fetch latest ~50 posts from each followed celebrity"
visibility_and_auth_filter: "Filter out stale pointers from unfollows, blocks, deleted posts, or private accounts"
merge_and_rank: "Merge candidates (~500 posts), execute ML ranking at read time, hydrate post metadata, and return top 50"Photo Storage Optimization: CDN Tiering and Image Compression
Managing petabyte-scale media ingestion requires hierarchical storage tiering combined with modern image compression formats to curtail CDN egress expenses.
scale_metrics:
daily_serves: "100B+ photos/day across multiple resolutions"
primary_cost_drivers: "Storage volume and CDN bandwidth"
resolution_variants:
thumbnail:
dimensions: "150x150"
size: "2 to 3 KB"
purpose: "Feed grid display and search previews"
low_res:
dimensions: "480x480"
size: "20 to 40 KB"
purpose: "Mobile feed timeline rendering"
standard:
dimensions: "1080x1080"
size: "100 to 200 KB"
purpose: "Full-screen detailed view"
original:
dimensions: "Unmodified source"
size: "Up to 10 MB"
purpose: "Master asset archived in object storage (never served directly)"
storage_tiers:
hot_tier:
layer: "CDN edge POPs"
content: "Thumbnails and low-res variants for posts created in the last 30 days"
traffic_share: "~80% of all image read requests"
benefit: "Dramatically reduced origin server egress load during peak popularity"
warm_tier:
layer: "S3 Standard storage"
content: "Standard and low-res variants for posts aged 30 to 180 days"
access_model: "Fetched on demand by CDN and cached on first request"
cold_tier:
layer: "S3 Glacier"
content: "Original upload files and old variants for content older than 180 days"
retrieval_sla: "1 to 12 hours asynchronous retrieval delay"
cost_advantage: "70% cheaper per gigabyte than S3 Standard"
compression_strategy:
format_migration: "JPEG to WebP achieves 30% to 40% reduction in byte size at equivalent visual quality"
lossy_settings: "quality=85 for thumbnails, quality=92 for standard 1080px images"
progressive_rendering: "Progressive encoding enables instant low-fidelity rendering during stream download"
aggregate_savings: "40% reduction in overall CDN egress expenses"
cdn_hit_rates:
thumbnails: "~95% cache hit rate"
standard_resolution: "~80% cache hit rate"
high_profile_impact: "Celebrity uploads exceeding 1B views are overwhelmingly served from edge caches, minimizing S3 origin egress"Explore Page: Content Discovery via Graph Signals
The Explore discovery engine surfaces high-affinity content from creators outside the user's direct social graph by blending interest clusters, computer vision signals, and engagement velocity.
discovery_objective:
target_inventory: "Content from accounts the user does NOT currently follow"
optimization_goal: "Maximize long-term engagement while introducing users to new creators"
ranking_signals:
interest_graph:
importance: "Primary signal"
mechanism: "Collaborative filtering clusters users with shared interaction patterns"
example: "Users who interact with wildlife photography also engage with specific outdoor creators"
content_understanding:
mechanism: "Computer vision classifiers categorize images into topics (food, travel, architecture)"
personalization: "Matches visual tags against user historical topic affinity vectors"
social_graph_proximity:
mechanism: "Second-degree social graph endorsements"
signal_strength: "Posts saved or liked by followed connections indicate strong community relevance"
trending_velocity:
mechanism: "Detects anomalous like and comment acceleration within the trailing 1-hour window"
redis_operation: "ZADD trending:explore <velocity_score> <post_id>"
candidate_generation_and_ranking:
step_1_candidate_pool: "Retrieve ~10,000 candidates from interest graph, social proximity, and trending queues"
step_2_ml_scoring: "Multi-task model predicts P(like), P(save), and P(follow)"
step_3_diversity_filter: "Enforce strict ceiling of at most 1 post per creator within the top 50 results"
step_4_brand_safety: "Filter content against safety heuristics before serving non-personalized slots"
step_5_novelty_adjustment: "Apply weight bonus to relevant creators the user has never viewed"
architectural_distinction_from_home_feed:
home_feed: "Optimizes for relationship satisfaction with known creators"
explore_feed: "Optimizes for network expansion and organic creator discovery"
novelty_bias: "Explore utilizes elevated novelty weights to prevent echo chambers"Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.