Interview Setup
Interview Prompt
Design a mentions and tagging system like Instagram or Twitter. Users can @mention others in posts and comments, tag people in media at specific coordinates, receive autocomplete suggestions, and get real-time notifications.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Text mentions only, or photo coordinate tags too? | Photo tags require normalized (x,y) coordinate storage and different notification copy than inline text @mentions. |
| Must blocked users be prevented from mentioning? | Privacy checks must run before database insertion and before notification event fan-out. |
| What is the autocomplete latency target? | 50K QPS at sub-50ms latency requires an in-memory prefix trie in Redis rather than database LIKE queries. |
| Are notifications synchronous or asynchronous? | Post publication must return immediately, so mention notifications must process via asynchronous Kafka consumers. |
Scope
In scope
- @mention parsing and token indexing
- Reverse index for mentions inbox
- Asynchronous notification triggering via Kafka
- Privacy and block list enforcement
- Username prefix autocomplete
- Photo coordinate tagging
Out of scope (state explicitly)
- Full ML ranking model training pipeline
- Direct messaging and chat channels
- Ad insertion and monetization
Functional Requirements
Confirm @mention parsing rules, notification delivery bounds, and mention inbox requirements. Core capabilities include text tokenization, user directory lookups, asynchronous notifications, and privacy permissions (such as controlling who can mention a user).
In the room: ask about edit-after-post semantics, specifically whether mentions get re-parsed on edit.
- Mention Users: Support @username syntax across posts, comments, and ephemeral stories.
- Media Tagging: Tag users in photos and videos at normalized coordinate positions.
- Real-Time Notifications: Dispatch push alerts to mentioned recipients within seconds.
- Mention Inbox Feed: Maintain an aggregated "Posts you are mentioned in" reverse index feed.
- Typeahead Autocomplete: Suggest valid usernames with sub-50ms latency as the user types "@".
- Mention Privacy Controls: Allow users to restrict mentions to followers or disable tagging entirely. For relationship checks, consult the Follower and Following System.
- Tag Removal: Allow recipients to detach tags from their profile without modifying the original media.
Non-Functional Requirements
Mention detection during publication must complete without delaying client responses, and notifications must deliver within seconds.
- Low Latency: Return autocomplete suggestions in under 50 ms and complete publish-time parsing in under 200 ms.
- Notification Speed: Deliver push notifications within 5 seconds under normal operational load.
- Scale: Ingest over 400M mention events daily across 200M published posts.
- Parsing Accuracy: Reliably extract handles containing periods and underscores while filtering invalid delimiters.
- Anti-Abuse Protections: Enforce rate limits on per-post and hourly mention counts to suppress mass spam.
Capacity Estimations
Posts per second and average mention density determine message queue partitioning and notification fan-out sizing.
| Metric | Calculation | Value |
|---|---|---|
| Posts with mentions / day | Given | 200M |
| Avg mentions per post | Given | 2 |
| Total mention events / day | 200M x 2 | 400M |
| Mention events / sec | Derived from daily volume ÷ 86400 (+ peak factor) | ~4.6K |
| Autocomplete queries / sec | Derived from daily volume ÷ 86400 (+ peak factor) | 50K |
| Username lookup latency | Given | < 50 ms |
Architecture Diagram
Parse @mentions at publish time on the write path, resolve against a username directory, write a reverse index for notification feeds, and enqueue async delivery via Kafka. Autocomplete operates as a separate read-heavy prefix index in Redis, rather than a full scan.
In the room: group notifications per user and post pair, because a viral thread with 50 tags should not trigger 50 separate push alerts.
The publish API stays synchronous only for parsing, validation, and index writes, while notifications, feed updates, and analytics fan out asynchronously through Kafka so authors get a sub-200ms response even when a post tags dozens of users.
Component Deep Dives
Mention Extraction and Validation
The write path must be fast and idempotent: parse mentions once at publish time, validate against permissions and block lists, persist rows in a separate mentions table, and hand off side effects to asynchronous consumers.
The client sends structured mention metadata alongside post text, where each entry includes the @username, character offset, and length so the server does not re-parse on every read. On publish, the server batch-resolves usernames to user IDs via a Redis hash, checks mention permissions (everyone, followers only, or nobody), and consults the block list before inserting rows. Invalid or blocked mentions are stored as plain text with no notification. Rate limits (20 mentions per post, 50 per hour per author) run in Redis sliding windows before any insert.
Edit-After-Publish Mention Diffing
Users frequently edit captions after publishing, requiring mention sets to diff cleanly without triggering duplicate alerts. On PATCH requests, compare parsed mention sets against stored records, notify only net-new recipients, and mark removed mentions inactive without mutating historical post body text.
User edits post text after publish:
Original: "Lunch with @alice"
Edited: "Lunch with @alice and @bob"
Diff algorithm (on PATCH /posts/{id}):
1. Parse old and new mention sets from stored offsets + fresh regex pass
2. added = new_mentions - old_mentions -> insert rows + Kafka notify
3. removed = old_mentions - new_mentions -> mark status='removed', no re-notify
4. unchanged = intersection -> skip
Idempotency: UNIQUE(content_id, mentioned_user) prevents duplicate rows on re-save.
Notification policy: notify only on newly added mentions, never on removal.
Rate limit edits: max 5 edits/post/hour to prevent mention-spam via edit loop.Mention Spam Rate Limits
Automated spam syndicates attempt to bypass feed filters by tagging high volumes of strangers. Defenses apply sliding window counters in Redis across author, target, and minute burst dimensions before insertion.
Per-author limits (Redis counters, sliding window): max 20 mentions per post max 50 mentions per hour per author max 200 mentions per day per author max 10 distinct users mentioned per minute (burst) Per-target protection: Celebrity digest mode: >100 mentions/min -> batch notifications Block list: mention stored but notification suppressed New account (<7 days): cap 30 mentions/day until trust score rises Abuse signals trigger shadow suppression: >80% mentions to non-followers Same @target spammed from >20 accounts in 1 hour Autocomplete enumeration: >100 prefix queries/min -> 429
Autocomplete Architecture
At 50K autocomplete queries per second, a database LIKE prefix scan will not survive, so the typeahead path requires sub-50ms latency. A two-tier strategy maintains responsive user interactions: the client caches the user following list from the Social Graph Store for instant suggestions among familiar connections, while the server handles global lookups via a Redis sorted set with lexicographic range queries:
ZRANGEBYLEX usernames "[al" "[al\xff" LIMIT 0 5
Results re-rank based on social proximity, mutual connections, popularity, and interaction recency. The full username index occupies approximately 20 GB of memory across 1B accounts.
Photo and Video Tagging
Photo tags differ structurally from text mentions: each tag records normalized coordinates in the 0 to 1 range so positions stay correct across thumbnail and full-resolution renders. Tags persist in a dedicated table indexed by tagged_user_idto power "photos of me" queries, dispatching alerts that read "Alice tagged you in a photo".
Event Bus Design (Kafka)
Once mention rows commit, the publish path emits a single event keyed by mentioned_user_id so all downstream consumers, including push notifications, the mentions inbox feed, and analytics, process in parallel without blocking the author. For messaging architecture patterns, review Message Queues Fundamentals.
Topic: mentions
Partitions: 64
Partition key: mentioned_user_id (routes notifications to correct shard)
Retention: 7 days
Replication factor: 3, min.insync.replicas: 2
Producer: Mention Service after validation + permission check
Event: { event_id, post_id, author_id, mentioned_user_ids[], content_type, timestamp }
Consumer groups:
1. notification: push and email dispatch to each mentioned user
2. mention-feed: compile the "tagged in" feed index per user
3. analytics: stream mention volume metrics to ClickHouse
Sync path: POST content -> store post -> validate mentions -> publish -> 201 Created.
Async path: notifications and feed indexes proceed without blocking post creation.
DLQ: mentions-dlq, alerting when notification consumer lag exceeds 30 seconds.API Design
Mentions and Autocomplete Endpoints
The API provides endpoints for publishing posts with structured mention metadata, querying user autocomplete suggestions, and fetching mention feeds.
Post with Mentions
POST /api/v1/posts
{
"text": "Great photo with @alice and @bob!",
"mentions": [
{"username": "alice", "offset": 17, "length": 6},
{"username": "bob", "offset": 28, "length": 4}
],
"media_tags": [
{"user_id": "u-alice", "x": 0.35, "y": 0.62}
]
}
Response: 201 Created { "post_id": "p-uuid" }Autocomplete
GET /api/v1/users/autocomplete?prefix=al&limit=5
Response: 200 OK
{
"suggestions": [
{"user_id": "u1", "username": "alice_smith", "name": "Alice Smith"},
{"user_id": "u2", "username": "alex_jones", "name": "Alex Jones"}
]
}Common Error Responses
Structured error representations for privacy blocks, disabled mention permissions, and rate limit violations.
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
MySQL: Mention Records
Relational tables persist content associations, character offsets, normalized media coordinates, and lifecycle statuses. Review Indexing and Query Optimization for query tuning.
CREATE TABLE mentions (
mention_id BIGINT PRIMARY KEY AUTO_INCREMENT,
content_type ENUM('post', 'comment', 'story') NOT NULL,
content_id BIGINT NOT NULL,
mentioned_user BIGINT NOT NULL,
mentioned_by BIGINT NOT NULL,
mention_type ENUM('text', 'photo_tag') DEFAULT 'text',
text_offset INT,
text_length INT,
tag_x FLOAT,
tag_y FLOAT,
status ENUM('active', 'removed') DEFAULT 'active',
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
INDEX idx_mentioned_user (mentioned_user, created_at DESC)
);Redis and Kafka Architecture
Redis maintains lexicographic autocomplete sets, privacy configurations, and unread badges, while Kafka partitions event streams by mentioned user ID.
Redis:
usernames -> Sorted Set (lexicographic autocomplete)
mention_settings:{uid} -> Hash { allow_from: "everyone" | "followers" | "nobody" }
blocked_by:{uid} -> SET of blocked user_ids
mentions_feed:{uid} -> List (TTL: 5 min)
Kafka topic: mention-events (partitioned by mentioned_user_id)Fault Tolerance
| Concern | Solution |
|---|---|
| Mention spam floods | Enforce multi-tiered rate limits: max 20 per post, 50 per hour, and 200 per day per author. |
| Deleted parent content | Cascade delete or softly mark associated mentions as inactive upon post removal. |
| User handle renames | Store permanent user_id values and resolve current usernames dynamically at render time. |
| Autocomplete index staleness | Update the Redis sorted set synchronously during account registration and handle change workflows. |
| Notification duplication | Deduplicate events by (user, post) pairs within a 5-minute buffering window. |
Additional Considerations
Notification Grouping
Buffer mentions for the same user and post pair in a 5-minute window. An initial mention triggers an immediate individual notification, while subsequent mentions collapse into aggregated copy: "Bob, Carol, and 45 others mentioned you." Buffering tracks via atomic Redis INCR counters with a 300-second TTL.
Broadcasting Mentions (@everyone and @channel)
For community or channel mentions, record a single group mention entity rather than creating thousands of individual rows. Enforce administrative privilege verification, and allow members to mute channel-level notifications.
Edit-After-Publish and Spam Controls
On caption edits, diff mention sets to notify only net-new users, capping post edits at 5 per hour. Sliding window counters in Redis prevent mention spam across accounts.
Interview Walkthrough
- 25-minute pacing strategy
Prioritize write-time parsing, reverse indexing, and asynchronous notification decoupling before addressing media tagging edge cases.
- Parse @mentions at write time with a tokenizer rather than at render time (5 min)
- Validate usernames against user directories before persistence (6 min)
- Build low-latency autocomplete with Redis sorted sets for prefix lookups (5 min)
- Fire mention notifications asynchronously via Kafka (5 min)
- Group notifications for identical user and post pairs in five-minute buffers (4 min)
- Parse @mentions at write time with a tokenizer rather than render time, storing (post_id, mentioned_user_id, char_offset) in a separate mentions table.
- Validate usernames against a user directory before persisting, rejecting or queuing mentions for deleted or renamed accounts.
- Build autocomplete with Redis sorted sets for prefix lookups (<2ms), falling back to Elasticsearch for fuzzy matching on typos.
- Fire notifications asynchronously via Kafka, never blocking the post-creation API on fan-out to all mentioned users.
- Group notifications for the same user and post pair in a 5-minute buffer, collapsing 50 mentions into "Bob, Carol, and 48 others mentioned you."
- Handle @everyone and @channel as a single group-mention record with admin permission checks rather than writing N individual mention rows.
- Render mentions client-side by overlaying clickable links at stored offsets, so username changes do not break historical post text.
- Common pitfall: storing only the raw @username string inline, because renames break historical links and deduplication becomes impossible across username changes.
Engineering Trade-offs
System Trade-offs
Mentions architectures balance synchronous vs asynchronous notification delivery, in-memory prefix tries against full-text search engines, and aggressive spam filters against friction.
Inline vs Separate Table for Mentions
A separate mentions table stores raw post text untouched and links (post_id, user_id, offset). On render, clients fetch mentions and overlay clickable anchors. Historical text remains intact across handle changes, and users who remove tags cleanly sever the association without corrupting captions.
Redis Sorted Set vs Elasticsearch for Autocomplete
Redis delivers ~2ms responses for exact prefix matches, consuming approximately 20 GB across 1B users. Elasticsearch requires ~20ms but supports fuzzy matching and rich scoring. The recommended production pattern pairs Redis for instant prefix hits with Elasticsearch fallback for typo tolerance.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.