Interview Setup
Interview Prompt
Design a Gmail-scale email service for 1B users handling 3.5M inbound emails/sec, 100K search queries/sec, and 15 EB of mailbox storage with spam filtering and threading.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Strong consistency on send (email must not be lost) or eventual for reads? | The send path requires strict durability guarantees, whereas inbox reads tolerate seconds of replication lag. |
| How is mailbox data sharded: by user_id or geographic region? | Storing 15 GB for 1B users totals 15 EB, making single-cluster sharding impossible. |
| Spam filtering inline (delay delivery) or async (deliver then move to spam)? |
|
| Search over full email body or metadata only? | Sustaining 100K search QPS across 15 EB of text requires dedicated inverted indexes partitioned by user. |
Scope
In scope
- SMTP relay
- Mailbox storage
- Search indexing
- Spam filtering (Bayesian + ML)
- Threading
- Attachment storage
Out of scope (state explicitly)
- Detailed frontend/UI pixel implementation
- Org structure, staffing, and hiring plan
Functional Requirements
Ask your interviewer which Gmail features are in scope, because while send/receive and inbox are table stakes, search, threading, and spam filtering each introduce major subsystems. Confirm attachment size limits and whether external SMTP interoperability is required.
- Send and receive emails via standard SMTP protocol with attachments up to 25 MB
- Folder and label organization: Inbox, Sent, Drafts, Spam, Trash, and user-defined labels
- Full-text search across subjects, message bodies, sender identities, and attachment contents
- Conversation threading: Group related replies into a cohesive chronological thread
- Spam and phishing defense combining heuristic rules and deep learning models
- Real-time push notifications alerting mobile and web clients of incoming messages
- Rich text composition supporting HTML formatting and embedded inline media
- Contact management with predictive recipient autocomplete
- Mailbox automation rules: Auto-label, archive, forward, and discard incoming traffic
- Calendar synchronization: Parse meeting invites and handle RSVP responses
Non-Functional Requirements
Durability is the headline non-functional requirement: once you accept an email for delivery, it must never disappear. Call out the read-heavy ratio (inbox loads vs sends) early, justifying why metadata indexes are kept separate from object storage.
- High Availability: 99.99% uptime because email is mission-critical infrastructure
- Zero Data Loss: 11 nines durability once a message is acknowledged at the SMTP gateway
- Low Delivery Latency: End-to-end delivery within 5 seconds for internal transfers
- Fast Search Queries: Sub-500ms p99 full-text search across millions of personal messages
- Massive Scalability: Support 1B+ active users and over 100B stored email records
- Strict Security: Enforce TLS in transit, encryption at rest, and automated phishing protection
Capacity Estimations
Work through storage and ingest rate before picking a database, because at 1B users and 100B+ emails, metadata sharding and blob tiering matter far more than micro-optimizing an isolated query.
| Metric | Calculation | Value |
|---|---|---|
| Users | Given (assumption documented in value) | 1B |
| Emails sent / day | 300B ÷ 86400 | 300B (50% spam) |
| Emails received per user / day | 50 ÷ 86400 | 50 (after spam filtering) |
| Avg email size | Given (typical workload assumption) | 50 KB (body) + 200 KB avg attachment |
| Storage per user | Given (assumption documented in value) | 15 GB |
| Total storage | 1B x 15 GB | 15 EB |
| Emails / sec (inbound) | From Emails / day ÷ 86400 (+ peak factor in value) | 3.5M |
| Search queries / sec | From Search queries / day ÷ 86400 (+ peak factor in value) | 100K |
Architecture Diagram
In the interview room, emphasize that message bodies live in object storage while metadata lives in a sharded index, ensuring inbox loads never scan full message text across billions of rows.
The system decouples into an outbound transmission pipeline and an inbound ingest pipeline connected by message queues. Outbound messages stream from client APIs into Kafka queues before SMTP workers execute DNS MX resolution and DKIM signing. Inbound messages pass through an MX gateway, undergo multi-layer spam filtering, and split into metadata in wide-column storage and payloads in object storage.
Component Deep Dives
This section explores the complete mail lifecycle across the send path, receive path, spam classification, and push notifications.
Email Send Flow
Outbound delivery validates recipient domains, writes attachments to object storage, signs cryptographically via DKIM, and coordinates SMTP handshakes with exponential backoff.
1. User clicks "Send" -> POST /api/messages/send 2. Validate: recipients exist, attachment size < 25MB, rate limit check 3. Store email body and attachments to Blob Store (S3 or GCS) 4. Store email metadata to Bigtable or Cassandra 5. Enqueue to send queue (Kafka topic: outgoing-emails) 6. SMTP Sender Worker picks up message from queue: a. DNS MX lookup: resolve recipient domain to locate receiving mail server b. Open TLS connection to receiving server using STARTTLS c. Authenticate: sign message headers with DKIM private key for sender domain d. Transmit email via standard SMTP commands e. Receiving server acknowledges receipt -> mark message as delivered f. If rejected -> generate delivery status notification (DSN bounce) to sender 7. On transient failure: retry with exponential backoff (from 1 minute up to 72 hours) After 72 hours of unsuccessful retries -> mark permanent failure and bounce to sender Optimization for internal delivery (sender@gmail to recipient@gmail): Bypass external SMTP entirely, writing directly into recipient mailbox with ~100ms latency.
Incoming Email Flow (SMTP Receive)
The inbound path applies connection-level filtering, evaluates cryptographic authentication headers, scores spam probability, and archives payloads.
1. External sender MTA establishes connection to our SMTP gateway (advertised via MX records). 2. Execute standard SMTP handshake: EHLO, MAIL FROM, and RCPT TO. 3. Pre-data validation checks: a. SPF verification: confirm whether the sending IP is authorized for the sender domain. b. Ingress rate limiting: if connection limits are exceeded, return temporary 421 code. c. Recipient mailbox validation: return 550 User Unknown if address does not exist. 4. Accept DATA payload, streaming email MIME content directly. 5. Verify DKIM cryptographic signature using sender DNS public key. 6. Evaluate DMARC alignment combining SPF and DKIM validation outcomes. 7. Execute multi-tier spam and phishing classification models: Score > 0.7 routes to Spam folder; 0.3 to 0.7 displays security banner; < 0.3 routes to Inbox. 8. Scan attachments for malicious payloads and viruses using streaming antimalware engines. 9. Persist MIME body to distributed Blob Store and write metadata to Bigtable. 10. Index searchable tokens asynchronously into Elasticsearch cluster. 11. Dispatch real-time push notification to active recipient devices. 12. Return 250 OK acknowledgment to sending MTA.
Spam Filtering Architecture
A three-layer filtering pipeline balances execution latency against classification accuracy, filtering 70% of malicious traffic at the TCP connection boundary.
Layer 1: Connection-level filters (< 1ms per connection) - IP reputation: verify against established DNS-based blocklists (Spamhaus ZEN). - Volume anomaly detection: throttle connections exceeding 1,000 emails per hour from single IP. - Forward-confirmed reverse DNS: ensure sending IP reverses to a legitimate hostname. Catches: ~70% of inbound spam volume before payload inspection. Layer 2: Content and cryptographic analysis (10 to 50ms per email) - SPF, DKIM, and DMARC verification: authenticate sender identity and domain ownership. - Heuristic rules: scan for known phishing text fragments, deceptive HTML, and obfuscated links. - URL intelligence: compare embedded links against real-time malicious domain registries. - Attachment analysis: scan binary attachments for malware signatures. Catches: additional ~20% of inbound spam. Layer 3: Machine learning classification (100 to 300ms per email) - Deep learning NLP: fine-tuned transformer model evaluating subject and body semantics. - Behavioral features: sender domain age, historical user interaction graph, and click ratios. - Confidence scoring: score > 0.85 moves to spam, 0.50 to 0.85 flags warning, < 0.50 routes to inbox. Catches: additional ~9% of sophisticated spam. Cumulative outcome: achieves a false positive rate below 0.01% on legitimate messages.
Push vs Pull Synchronization
Mobile and web clients leverage specialized push channels rather than periodic polling: IMAP IDLE maintains long-lived TCP connections for legacy desktop clients, Firebase Cloud Messaging and Apple Push Notification service deliver battery-efficient mobile wakeups, and WebSockets drive live browser inbox updates.
SPF, DKIM, and DMARC Authentication
Sender Policy Framework (SPF) publishes authorized IP ranges in DNS, DomainKeys Identified Mail (DKIM) attaches asymmetric cryptographic signatures to headers, and DMARC specifies recipient enforcement policies when authentication checks fail.
API Design
Mailbox and Message API Endpoints
Gmail-style client applications communicate over RESTful HTTPS endpoints, while standard SMTP and IMAP protocols serve external mail transfer agents. Outbound submissions enforce idempotency keys, and list operations rely on cursor pagination.
# Mailbox and message operations
GET /api/messages?label=INBOX&page_token=... -> List emails with cursor pagination
GET /api/messages/{id} -> Retrieve email metadata and snippet
POST /api/messages/send -> Submit new email for outbound delivery
PUT /api/messages/{id}/labels -> Mutate labels (add or remove)
PUT /api/messages/{id}/read -> Update read and unread status flags
DELETE /api/messages/{id} -> Soft-delete email by moving to Trash
POST /api/messages/{id}/reply -> Create in-thread reply
POST /api/messages/{id}/forward -> Forward email with attachment references
GET /api/threads/{thread_id} -> Retrieve complete conversation thread
GET /api/messages/search?q=from:alice+subject:meeting -> Search messages via inverted indexCommon Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpointData Model
Email Metadata Storage (Bigtable or Cassandra)
Applying sharding and partitioning by user ID isolates mailbox reads while wide-column clustering keys support descending chronological retrieval.
Row key: user_id:reverse_timestamp:email_id (reverse timestamp enables sequential scans of newest emails first) Columns: metadata:subject, metadata:from, metadata:to, metadata:cc, metadata:bcc, metadata:snippet, metadata:thread_id, metadata:labels, metadata:is_read, metadata:is_starred, metadata:has_attachment, refs:body_blob_id, refs:attachment_ids
Conversation Threading Architecture
Threads link related messages into continuous dialogues using RFC 5322 In-Reply-To and References headers. When headers are missing, the system groups messages sharing normalized subjects and participant sets within 7 days under a shared thread ID.
Fault Tolerance
Email Delivery Guarantees
Per RFC 5321, once an SMTP gateway returns a 250 OK code, it assumes legal responsibility to either deliver the message or generate a bounce notification. Once accepted, the message is persisted to a durable queue (Kafka, RF=3) before sending an acknowledgment. If downstream workers encounter temporary failures, they retry with exponential backoff for up to 72 hours without silently dropping data.
Data Loss Prevention
The storage architecture guarantees 11 nines of durability through triple-replicated object storage across availability zones, quorum writes in Cassandra metadata clusters, and decoupled search indexes that can be rebuilt idempotently from source blobs.
Additional Considerations
Attachment Content Deduplication
Content-addressable storage hashes attachment payloads via SHA-256, mapping duplicate attachments to a single underlying object and reclaiming up to 80% of storage in enterprise domains.
Content-addressable storage (attachment deduplication): 1. Compute cryptographic SHA-256 digest of attachment binary payload. 2. Query metadata index: if digest exists, link email to existing blob_id. 3. If digest does not exist, upload payload to Blob Store using the digest as the key. 4. Write attachment reference pointer into the email metadata row. Impact: 50,000 corporate users receiving the same 10 MB PDF store only 1 object copy in S3. Typical deduplication savings across enterprise swarms: 60% to 80% of attachment storage. Operational controls: - Reference counting: atomic counter increments and decrements track active references before deletion. - Privacy and multitenancy: payloads encrypted with envelope keys to maintain tenant isolation. - Integrity validation: verify SHA-256 hashes on read, repairing anomalies from secondary replicas.
Interview Walkthrough
- 25-minute cut
Focus on the federated SMTP model, tiered storage separation, and spam filtering.
- Federated SMTP model: anyone can run an MTA (5 min)
- Ingest: MTA through spam filter into blob storage pipeline (6 min)
- Bodies in blob storage with metadata in sharded index (5 min)
- Spam filtering as critical infrastructure (5 min)
- Attachment deduplication via content hash (4 min)
- Explain the federated SMTP model: anyone can run an MTA, meaning your system functions as one node in a global store-and-forward network rather than a closed proprietary silo.
- Separate ingest where the MTA receives, filters spam, and stores blobs and metadata from delivery where the MDA pushes to user mailboxes, acknowledging their fundamentally different scaling profiles.
- Store email bodies in blob storage; metadata (headers, labels, thread ID) in a wide-column store optimized for user-scoped queries.
- Highlight that spam filtering is critical infrastructure because more than 50% of all global email is spam, requiring multi-layer scoring (DNSBL, Bayesian, and ML models) before inbox delivery.
- Detail attachment dedup via content hash: identical files across millions of emails are stored once, saving 30% to 50% of blob storage.
- Explain that search indexes are rebuilt from metadata and blob pointers, maintaining the blob store as the single source of truth rather than the search index.
- Warn against storing full MIME bodies in the search database, which bloats indexes and degrades query latencies when message bodies belong in object storage.
Engineering Trade-offs
Email systems trade storage cost against search speed, balancing cold object storage against hot metadata indexes.
Storage Architecture: Bigtable vs Cassandra vs Sharded Relational
Scale requirement: 100B+ emails stored across 1B+ active user accounts. Option 1: Wide-Column Storage (Bigtable or Cassandra) [Preferred] Row key format: user_id:reverse_timestamp:email_id - Single-key prefix scans retrieve the user inbox in descending chronological order. - Linearly scales storage capacity across petabytes and exabytes. - Eliminates costly real-time sorting operations during inbox listing. - Limitation: requires auxiliary inverted search index (Elasticsearch) for text search. Option 2: Sharded Relational Database (MySQL or PostgreSQL) - Strict ACID guarantees and expressive join capabilities for labels and folders. - Cross-shard rebalancing and schema migrations become operational liabilities at scale. - High write amplification under massive inbound email streams. Design choice: Cassandra or Bigtable manages mailbox message rows, S3 stores MIME blobs, and Elasticsearch powers full-text queries.
Interview alignment: While simplified architectures use PostgreSQL for relational queries, at Gmail scale the mailbox hot path uses wide-column storage (Cassandra or Bigtable) for write-heavy per-user partitions, where PostgreSQL handles account metadata while Cassandra or Bigtable manages high-volume message rows.
Why Email Architecture Is Unique
- Federated protocol: anyone can run an SMTP server on the open internet without vendor lock-in.
- Store-and-forward routing: emails can relay asynchronously across intermediate MTAs with variable delays.
- Severe spam volume: over 50% of all traffic is adversarial, making multi-tier filtering mandatory.
- Four decades of RFC standards: strict backward compatibility constraints dictate MIME and header handling.
- Asynchronous delivery model: unlike instant messaging, email gracefully tolerates delivery delays from seconds to days.
- Mandatory delivery accounting: RFC regulations require explicit bounce generation whenever delivery fails permanently.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.