Interview Setup
Interview Prompt
Design a notification system that delivers push notifications, SMS, and emails to users across global distributed platforms based on business events such as order updates, social interactions, and urgent security alerts. The system must support priority levels, custom user preferences, quiet hours, and end-to-end delivery tracking.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| What delivery latency SLA is expected across different notification types? | Critical security alerts require priority queues with delivery under 5 seconds, whereas bulk marketing emails can tolerate several minutes of queue delay. |
| Is delivery expected to be strictly exactly-once or at-least-once with deduplication? | External notification providers duplicate webhooks and retry deliveries, making at-least-once transport combined with idempotency keys the standard architecture. |
| Does the system need to enforce per-user daily caps and localized quiet hours? | Daily caps prevent user notification fatigue, while quiet hours ensure compliance with regulations such as TCPA and GDPR. |
| Who produces notification events: a centralized monolith or hundreds of distributed microservices? | Microservice environments require standardized Kafka event schemas, schema registries, and robust ingress fan-in validation. |
Scope
In scope
- Multi-channel delivery across push notifications, SMS, and email
- Priority scheduling and asynchronous queue decoupling
- User preferences, quiet-hours suppression, and regulatory opt-outs
- Idempotent deduplication and per-user token-bucket rate limiting
- Delivery receipt tracking, provider failover, and dead letter queues
Out of scope (state explicitly)
- In-app notification client UI component libraries
- Full WYSIWYG email marketing visual campaign builders
- Machine learning models for individual send-time optimization
Functional Requirements
Scope the core capabilities with the interviewer upfront. The system must support multi-channel delivery across push, email, and SMS, handle user preferences and scheduling, and track delivery lifecycles reliably.
- Deliver notifications across multiple channels: Push (iOS, Android, Web), Email, and SMS.
- Support both real-time delivery (under 5 seconds) and scheduled future delivery.
- Enforce user preferences, allowing users to opt into or out of specific channels per notification category.
- Support dynamic templating with localized variable substitution across languages and device formats.
- Execute bulk broadcast campaigns (such as promotional updates to 10 million users) without starving transactional alerts.
- Track end-to-end delivery lifecycle status: queued, sent, delivered, read, failed, and bounced.
- Enforce rate limits and daily caps per user to prevent notification fatigue and spam.
- Support notification bundling and digests (such as collapsing multiple social likes into a single summary).
- Enforce priority levels: critical (immediate), high, normal, and low (batched).
Non-Functional Requirements
Delivery reliability, channel isolation, and sub-second ingress latency are foundational non-functional requirements. The system must handle third-party provider downtime gracefully.
- High Throughput: Support over 1 million notifications per minute cluster-wide with peak traffic exceeding 60,000 requests per second.
- Low Latency: Deliver real-time P0 notifications to end-user devices within 1 to 2 seconds under normal network conditions.
- High Reliability: Ensure at-least-once message delivery so that no valid notification is dropped silently.
- Horizontal Scalability: Scale to 100 million daily active users generating billions of notifications daily.
- Channel Fault Tolerance: An outage in an external provider (such as an SMS carrier) must never block or degrade other channels.
- Best-Effort Ordering: Preserve chronological delivery order for sequential notifications sent to an individual user.
- Idempotency: Prevent duplicate sends when client apps or upstream services retry requests.
- Extensibility: Support seamless addition of new channels (such as WhatsApp, Slack, or webhook integrations) without modifying core ingress logic.
Capacity Estimations
Evaluate daily notification volumes and queue depth requirements to size worker pools and storage retention buffers across peak operating windows.
| Metric | Calculation | Value |
|---|---|---|
| Daily active users (DAU) | Given product assumption | 100M |
| Total notifications per day | 100M DAU x 10 per user | 1B (10 per user daily average) |
| Average notification throughput | 1B ÷ 86,400 seconds | ~11,574 / sec (peak 5x: ~60,000 / sec) |
| Push notifications volume | 60% of total traffic | 600M per day (~7,000 / sec avg) |
| Email notifications volume | 30% of total traffic | 300M per day (~3,500 / sec avg) |
| SMS notifications volume | 10% of total traffic | 100M per day (~1,150 / sec avg) |
| Average notification payload size | Rendered message record with metadata | ~500 bytes |
| Daily storage ingestion rate | 1B x 500 bytes | 500 GB per day |
| Annual storage footprint | 500 GB x 365 days | ~182.5 TB per year |
Throughput and Storage Sizing Details
- Peak Ingress Rate: 1 billion daily notifications average roughly 11,574 requests per second. Factoring in a 5x peak traffic multiplier during flash sales yields a target peak ingress of approximately 60,000 notifications per second.
- Channel Distribution: Push accounts for 600 million daily messages (~7,000/sec avg), email accounts for 300 million (~3,500/sec avg), and SMS represents 100 million (~1,150/sec avg).
- Log Storage Footprint: Storing 500-byte notification records for 1 billion daily notifications generates 500 GB of logs daily, requiring roughly 45 TB for a 90-day retention window in Cassandra.
Architecture Diagram
Interview strategy: Emphasize the synchronous and asynchronous split early. The ingress API returns a 202 Accepted status immediately after validating, checking deduplication, and enqueueing the message into Kafka. Returning 202 Accepted acknowledges durable acceptance for asynchronous processing rather than guaranteed or completed delivery to the end user.
The architecture separates synchronous request intake from asynchronous channel delivery. Internal services submit notifications through a central API layer that validates requests, filters opt-outs and quiet hours, enforces Redis deduplication and rate limits, renders channel-specific templates, and enqueues messages onto channel-specific Kafka topics before returning 202 Accepted.
Independent worker pools consume from dedicated Kafka topics to communicate with external delivery providers such as APNs, FCM, AWS SES, and Twilio. As providers emit delivery receipts, an asynchronous Delivery Tracker service streams status updates into Cassandra for notification history and ClickHouse for analytics dashboards.
Component Deep Dives
Deconstructing the notification architecture reveals how the ingress gateway, template rendering engine, Kafka event bus, and worker pools maintain isolation across channels.
Notification Service (Ingress API Gateway)
The API layer serves as the single ingress point for all internal microservices. It validates request payloads, verifies recipient preferences and quiet hours, renders localized templates, and fans out messages to channel-specific Kafka topics before returning an immediate 202 Accepted response.
- Validation: Enforces schema correctness, verifies user existence, and confirms template availability.
- Preference Filtering: Queries cached user preferences to suppress channels the user has disabled for that notification category.
- Rate Limiting: Executes token-bucket checks in Redis to enforce per-user channel thresholds (such as a maximum of 10 push notifications per hour).
- Quiet Hours Evaluation: Converts the current timestamp to the user's local timezone and defers non-critical notifications until the quiet window concludes.
- Idempotency Enforcement: Checks the
request_idagainst Redis using aSET NXcommand with a 24-hour expiration to prevent duplicate enqueueing.
Template Rendering Pipeline
After preference and quiet-hours checks pass, the API layer resolves the corresponding template_id and user locale from PostgreSQL. It substitutes dynamic parameters (such as customer name and tracking links) to produce channel-tailored payloads: push headers under 256 characters, full responsive HTML email MIME bodies, and segment-optimized SMS text.
Pre-rendering templates at the API layer allows downstream channel workers to operate as lightweight delivery agents that never execute secondary database lookups.
User Preferences and Suppression Service
User preferences and global suppression lists reside in PostgreSQL to guarantee strong consistency, ensuring marketing opt-outs take effect immediately without violating regulatory compliance standards.
- Granular Settings: Configurable per notification category (marketing, transactional, social) and per channel (push, email, SMS).
- Suppression Registry: Maintains bounced email addresses, unsubscribed phone numbers, and invalidated push tokens to prevent repetitive failed deliveries.
Kafka Message Bus and Topic Topology
Kafka isolates high-volume marketing traffic from critical alerts and buffers transient downstream provider slowdowns. Topics are partitioned by user_id to preserve strict chronological ordering for each individual recipient.
- Dedicated Channel Topics:
push_notifications,email_notifications,sms_notifications, andnotification_status. - Durable Asynchronous Decoupling: Configured with a replication factor of 3 and
min.insync.replicas=2, supporting a 7-day retention period. Kafka guarantees at-least-once message delivery to worker pools, while final handset delivery depends on provider responses and retries. - Blast Isolation: If an external email provider experiences latency, only the email topic accumulates lag while push notifications and SMS alerts proceed uninterrupted.
Push Notification Worker Pool
Push workers manage persistent HTTP/2 connections to Apple Push Notification service (APNs) and Firebase Cloud Messaging (FCM), minimizing handshake latency across high-throughput bursts.
- Connection Pooling: Reuses warm HTTP/2 multiplexed connections to avoid per-message TLS handshake overhead.
- Token Invalidation Handling: When APNs returns an HTTP 410 Gone or FCM returns a NotRegistered error, workers flag the device token as invalid in PostgreSQL to suppress future attempts.
Email Worker Pool
Email workers ingest pre-compiled MIME payloads from Kafka and dispatch them through primary and secondary delivery providers with automated circuit breakers.
- Multi-Provider Routing: Routes bulk traffic through AWS SES and transactional alerts through SendGrid, failing over automatically if error rates rise.
- Deliverability Guardrails: Enforces SPF, DKIM, and DMARC record signing while managing dedicated IP warmup schedules to safeguard domain reputation.
SMS Worker Pool
Because SMS carries significant per-segment carrier costs, SMS workers apply strict rate caps and dynamically route messages to the most cost-effective provider by destination country code.
- Regional Routing: Directs domestic North American traffic through Twilio and international deliveries through regional aggregators like Sinch or MessageBird.
- Segment Optimization: Validates character encodings (GSM-7 vs UCS-2) to avoid accidental multi-segment billing overhead.
Delivery Tracker and Telemetry Service
Delivery Tracker processes worker dispatch receipts and asynchronous provider webhooks to record lifecycle milestones in Cassandra and ClickHouse.
- State Transitions: Tracks progression across
QUEUED,SENT,DELIVERED,READ,FAILED, andBOUNCED, distinguishing between successful provider acceptance and confirmed handset delivery based on provider capabilities. - Storage Split: Writes notification records to Cassandra for fast user-facing inbox lookups and streams analytics events to ClickHouse for delivery funnel dashboards.
Event Bus Configuration Reference
The Kafka event bus schema and consumer group topology structure message flow across partitions and dead-letter queues:
Topic: push_notifications
Partitions: 128
Partition key: user_id (preserves per-user delivery order)
Retention: 7 days
Replication factor: 3, min.insync.replicas: 2
Topic: email_notifications
Partitions: 64
Partition key: user_id
Retention: 7 days
Topic: sms_notifications
Partitions: 32
Partition key: user_id
Retention: 7 days
Topic: notification_status
Partitions: 32
Partition key: notification_id
Retention: 30 days (delivery webhook audit trail)
Producer: Notification Service (executes validation, deduplication, and template rendering)
Event schema: { notification_id, request_id, user_id, channel, template_id, priority, payload, created_at }
Consumer groups (per channel pipeline):
1. push-workers: Persistent HTTP/2 connection pooling to APNs and FCM
2. email-workers: SendGrid and AWS SES with automated circuit-breaker failover
3. sms-workers: Twilio and Nexmo with regional country routing
4. status-tracker: Streams delivery receipts into Cassandra and ClickHouse
Dead Letter Queue: {channel}_notifications-dlq after 3 retries, with lag alerts above 60s
Synchronous Hot Path: Validate request -> Redis deduplication on request_id -> Enqueue to channel topics -> Return 202 Accepted
Asynchronous Delivery Path: Channel worker pools deliver to providers -> Delivery Tracker processes provider webhooksAPI Design
The system exposes RESTful HTTP endpoints for internal service integration, status querying, user preference management, and historical inbox access.
Client API Type Definitions
TypeScript domain types defining notification payloads, priority tiers, and delivery status models:
export type NotificationPriority = "critical" | "high" | "normal" | "low";
export type NotificationChannel = "push" | "email" | "sms" | "in_app";
export type DeliveryStatus = "queued" | "sent" | "delivered" | "read" | "failed" | "bounced";
export interface SendNotificationRequest {
requestId: string;
userIds: string[];
notificationType: string;
priority: NotificationPriority;
channels?: NotificationChannel[];
templateId: string;
templateVars: Record<string, string>;
scheduledAt?: string | null;
metadata?: Record<string, string>;
}
export interface ChannelDeliveryStatus {
status: DeliveryStatus;
provider: string;
sentAt?: string;
deliveredAt?: string;
failureReason?: string;
}
export interface NotificationStatusResponse {
notificationId: string;
userId: string;
channels: Record<NotificationChannel, ChannelDeliveryStatus>;
createdAt: string;
}
export interface NotificationServiceClient {
send(request: SendNotificationRequest): Promise<{ notificationId: string; status: "queued" }>;
getStatus(notificationId: string): Promise<NotificationStatusResponse>;
getUserHistory(userId: string, page: number, limit: number): Promise<NotificationStatusResponse[]>;
}Send Notification Endpoint
Primary ingress endpoint accepting notification dispatch requests with idempotency keys:
POST /api/v1/notifications
Authorization: Bearer <service_token>
Content-Type: application/json
{
"request_id": "9f8b4c2e-6d1a-4f5e-8b3c-1a2b3c4d5e6f",
"user_ids": ["user_101", "user_102"],
"notification_type": "order_shipped",
"priority": "high",
"channels": ["push", "email"],
"template_id": "order_shipped_v2",
"template_vars": {
"order_id": "ORD-12345",
"tracking_url": "https://track.example.com/ORD-12345"
},
"scheduled_at": null,
"metadata": {
"campaign_id": "spring_launch_2026"
}
}
HTTP/1.1 202 Accepted
Content-Type: application/json
{
"notification_id": "notif_7a8b9c0d-1e2f-3a4b-5c6d-7e8f9a0b1c2d",
"status": "queued",
"channels_targeted": ["push", "email"],
"queued_at": "2026-03-13T10:00:00Z"
}Get Notification Status Endpoint
Retrieves per-channel delivery timestamps and provider feedback for a specific notification:
GET /api/v1/notifications/notif_7a8b9c0d-1e2f-3a4b-5c6d-7e8f9a0b1c2d
Authorization: Bearer <service_token>
HTTP/1.1 200 OK
Content-Type: application/json
{
"notification_id": "notif_7a8b9c0d-1e2f-3a4b-5c6d-7e8f9a0b1c2d",
"user_id": "user_101",
"channels": {
"push": {
"status": "delivered",
"provider": "APNs",
"delivered_at": "2026-03-13T10:00:01.240Z"
},
"email": {
"status": "sent",
"provider": "AWS_SES",
"sent_at": "2026-03-13T10:00:01.850Z"
}
}
}Update User Preferences Endpoint
Configures channel opt-ins, quiet-hours boundaries, and notification category rules:
PUT /api/v1/users/user_101/notification-preferences
Authorization: Bearer <user_token>
Content-Type: application/json
{
"channels": {
"push": true,
"email": true,
"sms": false
},
"quiet_hours": {
"start": "22:00",
"end": "08:00",
"timezone": "America/New_York"
},
"notification_types": {
"marketing": { "push": false, "email": true },
"social": { "push": true, "email": false },
"order_updates": { "push": true, "email": true, "sms": true }
}
}
HTTP/1.1 200 OK
Content-Type: application/json
{
"user_id": "user_101",
"updated_at": "2026-03-13T10:05:00Z",
"status": "updated"
}Get User Notification History Endpoint
Paginated query returning recent notifications delivered to a user:
GET /api/v1/users/user_101/notifications?page=1&limit=20
Authorization: Bearer <user_token>
HTTP/1.1 200 OK
Content-Type: application/json
{
"notifications": [
{
"notification_id": "notif_7a8b9c0d-1e2f-3a4b-5c6d-7e8f9a0b1c2d",
"type": "order_shipped",
"title": "Your order has shipped!",
"body": "Order ORD-12345 is on its way.",
"channel": "push",
"status": "read",
"created_at": "2026-03-13T10:00:00Z"
}
],
"pagination": {
"page": 1,
"limit": 20,
"total": 150
}
}Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpointData Model
Data storage balances relational consistency for preferences and templates with high-throughput distributed stores for logs and caching.
PostgreSQL: User Preferences and Device Tokens
Relational schemas enforce strict data integrity for user notification settings, channel opt-outs, and active mobile device tokens:
CREATE TABLE user_notification_preferences (
user_id UUID PRIMARY KEY,
push_enabled BOOLEAN DEFAULT TRUE,
email_enabled BOOLEAN DEFAULT TRUE,
sms_enabled BOOLEAN DEFAULT FALSE,
quiet_hours_start TIME,
quiet_hours_end TIME,
timezone VARCHAR(64),
language VARCHAR(10) DEFAULT 'en',
updated_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP
);
CREATE TABLE user_type_preferences (
user_id UUID,
notification_type VARCHAR(64),
push_enabled BOOLEAN DEFAULT TRUE,
email_enabled BOOLEAN DEFAULT TRUE,
sms_enabled BOOLEAN DEFAULT FALSE,
PRIMARY KEY (user_id, notification_type)
);
CREATE TABLE user_device_tokens (
user_id UUID,
device_token VARCHAR(256),
platform VARCHAR(32), -- 'ios', 'android', 'web'
app_version VARCHAR(32),
is_valid BOOLEAN DEFAULT TRUE,
last_active_at TIMESTAMP WITH TIME ZONE,
created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
PRIMARY KEY (user_id, device_token)
);Cassandra: Notification Delivery Log
Cassandra provides write-optimized time-series storage partitioned by user_id with a default 90-day time-to-live expiration:
CREATE TABLE notification_log (
user_id UUID,
created_at TIMESTAMP,
notification_id UUID,
notification_type VARCHAR,
channel VARCHAR,
title TEXT,
body TEXT,
status VARCHAR, -- queued, sent, delivered, read, failed, bounced
metadata MAP<TEXT, TEXT>,
PRIMARY KEY (user_id, created_at, notification_id)
) WITH CLUSTERING ORDER BY (created_at DESC)
AND default_time_to_live = 7776000; -- 90 days retentionMySQL: Parameterized Notification Templates
Templates are stored with composite primary keys supporting versioning, language localization, and channel formatting:
CREATE TABLE notification_templates (
template_id VARCHAR(128) NOT NULL,
version INT NOT NULL,
channel VARCHAR(16) NOT NULL, -- 'push', 'email', 'sms'
language VARCHAR(10) NOT NULL, -- 'en', 'es', 'de'
subject TEXT, -- used for email subject lines
title TEXT, -- used for push headers
body_template TEXT NOT NULL, -- parameterized text: "Hello {{user_name}}, ..."
html_template TEXT, -- parameterized HTML structure for email
active BOOLEAN DEFAULT TRUE,
created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
PRIMARY KEY (template_id, version, channel, language)
);Redis Key Structure: Deduplication and Rate Limiting
In-memory data structures enforce request deduplication windows and hourly token-bucket rate limits:
# Deduplication key (expires after 24 hours)
Key: notif:dedup:{request_id}
Value: 1
TTL: 86400
# Per-user hourly rate limiter counter
Key: notif:rate:{user_id}:{channel}:{hour}
Value: integer counter (incremented via INCR)
TTL: 3600
# Active device tokens cache
Key: device_tokens:{user_id}
Value: SET of serialized {device_token, platform, app_version, last_active}Fault Tolerance
External provider outages, duplicate deliveries, and burst-traffic spikes require layered resilience mechanisms.
Resilience Strategy Matrix
| Technique | Application Mechanism |
|---|---|
| Kafka Broker Durability | Replication factor of 3 with min.insync.replicas=2 ensures notifications survive broker failures without data loss. |
| Consumer Group Rebalancing | If a worker instance crashes, Kafka automatically rebalances partitions across surviving workers within seconds. |
| Exponential Backoff with Jitter | Transient 5xx errors from APNs, FCM, or SendGrid trigger retries with randomized backoff delays to prevent thundering herds. |
| Dead Letter Queue (DLQ) | Malformed payloads or unrecoverable provider rejections move to a channel DLQ after 3 failed retries for investigation. |
| Idempotent Processing Gate | Workers verify Redis deduplication keys before executing external provider API calls to prevent duplicate sends. |
| Provider Circuit Breakers | Monitors consecutive provider errors and trips open after 5 consecutive failures, routing traffic to secondary providers. |
Problem-Specific Failure Handling
1. Downstream Provider Outages (e.g., SendGrid Downtime)
- Circuit breakers trip after 5 consecutive request failures or when error rates exceed 50% across a 60-second window.
- Traffic automatically shifts to secondary providers such as AWS SES without requiring manual operator intervention.
- If all providers for a channel fail, messages buffer safely in Kafka topics until provider connectivity recovers.
2. Mobile Device Token Invalidation
- When APNs or FCM returns an invalid-token response after an app uninstallation, the worker marks the token invalid in the database.
- Future notifications skip the invalidated token, and if no valid tokens remain, the system falls back to email or SMS if allowed by user preferences.
3. Duplicate Notification Mitigation
- Kafka consumers commit offsets only after receiving provider acknowledgment, meaning that if a worker crashes before completing the commit, the message is re-delivered to another worker.
- Workers check the Redis deduplication cache using
notification_idbefore dispatching to the external provider. - Provider-native deduplication headers (such as
apns-collapse-idfor iOS) provide an additional safety net.
4. Broadcast Campaign Surges (Marketing Blasts)
- Bulk promotional campaigns publish to low-priority Kafka partitions with dedicated rate-limited worker pools.
- Transactional alerts (such as password resets and order confirmations) publish to dedicated high-priority topics with isolated consumer instances to prevent head-of-line blocking.
5. Offline Recipient Devices
- Push notifications dispatched while a user's device is disconnected are buffered by APNs and FCM until the device reconnects.
- The notification status remains in the
SENTstate until the device emits a delivery receipt, updating the record toDELIVERED.
Additional Considerations
Advanced topics evaluate notification bundling, priority queues, quiet-hours timezone logic, and compliance standards.
Notification Grouping and Digest Bundling
To prevent flooding users with high-frequency social events (such as 50 likes on a post), the system buffers events in a Redis sorted set for 5 minutes. When the timer fires, the engine collapses individual items into a single consolidated summary notification (for example, "User A, User B, and 48 others liked your photo").
Priority Queue Routing Architecture
Kafka topics are partitioned by priority tier, ensuring critical alerts receive immediate processing resources:
Topic Priority Allocation: notifications_critical: Dedicated non-batched worker pool with p99 delivery latency under 5 seconds notifications_high: Standard transactional processing for order updates and account changes notifications_normal: General activity alerts and social updates notifications_low: Batched marketing campaigns and periodic digest summaries
Delivery Funnel Analytics
- Tracks real-time delivery conversion rates across providers to detect regional carrier anomalies.
- Captures email open rates and link click-through metrics via tracking pixels and redirect gateways.
- Streams status events into ClickHouse for operational monitoring dashboards and SLA tracking.
Quiet Hours and Timezone Processing
- Maintains user timezone preferences and checks whether the current local time falls within configured quiet hours.
- Non-critical notifications arriving during quiet hours are routed to a delayed scheduling queue and dispatched once quiet hours end.
- Critical P0 security alerts bypass quiet-hours restrictions to ensure urgent safety notifications deliver immediately.
Regulatory Compliance and Privacy Standards
- Includes mandatory one-click unsubscribe links and physical mailing headers in all marketing emails (CAN-SPAM and GDPR).
- Enforces verifiable opt-in verification and opt-out keywords (such as STOP) for all SMS workflows (TCPA compliance).
- Sanitizes personally identifiable information (PII) from payload logs before streaming records to analytics data warehouses.
Related Problems and Concepts
Multi-channel fan-out patterns connect directly to offline push notifications in Design a Real-Time Chat System and scheduled reminders in Design a Shared Calendar. Distributed idempotency and deduplication mechanisms mirror patterns in Design a Payment Gateway. Review foundational architectural concepts in Message Queues Fundamentals, Circuit Breaker, Retries, and Bulkheads, Scaling 0 to 1M Users, and System Design Interview Patterns.
Interview Walkthrough
- 25-Minute Interview Strategy
Focus on the synchronous ingress split, Kafka channel isolation, and deduplication before diving into worker pool failover.
- Requirements and Channel Scope Definition (5 min)
- High-Level Architecture and Synchronous/Asynchronous Split (6 min)
- Kafka Topic Partitioning and Deduplication Pipeline (5 min)
- Channel Worker Pools and Provider Circuit Breakers (5 min)
- User Preferences, Quiet Hours, and Compliance Guardrails (4 min)
- Clarify channel requirements upfront across push, email, SMS, and in-app, highlighting their differing latency targets, cost models, and delivery guarantees.
- Separate the synchronous ingress path (returning 202 Accepted) from asynchronous delivery workers using durable Kafka message queues.
- Design dedicated worker pools per channel with provider-specific connection pooling, rate limiters, and automated circuit breakers.
- Evaluate user preferences and suppression lists at the API gateway before enqueueing to guarantee immediate compliance with opt-out rules.
- Apply Redis-backed idempotency keys on the enqueue API to ensure upstream retries do not produce duplicate notifications.
- Walk through capacity calculations, explaining how priority partitioning prevents high-volume marketing blasts from starving critical transactional alerts.
- Highlight the common pitfall of calling external providers synchronously in web request handlers, which causes cascading timeouts when providers slow down.
Engineering Trade-offs
Designing a large-scale notification platform requires evaluating delivery patterns, queue topology isolation, transport reliability semantics, and multi-provider failover strategies.
Push vs Pull for Notification Delivery
Choosing between server-initiated push delivery and client-initiated polling determines real-time latency and network efficiency.
| Delivery Paradigm | Operational Mechanism | Advantages | Disadvantages |
|---|---|---|---|
| Server-Initiated Push ⭐ | Servers push messages to clients via APNs, FCM, or persistent WebSockets | Sub-second real-time delivery with zero wasted bandwidth when idle | Requires persistent connection management and remains subject to provider throttling |
| Client-Initiated Pull | Client applications periodically poll the server for new notification items | Simple server implementation where client applications control query frequency | Wastes bandwidth on empty polls and introduces latency bounded by polling interval |
| Hybrid Signal and Fetch | Server sends a lightweight push notification trigger while client fetches full payload upon receipt | Minimizes push payload overhead while allowing clients to retrieve updated context | Introduces secondary network round-trip on notification open |
Channel-Specific Topics vs Unified Topic Architecture
Partitioning Kafka topics by delivery channel prevents slow providers from introducing head-of-line blocking across the system.
| Queue Architecture | Structural Design | Trade-off Assessment |
|---|---|---|
| Single Unified Topic | All notifications stream into a single shared topic consumed by a unified worker pool | A backlog in slow SMS delivery (2s per request) or bulk email blasts blocks fast push notifications (50ms) in the same queue. |
| Dedicated Channel Topics ⭐ | Separate Kafka topics for push, email, and SMS with independent consumer groups | Enables independent horizontal scaling and ensures SendGrid slowdowns isolate strictly to email without degrading push alerts. |
Delivery Guarantees: At-Least-Once vs Exactly-Once
Distributed notification systems balance delivery guarantees against latency overhead and deduplication complexity.
| Delivery Guarantee | Implementation Mechanism | Recommended Workload |
|---|---|---|
| At-Most-Once | Dispatch and forget with zero retries on provider failure | Acceptable for non-critical marketing broadcasts where missing a delivery is preferable to duplicate spam. |
| At-Least-Once with Dedup ⭐ | Retry on provider failure combined with Redis deduplication keys and client-side deduplication | Recommended for transactional notifications (order confirmations, security alerts, one-time passwords). |
| Strict Exactly-Once | Two-phase commit and distributed locking across provider calls | Impractical at scale due to external provider boundaries, but simulated effectively via at-least-once transport with idempotency keys. |
Provider Failover and Routing Strategy
Managing multiple external delivery providers protects against third-party outages while optimizing delivery costs and deliverability rates.
| Channel | Primary Provider | Secondary Fallback | Routing Criteria |
|---|---|---|---|
| Email Delivery | AWS SES (cost-effective high-volume tier) | SendGrid or Mailgun for transactional failover and specialized deliverability tools | — |
| SMS Messaging | Twilio (high deliverability in North America) | Sinch or MessageBird for regional international cost optimization and carrier failover | — |
| Push Notifications | APNs (iOS) and FCM (Android / Web) | Platform-specific targets with zero cross-platform provider substitution | — |
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.