System Design Problem

Design a Notification System (Push, Email, SMS)

Commonly Asked By:AppleMetaAmazonGoogleUber

Interview Setup

Interview Prompt

Design a notification system that delivers push notifications, SMS, and emails to users across global distributed platforms based on business events such as order updates, social interactions, and urgent security alerts. The system must support priority levels, custom user preferences, quiet hours, and end-to-end delivery tracking.

Clarifying Questions (ask before designing)

QuestionWhy it matters
What delivery latency SLA is expected across different notification types?Critical security alerts require priority queues with delivery under 5 seconds, whereas bulk marketing emails can tolerate several minutes of queue delay.
Is delivery expected to be strictly exactly-once or at-least-once with deduplication?External notification providers duplicate webhooks and retry deliveries, making at-least-once transport combined with idempotency keys the standard architecture.
Does the system need to enforce per-user daily caps and localized quiet hours?Daily caps prevent user notification fatigue, while quiet hours ensure compliance with regulations such as TCPA and GDPR.
Who produces notification events: a centralized monolith or hundreds of distributed microservices?Microservice environments require standardized Kafka event schemas, schema registries, and robust ingress fan-in validation.

Scope

In scope

  • Multi-channel delivery across push notifications, SMS, and email
  • Priority scheduling and asynchronous queue decoupling
  • User preferences, quiet-hours suppression, and regulatory opt-outs
  • Idempotent deduplication and per-user token-bucket rate limiting
  • Delivery receipt tracking, provider failover, and dead letter queues

Out of scope (state explicitly)

  • In-app notification client UI component libraries
  • Full WYSIWYG email marketing visual campaign builders
  • Machine learning models for individual send-time optimization

Functional Requirements

Scope the core capabilities with the interviewer upfront. The system must support multi-channel delivery across push, email, and SMS, handle user preferences and scheduling, and track delivery lifecycles reliably.

  • Deliver notifications across multiple channels: Push (iOS, Android, Web), Email, and SMS.
  • Support both real-time delivery (under 5 seconds) and scheduled future delivery.
  • Enforce user preferences, allowing users to opt into or out of specific channels per notification category.
  • Support dynamic templating with localized variable substitution across languages and device formats.
  • Execute bulk broadcast campaigns (such as promotional updates to 10 million users) without starving transactional alerts.
  • Track end-to-end delivery lifecycle status: queued, sent, delivered, read, failed, and bounced.
  • Enforce rate limits and daily caps per user to prevent notification fatigue and spam.
  • Support notification bundling and digests (such as collapsing multiple social likes into a single summary).
  • Enforce priority levels: critical (immediate), high, normal, and low (batched).

Non-Functional Requirements

Delivery reliability, channel isolation, and sub-second ingress latency are foundational non-functional requirements. The system must handle third-party provider downtime gracefully.

  • High Throughput: Support over 1 million notifications per minute cluster-wide with peak traffic exceeding 60,000 requests per second.
  • Low Latency: Deliver real-time P0 notifications to end-user devices within 1 to 2 seconds under normal network conditions.
  • High Reliability: Ensure at-least-once message delivery so that no valid notification is dropped silently.
  • Horizontal Scalability: Scale to 100 million daily active users generating billions of notifications daily.
  • Channel Fault Tolerance: An outage in an external provider (such as an SMS carrier) must never block or degrade other channels.
  • Best-Effort Ordering: Preserve chronological delivery order for sequential notifications sent to an individual user.
  • Idempotency: Prevent duplicate sends when client apps or upstream services retry requests.
  • Extensibility: Support seamless addition of new channels (such as WhatsApp, Slack, or webhook integrations) without modifying core ingress logic.

Capacity Estimations

Evaluate daily notification volumes and queue depth requirements to size worker pools and storage retention buffers across peak operating windows.

MetricCalculationValue
Daily active users (DAU)Given product assumption100M
Total notifications per day100M DAU x 10 per user1B (10 per user daily average)
Average notification throughput1B ÷ 86,400 seconds~11,574 / sec (peak 5x: ~60,000 / sec)
Push notifications volume60% of total traffic600M per day (~7,000 / sec avg)
Email notifications volume30% of total traffic300M per day (~3,500 / sec avg)
SMS notifications volume10% of total traffic100M per day (~1,150 / sec avg)
Average notification payload sizeRendered message record with metadata~500 bytes
Daily storage ingestion rate1B x 500 bytes500 GB per day
Annual storage footprint500 GB x 365 days~182.5 TB per year

Throughput and Storage Sizing Details

  • Peak Ingress Rate: 1 billion daily notifications average roughly 11,574 requests per second. Factoring in a 5x peak traffic multiplier during flash sales yields a target peak ingress of approximately 60,000 notifications per second.
  • Channel Distribution: Push accounts for 600 million daily messages (~7,000/sec avg), email accounts for 300 million (~3,500/sec avg), and SMS represents 100 million (~1,150/sec avg).
  • Log Storage Footprint: Storing 500-byte notification records for 1 billion daily notifications generates 500 GB of logs daily, requiring roughly 45 TB for a 90-day retention window in Cassandra.

Architecture Diagram

Interview strategy: Emphasize the synchronous and asynchronous split early. The ingress API returns a 202 Accepted status immediately after validating, checking deduplication, and enqueueing the message into Kafka. Returning 202 Accepted acknowledges durable acceptance for asynchronous processing rather than guaranteed or completed delivery to the end user.

The architecture separates synchronous request intake from asynchronous channel delivery. Internal services submit notifications through a central API layer that validates requests, filters opt-outs and quiet hours, enforces Redis deduplication and rate limits, renders channel-specific templates, and enqueues messages onto channel-specific Kafka topics before returning 202 Accepted.

Independent worker pools consume from dedicated Kafka topics to communicate with external delivery providers such as APNs, FCM, AWS SES, and Twilio. As providers emit delivery receipts, an asynchronous Delivery Tracker service streams status updates into Cassandra for notification history and ClickHouse for analytics dashboards.

Loading...

Component Deep Dives

Deconstructing the notification architecture reveals how the ingress gateway, template rendering engine, Kafka event bus, and worker pools maintain isolation across channels.

Notification Service (Ingress API Gateway)

The API layer serves as the single ingress point for all internal microservices. It validates request payloads, verifies recipient preferences and quiet hours, renders localized templates, and fans out messages to channel-specific Kafka topics before returning an immediate 202 Accepted response.

  • Validation: Enforces schema correctness, verifies user existence, and confirms template availability.
  • Preference Filtering: Queries cached user preferences to suppress channels the user has disabled for that notification category.
  • Rate Limiting: Executes token-bucket checks in Redis to enforce per-user channel thresholds (such as a maximum of 10 push notifications per hour).
  • Quiet Hours Evaluation: Converts the current timestamp to the user's local timezone and defers non-critical notifications until the quiet window concludes.
  • Idempotency Enforcement: Checks the request_id against Redis using a SET NX command with a 24-hour expiration to prevent duplicate enqueueing.

Template Rendering Pipeline

After preference and quiet-hours checks pass, the API layer resolves the corresponding template_id and user locale from PostgreSQL. It substitutes dynamic parameters (such as customer name and tracking links) to produce channel-tailored payloads: push headers under 256 characters, full responsive HTML email MIME bodies, and segment-optimized SMS text.

Pre-rendering templates at the API layer allows downstream channel workers to operate as lightweight delivery agents that never execute secondary database lookups.

User Preferences and Suppression Service

User preferences and global suppression lists reside in PostgreSQL to guarantee strong consistency, ensuring marketing opt-outs take effect immediately without violating regulatory compliance standards.

  • Granular Settings: Configurable per notification category (marketing, transactional, social) and per channel (push, email, SMS).
  • Suppression Registry: Maintains bounced email addresses, unsubscribed phone numbers, and invalidated push tokens to prevent repetitive failed deliveries.

Kafka Message Bus and Topic Topology

Kafka isolates high-volume marketing traffic from critical alerts and buffers transient downstream provider slowdowns. Topics are partitioned by user_id to preserve strict chronological ordering for each individual recipient.

  • Dedicated Channel Topics: push_notifications, email_notifications, sms_notifications, and notification_status.
  • Durable Asynchronous Decoupling: Configured with a replication factor of 3 and min.insync.replicas=2, supporting a 7-day retention period. Kafka guarantees at-least-once message delivery to worker pools, while final handset delivery depends on provider responses and retries.
  • Blast Isolation: If an external email provider experiences latency, only the email topic accumulates lag while push notifications and SMS alerts proceed uninterrupted.

Push Notification Worker Pool

Push workers manage persistent HTTP/2 connections to Apple Push Notification service (APNs) and Firebase Cloud Messaging (FCM), minimizing handshake latency across high-throughput bursts.

  • Connection Pooling: Reuses warm HTTP/2 multiplexed connections to avoid per-message TLS handshake overhead.
  • Token Invalidation Handling: When APNs returns an HTTP 410 Gone or FCM returns a NotRegistered error, workers flag the device token as invalid in PostgreSQL to suppress future attempts.

Email Worker Pool

Email workers ingest pre-compiled MIME payloads from Kafka and dispatch them through primary and secondary delivery providers with automated circuit breakers.

  • Multi-Provider Routing: Routes bulk traffic through AWS SES and transactional alerts through SendGrid, failing over automatically if error rates rise.
  • Deliverability Guardrails: Enforces SPF, DKIM, and DMARC record signing while managing dedicated IP warmup schedules to safeguard domain reputation.

SMS Worker Pool

Because SMS carries significant per-segment carrier costs, SMS workers apply strict rate caps and dynamically route messages to the most cost-effective provider by destination country code.

  • Regional Routing: Directs domestic North American traffic through Twilio and international deliveries through regional aggregators like Sinch or MessageBird.
  • Segment Optimization: Validates character encodings (GSM-7 vs UCS-2) to avoid accidental multi-segment billing overhead.

Delivery Tracker and Telemetry Service

Delivery Tracker processes worker dispatch receipts and asynchronous provider webhooks to record lifecycle milestones in Cassandra and ClickHouse.

  • State Transitions: Tracks progression across QUEUED, SENT, DELIVERED, READ, FAILED, and BOUNCED, distinguishing between successful provider acceptance and confirmed handset delivery based on provider capabilities.
  • Storage Split: Writes notification records to Cassandra for fast user-facing inbox lookups and streams analytics events to ClickHouse for delivery funnel dashboards.

Event Bus Configuration Reference

The Kafka event bus schema and consumer group topology structure message flow across partitions and dead-letter queues:

Topic: push_notifications
  Partitions: 128
  Partition key: user_id (preserves per-user delivery order)
  Retention: 7 days
  Replication factor: 3, min.insync.replicas: 2

Topic: email_notifications
  Partitions: 64
  Partition key: user_id
  Retention: 7 days

Topic: sms_notifications
  Partitions: 32
  Partition key: user_id
  Retention: 7 days

Topic: notification_status
  Partitions: 32
  Partition key: notification_id
  Retention: 30 days (delivery webhook audit trail)

Producer: Notification Service (executes validation, deduplication, and template rendering)
  Event schema: { notification_id, request_id, user_id, channel, template_id, priority, payload, created_at }

Consumer groups (per channel pipeline):
  1. push-workers: Persistent HTTP/2 connection pooling to APNs and FCM
  2. email-workers: SendGrid and AWS SES with automated circuit-breaker failover
  3. sms-workers: Twilio and Nexmo with regional country routing
  4. status-tracker: Streams delivery receipts into Cassandra and ClickHouse
  Dead Letter Queue: {channel}_notifications-dlq after 3 retries, with lag alerts above 60s

Synchronous Hot Path: Validate request -> Redis deduplication on request_id -> Enqueue to channel topics -> Return 202 Accepted
Asynchronous Delivery Path: Channel worker pools deliver to providers -> Delivery Tracker processes provider webhooks

API Design

The system exposes RESTful HTTP endpoints for internal service integration, status querying, user preference management, and historical inbox access.

Client API Type Definitions

TypeScript domain types defining notification payloads, priority tiers, and delivery status models:

TYPESCRIPT
export type NotificationPriority = "critical" | "high" | "normal" | "low";
export type NotificationChannel = "push" | "email" | "sms" | "in_app";
export type DeliveryStatus = "queued" | "sent" | "delivered" | "read" | "failed" | "bounced";

export interface SendNotificationRequest {
  requestId: string;
  userIds: string[];
  notificationType: string;
  priority: NotificationPriority;
  channels?: NotificationChannel[];
  templateId: string;
  templateVars: Record<string, string>;
  scheduledAt?: string | null;
  metadata?: Record<string, string>;
}

export interface ChannelDeliveryStatus {
  status: DeliveryStatus;
  provider: string;
  sentAt?: string;
  deliveredAt?: string;
  failureReason?: string;
}

export interface NotificationStatusResponse {
  notificationId: string;
  userId: string;
  channels: Record<NotificationChannel, ChannelDeliveryStatus>;
  createdAt: string;
}

export interface NotificationServiceClient {
  send(request: SendNotificationRequest): Promise<{ notificationId: string; status: "queued" }>;
  getStatus(notificationId: string): Promise<NotificationStatusResponse>;
  getUserHistory(userId: string, page: number, limit: number): Promise<NotificationStatusResponse[]>;
}

Send Notification Endpoint

Primary ingress endpoint accepting notification dispatch requests with idempotency keys:

HTTP
POST /api/v1/notifications
Authorization: Bearer <service_token>
Content-Type: application/json

{
  "request_id": "9f8b4c2e-6d1a-4f5e-8b3c-1a2b3c4d5e6f",
  "user_ids": ["user_101", "user_102"],
  "notification_type": "order_shipped",
  "priority": "high",
  "channels": ["push", "email"],
  "template_id": "order_shipped_v2",
  "template_vars": {
    "order_id": "ORD-12345",
    "tracking_url": "https://track.example.com/ORD-12345"
  },
  "scheduled_at": null,
  "metadata": {
    "campaign_id": "spring_launch_2026"
  }
}

HTTP/1.1 202 Accepted
Content-Type: application/json

{
  "notification_id": "notif_7a8b9c0d-1e2f-3a4b-5c6d-7e8f9a0b1c2d",
  "status": "queued",
  "channels_targeted": ["push", "email"],
  "queued_at": "2026-03-13T10:00:00Z"
}

Get Notification Status Endpoint

Retrieves per-channel delivery timestamps and provider feedback for a specific notification:

HTTP
GET /api/v1/notifications/notif_7a8b9c0d-1e2f-3a4b-5c6d-7e8f9a0b1c2d
Authorization: Bearer <service_token>

HTTP/1.1 200 OK
Content-Type: application/json

{
  "notification_id": "notif_7a8b9c0d-1e2f-3a4b-5c6d-7e8f9a0b1c2d",
  "user_id": "user_101",
  "channels": {
    "push": {
      "status": "delivered",
      "provider": "APNs",
      "delivered_at": "2026-03-13T10:00:01.240Z"
    },
    "email": {
      "status": "sent",
      "provider": "AWS_SES",
      "sent_at": "2026-03-13T10:00:01.850Z"
    }
  }
}

Update User Preferences Endpoint

Configures channel opt-ins, quiet-hours boundaries, and notification category rules:

HTTP
PUT /api/v1/users/user_101/notification-preferences
Authorization: Bearer <user_token>
Content-Type: application/json

{
  "channels": {
    "push": true,
    "email": true,
    "sms": false
  },
  "quiet_hours": {
    "start": "22:00",
    "end": "08:00",
    "timezone": "America/New_York"
  },
  "notification_types": {
    "marketing": { "push": false, "email": true },
    "social": { "push": true, "email": false },
    "order_updates": { "push": true, "email": true, "sms": true }
  }
}

HTTP/1.1 200 OK
Content-Type: application/json

{
  "user_id": "user_101",
  "updated_at": "2026-03-13T10:05:00Z",
  "status": "updated"
}

Get User Notification History Endpoint

Paginated query returning recent notifications delivered to a user:

HTTP
GET /api/v1/users/user_101/notifications?page=1&limit=20
Authorization: Bearer <user_token>

HTTP/1.1 200 OK
Content-Type: application/json

{
  "notifications": [
    {
      "notification_id": "notif_7a8b9c0d-1e2f-3a4b-5c6d-7e8f9a0b1c2d",
      "type": "order_shipped",
      "title": "Your order has shipped!",
      "body": "Order ORD-12345 is on its way.",
      "channel": "push",
      "status": "read",
      "created_at": "2026-03-13T10:00:00Z"
    }
  ],
  "pagination": {
    "page": 1,
    "limit": 20,
    "total": 150
  }
}

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpoint

Data Model

Data storage balances relational consistency for preferences and templates with high-throughput distributed stores for logs and caching.

PostgreSQL: User Preferences and Device Tokens

Relational schemas enforce strict data integrity for user notification settings, channel opt-outs, and active mobile device tokens:

SQL
CREATE TABLE user_notification_preferences (
    user_id             UUID PRIMARY KEY,
    push_enabled        BOOLEAN DEFAULT TRUE,
    email_enabled       BOOLEAN DEFAULT TRUE,
    sms_enabled         BOOLEAN DEFAULT FALSE,
    quiet_hours_start   TIME,
    quiet_hours_end     TIME,
    timezone            VARCHAR(64),
    language            VARCHAR(10) DEFAULT 'en',
    updated_at          TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP
);

CREATE TABLE user_type_preferences (
    user_id             UUID,
    notification_type   VARCHAR(64),
    push_enabled        BOOLEAN DEFAULT TRUE,
    email_enabled       BOOLEAN DEFAULT TRUE,
    sms_enabled         BOOLEAN DEFAULT FALSE,
    PRIMARY KEY (user_id, notification_type)
);

CREATE TABLE user_device_tokens (
    user_id             UUID,
    device_token        VARCHAR(256),
    platform            VARCHAR(32), -- 'ios', 'android', 'web'
    app_version         VARCHAR(32),
    is_valid            BOOLEAN DEFAULT TRUE,
    last_active_at      TIMESTAMP WITH TIME ZONE,
    created_at          TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
    PRIMARY KEY (user_id, device_token)
);

Cassandra: Notification Delivery Log

Cassandra provides write-optimized time-series storage partitioned by user_id with a default 90-day time-to-live expiration:

SQL
CREATE TABLE notification_log (
    user_id           UUID,
    created_at        TIMESTAMP,
    notification_id   UUID,
    notification_type VARCHAR,
    channel           VARCHAR,
    title             TEXT,
    body              TEXT,
    status            VARCHAR, -- queued, sent, delivered, read, failed, bounced
    metadata          MAP<TEXT, TEXT>,
    PRIMARY KEY (user_id, created_at, notification_id)
) WITH CLUSTERING ORDER BY (created_at DESC)
  AND default_time_to_live = 7776000; -- 90 days retention

MySQL: Parameterized Notification Templates

Templates are stored with composite primary keys supporting versioning, language localization, and channel formatting:

SQL
CREATE TABLE notification_templates (
    template_id     VARCHAR(128) NOT NULL,
    version         INT NOT NULL,
    channel         VARCHAR(16) NOT NULL, -- 'push', 'email', 'sms'
    language        VARCHAR(10) NOT NULL, -- 'en', 'es', 'de'
    subject         TEXT,                 -- used for email subject lines
    title           TEXT,                 -- used for push headers
    body_template   TEXT NOT NULL,        -- parameterized text: "Hello {{user_name}}, ..."
    html_template   TEXT,                 -- parameterized HTML structure for email
    active          BOOLEAN DEFAULT TRUE,
    created_at      TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
    PRIMARY KEY (template_id, version, channel, language)
);

Redis Key Structure: Deduplication and Rate Limiting

In-memory data structures enforce request deduplication windows and hourly token-bucket rate limits:

# Deduplication key (expires after 24 hours)
Key:   notif:dedup:{request_id}
Value: 1
TTL:   86400

# Per-user hourly rate limiter counter
Key:   notif:rate:{user_id}:{channel}:{hour}
Value: integer counter (incremented via INCR)
TTL:   3600

# Active device tokens cache
Key:   device_tokens:{user_id}
Value: SET of serialized {device_token, platform, app_version, last_active}

Fault Tolerance

External provider outages, duplicate deliveries, and burst-traffic spikes require layered resilience mechanisms.

Resilience Strategy Matrix

TechniqueApplication Mechanism
Kafka Broker DurabilityReplication factor of 3 with min.insync.replicas=2 ensures notifications survive broker failures without data loss.
Consumer Group RebalancingIf a worker instance crashes, Kafka automatically rebalances partitions across surviving workers within seconds.
Exponential Backoff with JitterTransient 5xx errors from APNs, FCM, or SendGrid trigger retries with randomized backoff delays to prevent thundering herds.
Dead Letter Queue (DLQ)Malformed payloads or unrecoverable provider rejections move to a channel DLQ after 3 failed retries for investigation.
Idempotent Processing GateWorkers verify Redis deduplication keys before executing external provider API calls to prevent duplicate sends.
Provider Circuit BreakersMonitors consecutive provider errors and trips open after 5 consecutive failures, routing traffic to secondary providers.

Problem-Specific Failure Handling

1. Downstream Provider Outages (e.g., SendGrid Downtime)

  • Circuit breakers trip after 5 consecutive request failures or when error rates exceed 50% across a 60-second window.
  • Traffic automatically shifts to secondary providers such as AWS SES without requiring manual operator intervention.
  • If all providers for a channel fail, messages buffer safely in Kafka topics until provider connectivity recovers.

2. Mobile Device Token Invalidation

  • When APNs or FCM returns an invalid-token response after an app uninstallation, the worker marks the token invalid in the database.
  • Future notifications skip the invalidated token, and if no valid tokens remain, the system falls back to email or SMS if allowed by user preferences.

3. Duplicate Notification Mitigation

  • Kafka consumers commit offsets only after receiving provider acknowledgment, meaning that if a worker crashes before completing the commit, the message is re-delivered to another worker.
  • Workers check the Redis deduplication cache using notification_id before dispatching to the external provider.
  • Provider-native deduplication headers (such as apns-collapse-id for iOS) provide an additional safety net.

4. Broadcast Campaign Surges (Marketing Blasts)

  • Bulk promotional campaigns publish to low-priority Kafka partitions with dedicated rate-limited worker pools.
  • Transactional alerts (such as password resets and order confirmations) publish to dedicated high-priority topics with isolated consumer instances to prevent head-of-line blocking.

5. Offline Recipient Devices

  • Push notifications dispatched while a user's device is disconnected are buffered by APNs and FCM until the device reconnects.
  • The notification status remains in the SENT state until the device emits a delivery receipt, updating the record to DELIVERED.

Additional Considerations

Advanced topics evaluate notification bundling, priority queues, quiet-hours timezone logic, and compliance standards.

Notification Grouping and Digest Bundling

To prevent flooding users with high-frequency social events (such as 50 likes on a post), the system buffers events in a Redis sorted set for 5 minutes. When the timer fires, the engine collapses individual items into a single consolidated summary notification (for example, "User A, User B, and 48 others liked your photo").

Priority Queue Routing Architecture

Kafka topics are partitioned by priority tier, ensuring critical alerts receive immediate processing resources:

Topic Priority Allocation:
  notifications_critical: Dedicated non-batched worker pool with p99 delivery latency under 5 seconds
  notifications_high:     Standard transactional processing for order updates and account changes
  notifications_normal:   General activity alerts and social updates
  notifications_low:      Batched marketing campaigns and periodic digest summaries

Delivery Funnel Analytics

  • Tracks real-time delivery conversion rates across providers to detect regional carrier anomalies.
  • Captures email open rates and link click-through metrics via tracking pixels and redirect gateways.
  • Streams status events into ClickHouse for operational monitoring dashboards and SLA tracking.

Quiet Hours and Timezone Processing

  • Maintains user timezone preferences and checks whether the current local time falls within configured quiet hours.
  • Non-critical notifications arriving during quiet hours are routed to a delayed scheduling queue and dispatched once quiet hours end.
  • Critical P0 security alerts bypass quiet-hours restrictions to ensure urgent safety notifications deliver immediately.

Regulatory Compliance and Privacy Standards

  • Includes mandatory one-click unsubscribe links and physical mailing headers in all marketing emails (CAN-SPAM and GDPR).
  • Enforces verifiable opt-in verification and opt-out keywords (such as STOP) for all SMS workflows (TCPA compliance).
  • Sanitizes personally identifiable information (PII) from payload logs before streaming records to analytics data warehouses.

Related Problems and Concepts

Multi-channel fan-out patterns connect directly to offline push notifications in Design a Real-Time Chat System and scheduled reminders in Design a Shared Calendar. Distributed idempotency and deduplication mechanisms mirror patterns in Design a Payment Gateway. Review foundational architectural concepts in Message Queues Fundamentals, Circuit Breaker, Retries, and Bulkheads, Scaling 0 to 1M Users, and System Design Interview Patterns.

Interview Walkthrough

  • 25-Minute Interview Strategy

    Focus on the synchronous ingress split, Kafka channel isolation, and deduplication before diving into worker pool failover.

    • Requirements and Channel Scope Definition (5 min)
    • High-Level Architecture and Synchronous/Asynchronous Split (6 min)
    • Kafka Topic Partitioning and Deduplication Pipeline (5 min)
    • Channel Worker Pools and Provider Circuit Breakers (5 min)
    • User Preferences, Quiet Hours, and Compliance Guardrails (4 min)
  • Clarify channel requirements upfront across push, email, SMS, and in-app, highlighting their differing latency targets, cost models, and delivery guarantees.
  • Separate the synchronous ingress path (returning 202 Accepted) from asynchronous delivery workers using durable Kafka message queues.
  • Design dedicated worker pools per channel with provider-specific connection pooling, rate limiters, and automated circuit breakers.
  • Evaluate user preferences and suppression lists at the API gateway before enqueueing to guarantee immediate compliance with opt-out rules.
  • Apply Redis-backed idempotency keys on the enqueue API to ensure upstream retries do not produce duplicate notifications.
  • Walk through capacity calculations, explaining how priority partitioning prevents high-volume marketing blasts from starving critical transactional alerts.
  • Highlight the common pitfall of calling external providers synchronously in web request handlers, which causes cascading timeouts when providers slow down.

Engineering Trade-offs

Designing a large-scale notification platform requires evaluating delivery patterns, queue topology isolation, transport reliability semantics, and multi-provider failover strategies.

Push vs Pull for Notification Delivery

Choosing between server-initiated push delivery and client-initiated polling determines real-time latency and network efficiency.

Delivery ParadigmOperational MechanismAdvantagesDisadvantages
Server-Initiated Push ⭐Servers push messages to clients via APNs, FCM, or persistent WebSocketsSub-second real-time delivery with zero wasted bandwidth when idleRequires persistent connection management and remains subject to provider throttling
Client-Initiated PullClient applications periodically poll the server for new notification itemsSimple server implementation where client applications control query frequencyWastes bandwidth on empty polls and introduces latency bounded by polling interval
Hybrid Signal and FetchServer sends a lightweight push notification trigger while client fetches full payload upon receiptMinimizes push payload overhead while allowing clients to retrieve updated contextIntroduces secondary network round-trip on notification open

Channel-Specific Topics vs Unified Topic Architecture

Partitioning Kafka topics by delivery channel prevents slow providers from introducing head-of-line blocking across the system.

Queue ArchitectureStructural DesignTrade-off Assessment
Single Unified TopicAll notifications stream into a single shared topic consumed by a unified worker poolA backlog in slow SMS delivery (2s per request) or bulk email blasts blocks fast push notifications (50ms) in the same queue.
Dedicated Channel Topics ⭐Separate Kafka topics for push, email, and SMS with independent consumer groupsEnables independent horizontal scaling and ensures SendGrid slowdowns isolate strictly to email without degrading push alerts.

Delivery Guarantees: At-Least-Once vs Exactly-Once

Distributed notification systems balance delivery guarantees against latency overhead and deduplication complexity.

Delivery GuaranteeImplementation MechanismRecommended Workload
At-Most-OnceDispatch and forget with zero retries on provider failureAcceptable for non-critical marketing broadcasts where missing a delivery is preferable to duplicate spam.
At-Least-Once with Dedup ⭐Retry on provider failure combined with Redis deduplication keys and client-side deduplicationRecommended for transactional notifications (order confirmations, security alerts, one-time passwords).
Strict Exactly-OnceTwo-phase commit and distributed locking across provider callsImpractical at scale due to external provider boundaries, but simulated effectively via at-least-once transport with idempotency keys.

Provider Failover and Routing Strategy

Managing multiple external delivery providers protects against third-party outages while optimizing delivery costs and deliverability rates.

ChannelPrimary ProviderSecondary FallbackRouting Criteria
Email DeliveryAWS SES (cost-effective high-volume tier)SendGrid or Mailgun for transactional failover and specialized deliverability tools—
SMS MessagingTwilio (high deliverability in North America)Sinch or MessageBird for regional international cost optimization and carrier failover—
Push NotificationsAPNs (iOS) and FCM (Android / Web)Platform-specific targets with zero cross-platform provider substitution—

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...