System Design Problem

Design an On-Call Escalation System (like PagerDuty / OpsGenie)

Commonly Asked By:PagerDutyAtlassianMicrosoftGoogle

Interview Setup

Interview Prompt

Design an on call escalation system like PagerDuty. Ingest alerts from monitoring systems, route to the right on call engineer via schedules and escalation policies, notify across channels, and escalate if unacknowledged.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Alert volume during steady state vs cascading outage peak?A peak rate of 100,000 alerts per minute during site wide incidents necessitates aggressive grouping and deduplication because the system cannot page responders for every raw event.
Acknowledgment timeout and escalation chain depth?A standard progression starts with a 5 minute timeout for the primary on call, advances to a 10 minute timeout for the secondary on call, and then escalates to the engineering manager, which drives the state machine timers.
Notification channels and delivery guarantees?SMS, voice, and push notifications have varying latency characteristics, requiring a durable queue to buffer dispatches if downstream APIs experience outages.
Schedule complexity including weekly rotations, overrides, and follow the sun coverage?Timezone conversions and daylight saving transitions invalidate naive cron-based schedule calculations.

Scope

In scope

  • Escalation policy state machine
  • Schedule management
  • Acknowledgment timeout
  • Multi-channel alert
  • Capacity estimation with shown math

Out of scope (state explicitly)

  • Full incident management war-room UI
  • Building a single monitoring tool vs integrating with upstream systems
  • Log search and distributed trace analysis

Functional Requirements

System Scope and Core Capabilities

The system manages the end to end incident lifecycle. It ingests alarms from monitoring backends, evaluates on-call rotation schedules, enforces escalation policies, and dispatches multi channel notifications until an engineer acknowledges.

Interview perspective: Highlighting alert fatigue and deduplication early is essential because firing hundreds of pages for a single root cause represents a major system design flaw.

  • Alert Ingestion: Accept incoming alarm webhooks from Prometheus Alertmanager, Datadog, CloudWatch, and custom monitoring backends with signed authentication, replay protection, idempotency, and per-source rate limits.
  • On-Call Rotations: Define calendar schedules including weekly, daily, and follow the sun rotations with primary and secondary responders.
  • Escalation Chains: Automatically escalate unacknowledged alerts to secondary engineers and engineering management after configurable timeouts.
  • Multi Channel Alerts: Page responders across mobile push notifications, SMS text messages, interactive voice calls, email, and Slack.
  • Acknowledge and Resolve: Expose explicit acknowledgment APIs that mark alerts acknowledged or resolved and terminate downstream escalation timers.
  • Alert Aggregation: Correlate related alerts into unified incidents to prevent notification fatigue during cascading failures.
  • Maintenance Silencing: Support scheduled maintenance windows that suppress active pages while preserving audit events.
  • Incident Timeline: Record complete event histories from initial trigger to final resolution for mean time to acknowledge (MTTA) and mean time to resolve (MTTR) analytics.

Non-Functional Requirements

Reliability and Operational Objectives

Notification delivery within thirty seconds and no intentional loss after an accepted alert is durably persisted form the core service level objectives. The reliability of the paging path itself is paramount because failure in the alert infrastructure blinds engineering teams to production outages.

  • Five Nines Availability: Deliver 99.999% availability for the paging path because monitoring failures should not prevent alert dispatch after an alert has been durably accepted.
  • Low Propagation Latency: Route P0 and P1 alerts from initial webhook ingestion to the first responder notification attempt within 30 seconds at p99. Lower priority alerts can use the configured grouping window.
  • Durable Notification Delivery: Provide at least once provider dispatch attempts after an alert is durably accepted, with provider delivery tracked separately because external channels can fail or delay indefinitely.
  • Cascading Outage Scale: Ingest and process more than 100,000 alerts per minute during severe site wide infrastructure collapses.
  • Alert Deduplication: Maintain a 5 minute deduplication window per incident fingerprint and priority to suppress redundant notifications without requiring distributed transport exactly once semantics. P0 and P1 alerts bypass the grouping delay but still use the same idempotency controls.
  • Immutable Audit Logging: Persist an append only audit trail capturing alert timestamps, notification dispatches, and responder acknowledgment actions.

Capacity Estimations

Traffic Projections and Notification Fan Out

While steady state alert volume is modest compared with high throughput metric streams, cascading outages create massive transient bursts. Multi channel fan out further multiplies notifications across push, SMS, and voice channels.

MetricCalculationValue
Alerts / day (normal)Given (steady state operations)10,000
Alerts / sec (avg)10,000 ÷ 86,400~0.12
Alerts / sec (cascade peak)100K alerts/min ÷ 60, before grouping~1,667 ungrouped alerts per second. The design target is well under 1 effective page per second for a single correlated storm after deduplication and grouping
Active On-Call SchedulesDerived50,000 teams
Initial notifications / day10K alerts x ~2.5 channels avg~25,000
Escalations / day~5% of alerts reach L1+~500

Scale Derivations and Peak Sizing

System components must handle both steady state processing and severe burst conditions. The peak example is a capacity target for a modeled correlated outage, not a guarantee that every 100,000 alerts per minute event collapses to fewer than one page per second:

YAML
# Capacity Math and Peak Volume Projections
steady_state_traffic:
  daily_alerts: 10000 alerts per day
  average_ingestion_qps: 10000 / 86400 ≈ 0.12 alerts per second
  active_schedules: 50000 teams
  daily_notifications: 10000 alerts * 2.5 channels average ≈ 25000 notifications per day
  daily_escalations: ~5% of alerts reach Level 1 or higher ≈ 500 escalations per day

cascading_outage_peak:
  peak_ingestion_rate: 100000 alerts per minute (≈ 1667 alerts per second) before grouping
  deduplicated_effective_pages: Under 1 page per second for a modeled correlated storm after fingerprint deduplication and sliding window grouping
  telecom_concurrency: High outbound concurrency required with Twilio and MessageBird to prevent queue backlogs during emergency voice dispatch

Architecture Diagram

End to End Alert Flow Architecture

The architecture receives alert webhooks from external monitoring providers. It authenticates and normalizes each payload, deduplicates repeated signals, evaluates the active on call schedule, and dispatches multi channel notifications with automated escalation when acknowledgments do not arrive.

Loading...

Component Deep Dives

1. Stateful Escalation Policy State Machine

Escalating alarms through responder levels must occur reliably regardless of human intervention. The state machine transitions alerts through deterministic states with persisted audit timestamps:

YAML
team_alert_lifecycle:
  team_id: "backend-infra"
  escalation_policy:
    level_0:
      target: "Primary on call (Alice)"
      channels: ["push", "sms"]
      ack_timeout_seconds: 300 # 5 minutes
    level_1:
      target: "Secondary on call (Bob)"
      channels: ["push", "sms", "voice_call"]
      ack_timeout_seconds: 600 # 10 minutes
    level_2:
      target: "Engineering Manager (Carol)"
      channels: ["voice_call"]
      ack_timeout_seconds: 900 # 15 minutes
    level_3:
      target: "VP of Engineering (Dave)"
      channels: ["voice_call", "sms"]

  execution_trace:
    t0_00: "Alert ingested and dispatched to Alice via Push and SMS"
    t5_00: "No ACK received from Alice, escalating to Level 1 and dispatching to Bob via Push, SMS, and Voice"
    t15_00: "No ACK received from Bob, escalating to Level 2 and dispatching to Carol via Voice"
    t30_00: "No ACK received from Carol, escalating to Level 3 and dispatching to Dave via Voice and SMS"

  termination_condition: "Immediate cancellation of pending timers upon receipt of valid responder acknowledgment or a resolved source event"

2. Escalation Timer Implementation with Redis Sorted Sets ⭐

Scanning database tables on a periodic interval creates heavy query overhead and high contention. We use Redis Sorted Sets as a fast execution index with approximately 100 millisecond polling and short leases. PostgreSQL remains the durable source for escalation state and next action times:

# Redis Sorted Set Escalation Timer Implementation
# Key: escalation_timers
# Score: Unix epoch timestamp when the escalation action is due
# Member: {alert_id}:{escalation_level}
# PostgreSQL is the durable source of escalation state and next_action_at timestamps.
# Redis is the execution index and can be rebuilt after Redis loss.

def enqueue_escalation_timer(redis_client, alert_id: str, level: int, delay_seconds: int):
    fire_timestamp = current_epoch_time() + delay_seconds
    redis_client.zadd("escalation_timers", {f"{alert_id}:{level}": fire_timestamp})

def worker_poll_loop(redis_client):
    while True:
        now = current_epoch_time()
        due_timers = redis_client.zrangebyscore("escalation_timers", 0, now, start=0, num=100)

        for item in due_timers:
            # Keep the timer until dispatch succeeds. A short lease prevents
            # competing workers from processing the same timer concurrently.
            lease_key = "timer:claim:" + item
            owner_token = random_token()
            claimed = redis_client.set(lease_key, owner_token, nx=True, px=5000)
            if not claimed:
                continue

            try:
                alert_id, level_str = item.split(":")
                level = int(level_str)
                # Revalidate the alert state in PostgreSQL before the provider call.
                # If the alert was acknowledged, resolved, or rescheduled, do not page.
                if not is_escalation_still_due(alert_id, level):
                    redis_client.zrem("escalation_timers", item)
                    continue
                # The dispatch transaction creates or reuses a stable notification_attempt
                # idempotency key and persists the next_action_at transition in PostgreSQL.
                # Provider calls use that same logical idempotency identity where supported.
                trigger_escalation_idempotently(alert_id, level, stable_notification_key(alert_id, level))
                redis_client.zrem("escalation_timers", item)

                if has_next_escalation_level(alert_id):
                    next_level = get_next_level(alert_id)
                    timeout = get_level_timeout(alert_id, next_level)
                    enqueue_escalation_timer(redis_client, alert_id, next_level, timeout)
            finally:
                release_lease_if_owner(redis_client, lease_key, owner_token)

        sleep(0.1)  # Approximately 100 ms scheduling precision target. Local elapsed-time checks use a monotonic clock.

3. Alert Grouping and Noise Reduction

During cascading infrastructure failures, hundreds of downstream microservices fire alerts simultaneously. A multi layer grouping and deduplication pipeline merges related alerts into a single actionable incident:

YAML
alert_grouping_and_deduplication:
  problem_context:
    without_grouping: "Cascading outage causes 500 downstream microservices to fire alerts, overwhelming on call with alarm fatigue"

  sliding_window_strategy:
    window_duration_seconds: 300 # 5 minute sliding correlation window
    emergency_bypass: "P0 and P1 notifications dispatch immediately"
    correlation_after: "Correlation continues in parallel so later related alerts can be consolidated"
    correlation_keys:
      - "cluster"
      - "environment"
      - "alert_name"
      - "dependency_id"
    note: "Service remains an incident attribute, but dependency-aware correlation allows related alerts from multiple services to be grouped"
    owner_lifecycle: "The canonical owner alert carries escalation state. If it resolves while related alerts remain active, the grouper elects another active alert as owner in the same transaction before the next escalation action is scheduled."
    priority_rule: "The correlated incident inherits the highest active member priority. A transition to P0 or P1 uses the urgent path immediately while the correlation window continues."
    outcome:
      incident_title: "Database connection timeout (503 occurrences across 50 downstream microservices)"
      action: "Dispatch one consolidated page to the on call responder. Acknowledging the incident silences grouped notifications while preserving all underlying alert events"

4. Event Bus Architecture and Kafka Partitioning

Decoupling synchronous webhook ingestion from asynchronous schedule matching and notification dispatch prevents downstream telecom latencies from blocking incoming alarms:

YAML
topics:
  alert-events:
    partitions: 32
    partition_key: dedup_key # preserves ordering for repeated events with the same fingerprint
    retention_days: 7

  grouped-alerts:
    partitions: 32
    partition_key: incident_correlation_key # brings alerts for the same incident onto one correlation stream
    retention_days: 7
    payload_fields: [incident_correlation_key, incident_owner_alert_id, grouped_alert_count]

  urgent-alerts:
    partitions: 32
    partition_key: incident_correlation_key
    retention_days: 7

  escalation-actions:
    partitions: 32
    partition_key: alert_id
    retention_days: 7

  notification-events:
    partitions: 32
    partition_key: alert_id
    retention_days: 7
    producers: [Notification Router, Provider Callback Handler]
    publication_rule: Notification state and each notification event are written through the transactional outbox before CDC publication

cluster_configuration:
  replication_factor: 3
  min_insync_replicas: 2

producer:
  service: Alert Ingestion Service
  execution_point: PostgreSQL transaction writes alert state and the outbox row. CDC publishes the event asynchronously
  payload_schema:
    alert_id: "8a219b1b-640a-4289-9812-42171542fca1"
    dedup_key: "high_cpu_payment_us_east"
    source_event_id: "evt_20260314_000001"
    event_type: "firing"
    idempotency_key: "prometheus:evt_20260314_000001"
    severity: "critical"
    service: "payment-service"
    labels:
      env: "production"
      cluster: "us-east-1"
    status: "triggered"
    source: "prometheus"
    alert_version: 7
    priority: "P1"
    timestamp: "2026-03-14T10:00:00Z"

consumer_groups:
  alert_grouper:
    consumes: alert-events
    filter: "priority in {P2, P3}"
    produces: grouped-alerts
    state_store: Durable keyed correlation state with a changelog so a grouper restart does not lose an open five minute window
    description: Coalesces exact duplicate source events, applies the initial 30 to 60 second notification grouping, continues broader incident correlation for five minutes, selects a canonical owner alert, and emits grouped incident state for downstream routing
  schedule_router:
    consumes: [grouped-alerts, urgent-alerts]
    produces: escalation-actions
    description: Resolves the effective team policy and on call schedule for the canonical incident owner, validates silence state, and creates or advances the escalation state machine
  urgent_router:
    consumes: alert-events
    filter: "priority in {P0, P1}"
    produces: urgent-alerts
    description: Bypasses the grouping delay for urgent alerts, revalidates the active incident state, and emits only required state transitions to the immediate notification path
  analytics_sink:
    consumes: [alert-events, notification-events]
    description: Streams alert state transitions and notification delivery events to ClickHouse for MTTA, MTTR, volume, and delivery analysis
  notification_router:
    consumes: escalation-actions
    produces: notification-events
    description: Creates idempotent notification attempts, dispatches to providers, and records provider callback state through the durable notification store
  consumer_idempotency:
    rule: Consumers key side effects by source_event_id and alert_version or by notification_attempt idempotency_key before committing a state transition

execution_paths:
  synchronous_path: Webhook Ingest -> Authenticate and normalize -> Advisory dedup check in Redis -> Validate the active incident in PostgreSQL on cache hit or miss -> Persist PostgreSQL alert state and outbox row atomically -> HTTP 202 Accepted
  asynchronous_path: PostgreSQL CDC/outbox -> alert-events -> Alert Grouper -> grouped-alerts -> Schedule Router -> escalation-actions -> Escalation Timers in Redis ZSET -> Notification Router -> provider callbacks -> notification-events
  urgent_path: alert-events -> Urgent Router for P0/P1 -> urgent-alerts -> Schedule Router -> escalation-actions, while normal correlation continues independently
  dead_letter_queue: "Failed events move to alert-events-dlq or grouped-alerts-dlq after 3 retries. Operators can replay a DLQ event after the root cause is fixed because the original source event ID remains the idempotency key. A separate lag alert pages if consumer lag exceeds 30 seconds."

API Design

Core API Interface Signatures

Typed TypeScript interfaces define the programmatic contract between external monitoring systems, responder clients, and internal escalation workers:

TYPESCRIPT
// Core API Interface Contracts for On-Call Escalation Engine
interface IngestAlertRequest {
  source: string;
  alertName: string;
  severity: "critical" | "warning" | "info";
  priority: "P0" | "P1" | "P2" | "P3";
  service: string;
  labels: Record<string, string>;
  description?: string;
  dedupKey?: string; // Optional provider fingerprint. If absent, the service derives a canonical fingerprint from normalized alert fields.
  sourceEventId: string;
  eventType: "firing" | "resolved";
  occurredAt: string;
}

interface IngestAlertResponse {
  alertId: string;
  status: "triggered" | "resolved" | "deduplicated";
  acceptedAt: string;
}

interface AcknowledgeRequest {
  responderId: string;
  channel: "mobile" | "web" | "sms" | "voice";
  expectedVersion: number;
}

interface AcknowledgeResponse {
  alertId: string;
  status: "acknowledged" | "already_acknowledged";
  acknowledgedAt: string;
  version: number;
}

interface ResolveRequest {
  responderId: string;
  resolutionCode: string;
  expectedVersion: number;
}

interface ResolveResponse {
  alertId: string;
  status: "resolved";
  resolvedAt: string;
  version: number;
}

interface ActiveOnCallResponder {
  userId: string;
  email: string;
  phone?: string;
  layer: number;
  rotationPriority: number;
  scheduleVersion: number;
}

interface ActiveOnCallResponse {
  teamId: string;
  effectiveAt: string;
  scheduleVersion: number;
  coverageStatus: "covered" | "gap";
  primary?: ActiveOnCallResponder;
  secondary?: ActiveOnCallResponder;
}

interface ScheduleOverride {
  targetUserId: string;
  coveringUserId: string;
  startTime: string;
  endTime: string;
  reason: string;
}

interface OverrideResponse {
  overrideId: string;
  status: "scheduled" | "active";
  effectiveFrom: string;
  effectiveUntil: string;
}

interface SnoozeRequest {
  responderId: string;
  durationSeconds: number;
  expectedVersion: number;
}

interface SnoozeResponse {
  alertId: string;
  status: "snoozed";
  snoozedUntil: string;
  version: number;
}

interface SilenceMatcher {
  label: string;
  value: string;
}

interface CreateSilenceRequest {
  teamId: string;
  matchers: SilenceMatcher[];
  startsAt: string;
  endsAt: string;
  comment?: string;
}

interface CreateSilenceResponse {
  silenceId: string;
  status: "scheduled" | "active";
  startsAt: string;
  endsAt: string;
  version: number;
}

export interface OnCallEscalationService {
  // Ingest incoming webhook alert from monitoring systems with deduplication
  ingestAlert(payload: IngestAlertRequest): Promise<IngestAlertResponse>;

  // Acknowledge an active alert, halting downstream escalation timers
  acknowledgeAlert(alertId: string, request: AcknowledgeRequest): Promise<AcknowledgeResponse>;

  // Resolve an active alert when operational health is restored
  resolveAlert(alertId: string, request: ResolveRequest): Promise<ResolveResponse>;

  // Snooze an active alert for a bounded period without resolving it
  snoozeAlert(alertId: string, request: SnoozeRequest): Promise<SnoozeResponse>;

  // Retrieve active on call responders for a team at a specific query timestamp
  getActiveOnCall(teamId: string, queryTimestamp?: string): Promise<ActiveOnCallResponse>;

  // Create a time bounded schedule override for shift coverage or swaps
  createScheduleOverride(scheduleId: string, override: ScheduleOverride): Promise<OverrideResponse>;

  // Create a time bounded maintenance silence for matching alert labels
  createSilence(request: CreateSilenceRequest): Promise<CreateSilenceResponse>;
}

1. Ingest Alert Webhook

External monitoring agents post alert payloads to the ingestion gateway:

HTTP
POST /api/v1/alerts
Content-Type: application/json
Authorization: Bearer <integration_token>
Idempotency-Key: prometheus:evt_20260314_000001
X-Signature: sha256=<hmac>

{
  "source": "prometheus",
  "alert_name": "HighCPUUsage",
  "severity": "critical",
  "priority": "P1",
  "service": "payment-service",
  "labels": {
    "env": "production",
    "cluster": "us-east-1"
  },
  "description": "CPU usage > 90% for 5 consecutive minutes",
  "dedup_key": "high_cpu_payment_us_east",
  "source_event_id": "evt_20260314_000001",
  "event_type": "firing",
  "occurred_at": "2026-03-14T10:00:00Z"
}

Response: 202 Accepted
{
  "alert_id": "8a219b1b-640a-4289-9812-42171542fca1",
  "status": "triggered",
  "accepted_at": "2026-03-14T10:00:00Z"
}

2. Acknowledge Active Alert

Responders acknowledge alerts via mobile apps, SMS replies, or web dashboards to stop escalation timers:

HTTP
POST /api/v1/alerts/8a219b1b-640a-4289-9812-42171542fca1/acknowledge
Content-Type: application/json
Authorization: Bearer <responder_token>

{
  "responder_id": "user-881",
  "channel": "mobile",
  "expected_version": 7
}

Response: 200 OK
{
  "alert_id": "8a219b1b-640a-4289-9812-42171542fca1",
  "status": "acknowledged",
  "acknowledged_at": "2026-03-14T10:05:12Z",
  "version": 8
}

3. Snooze Active Alert

Responders can temporarily suppress repeat notifications without resolving the incident. The operation uses optimistic version checking and reschedules the next escalation action after the snooze period.

HTTP
POST /api/v1/alerts/8a219b1b-640a-4289-9812-42171542fca1/snooze
Content-Type: application/json
Authorization: Bearer <responder_token>

{
  "responder_id": "user-881",
  "duration_seconds": 3600,
  "expected_version": 7
}

Response: 200 OK
{
  "alert_id": "8a219b1b-640a-4289-9812-42171542fca1",
  "status": "snoozed",
  "snoozed_until": "2026-03-14T11:05:12Z",
  "version": 8
}

4. Retrieve Active On-Call Schedule

The routing service inspects active schedule rosters for a team at a specific point in time:

HTTP
GET /api/v1/schedules/team-backend-infra/on-call?at=2026-03-14T10:00:00Z
Authorization: Bearer <responder_token>

Response: 200 OK
{
  "team_id": "team-backend-infra",
  "primary": {
    "user_id": "user-881",
    "email": "alice@company.com",
    "phone": "+15550199",
    "layer": 0,
    "rotation_priority": 100,
    "schedule_version": 42
  },
  "secondary": {
    "user_id": "user-902",
    "email": "bob@company.com",
    "phone": "+15550212",
    "layer": 1,
    "rotation_priority": 90,
    "schedule_version": 42
  },
  "effective_at": "2026-03-14T10:00:00Z",
  "schedule_version": 42,
  "coverage_status": "covered"
}

Standard HTTP Error Responses

Standardized HTTP error responses communicate validation failures, authentication issues, and optimistic locking race conditions:

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
422 Unprocessable Entity: password does not meet security requirements or MFA verification is required
423 Locked: account temporarily locked after repeated failed login attempts
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpoint
401 Unauthorized: webhook signature, integration token, or authentication credentials are invalid
404 Not Found: requested alert_id, team_id, or schedule_id does not exist
409 Conflict: alert state, expected version, or schedule override is stale because another operation changed it first
422 Unprocessable Entity: invalid escalation policy level, schedule interval, or unsupported notification channel configuration
429 Too Many Requests: source or responder has exceeded the configured ingestion or notification rate limit
503 Service Unavailable: a required scheduling, queue, or database dependency is temporarily unavailable. Previously accepted alerts remain durable while downstream provider failures are retried asynchronously

Data Model

PostgreSQL Database Schema

PostgreSQL provides durable transactional storage for alert states, pinned escalation policy versions, normalized schedule assignments, notification delivery state, incident event history, and temporary calendar overrides:

SQL
-- Authoritative alert and incident state
CREATE TABLE alerts (
    alert_id              UUID PRIMARY KEY,
    source                VARCHAR(64) NOT NULL,
    latest_source_event_id VARCHAR(256),
    source_state          VARCHAR(16) NOT NULL DEFAULT 'firing', -- firing or resolved
    last_source_event_at  TIMESTAMP WITH TIME ZONE,
    dedup_key             VARCHAR(256) NOT NULL,
    incident_correlation_key VARCHAR(256),
    incident_owner        BOOLEAN NOT NULL DEFAULT TRUE,
    alert_name            VARCHAR(256) NOT NULL,
    severity              VARCHAR(20) NOT NULL,
    priority              VARCHAR(4) NOT NULL,
    service               VARCHAR(128) NOT NULL,
    labels                JSONB,
    description           TEXT,
    status                VARCHAR(32) NOT NULL DEFAULT 'triggered', -- triggered, notified_primary, escalated_secondary, escalated_manager, escalated_vp, acknowledged, snoozed, resolved
    escalation_level      INT NOT NULL DEFAULT 0,
    team_id               UUID NOT NULL,
    policy_id             UUID,
    policy_version        BIGINT,
    occurrence_count      BIGINT NOT NULL DEFAULT 1,
    last_seen_at          TIMESTAMP WITH TIME ZONE NOT NULL DEFAULT NOW(),
    acknowledged_by      VARCHAR(256),
    acknowledged_at      TIMESTAMP WITH TIME ZONE,
    resolved_by           VARCHAR(256),
    resolved_at           TIMESTAMP WITH TIME ZONE,
    snoozed_until         TIMESTAMP WITH TIME ZONE,
    next_action_at        TIMESTAMP WITH TIME ZONE,
    version               BIGINT NOT NULL DEFAULT 0,
    created_at            TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
    updated_at            TIMESTAMP WITH TIME ZONE DEFAULT NOW()
);

-- The database is authoritative for active incident deduplication by team, fingerprint, and priority.
-- Redis is an advisory accelerator and may be rebuilt after cache loss.
CREATE UNIQUE INDEX idx_alerts_active_dedup
    ON alerts(team_id, dedup_key, priority)
    WHERE status <> 'resolved';

-- Related alerts can be correlated into one incident for paging while each source alert remains auditable.
-- The canonical owner alert carries the escalation state and next action for that correlated incident.
CREATE UNIQUE INDEX idx_alerts_active_incident_owner
    ON alerts(team_id, incident_correlation_key)
    WHERE status <> 'resolved' AND incident_owner = TRUE AND incident_correlation_key IS NOT NULL;

CREATE INDEX idx_alerts_active_team
    ON alerts(team_id, status, updated_at)
    WHERE status <> 'resolved';

CREATE INDEX idx_alerts_next_action
    ON alerts(next_action_at)
    WHERE next_action_at IS NOT NULL AND status NOT IN ('acknowledged', 'resolved');

CREATE INDEX idx_alerts_status
    ON alerts(status);

-- Immutable state transition history for incident timelines and MTTA / MTTR reconstruction.
CREATE TABLE alert_events (
    event_id              UUID PRIMARY KEY,
    alert_id              UUID NOT NULL REFERENCES alerts(alert_id),
    source                 VARCHAR(64),
    source_event_id       VARCHAR(256),
    event_type            VARCHAR(32) NOT NULL,
    actor_id              VARCHAR(256),
    channel               VARCHAR(32),
    occurred_at           TIMESTAMP WITH TIME ZONE,
    state_version         BIGINT,
    payload               JSONB,
    created_at             TIMESTAMP WITH TIME ZONE DEFAULT NOW()
);

CREATE UNIQUE INDEX idx_alert_events_source_event
    ON alert_events(source, source_event_id)
    WHERE source IS NOT NULL AND source_event_id IS NOT NULL;

CREATE INDEX idx_alert_events_alert_time
    ON alert_events(alert_id, created_at);

-- Provider callbacks are normalized into notification-events so analytics can measure dispatch, delivery, and failure independently of alert ingestion.

-- Durable notification state and provider reconciliation data.
CREATE TABLE notification_attempts (
    attempt_id            UUID PRIMARY KEY,
    alert_id              UUID NOT NULL REFERENCES alerts(alert_id),
    escalation_level      INT NOT NULL,
    recipient_user_id     VARCHAR(128) NOT NULL,
    channel               VARCHAR(32) NOT NULL,
    idempotency_key       VARCHAR(256) NOT NULL UNIQUE,
    provider              VARCHAR(64),
    provider_message_id   VARCHAR(256),
    status                VARCHAR(32) NOT NULL,
    attempt_count         INT NOT NULL DEFAULT 1,
    next_retry_at         TIMESTAMP WITH TIME ZONE,
    sent_at               TIMESTAMP WITH TIME ZONE,
    delivered_at          TIMESTAMP WITH TIME ZONE,
    acknowledged_at       TIMESTAMP WITH TIME ZONE,
    created_at            TIMESTAMP WITH TIME ZONE DEFAULT NOW()
);

CREATE INDEX idx_notification_attempts_alert
    ON notification_attempts(alert_id, escalation_level, channel);

CREATE INDEX idx_notification_attempts_provider_message
    ON notification_attempts(provider, provider_message_id)
    WHERE provider_message_id IS NOT NULL;

-- Provider callback handlers update notification_attempts and write the corresponding notification event to the outbox in the same transaction.
-- CDC publishes notification-events asynchronously after the state transition commits.

-- Versioned escalation policies. Active incidents retain the policy version selected when routed.
CREATE TABLE escalation_policies (
    policy_id             UUID PRIMARY KEY,
    team_id               UUID NOT NULL,
    version               BIGINT NOT NULL,
    levels                JSONB NOT NULL, -- [{level: 0, wait_min: 5, channels: ["push", "sms"]}, ...]
    status                VARCHAR(16) NOT NULL DEFAULT 'active',
    created_at            TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
    UNIQUE (team_id, version)
);

-- Base recurring schedule definition.
CREATE TABLE on_call_schedules (
    schedule_id           UUID PRIMARY KEY,
    team_id               UUID NOT NULL,
    rotation_type         VARCHAR(20) NOT NULL,
    recurrence_rule       JSONB NOT NULL,
    participants          JSONB NOT NULL,
    timezone              VARCHAR(64) DEFAULT 'UTC',
    schedule_version      BIGINT NOT NULL DEFAULT 1,
    updated_at            TIMESTAMP WITH TIME ZONE DEFAULT NOW()
);

-- Materialized time bounded assignments support indexed point in time schedule lookup.
CREATE TABLE schedule_assignments (
    assignment_id         UUID PRIMARY KEY,
    schedule_id           UUID NOT NULL REFERENCES on_call_schedules(schedule_id),
    team_id               UUID NOT NULL,
    user_id               VARCHAR(128) NOT NULL,
    layer                 INT NOT NULL,
    priority              INT NOT NULL,
    start_time            TIMESTAMP WITH TIME ZONE NOT NULL,
    end_time              TIMESTAMP WITH TIME ZONE NOT NULL,
    schedule_version      BIGINT NOT NULL,
    CHECK (end_time > start_time)
);

CREATE INDEX idx_schedule_assignments_lookup
    ON schedule_assignments(team_id, layer, priority, start_time DESC);

-- Schedule overrides define temporary coverage changes. Overlap validation is serialized per schedule so concurrent writes cannot create ambiguous active overrides.
CREATE TABLE schedule_overrides (
    override_id           UUID PRIMARY KEY,
    schedule_id           UUID NOT NULL REFERENCES on_call_schedules(schedule_id),
    target_user_id        VARCHAR(128) NOT NULL,
    covering_user_id      VARCHAR(128) NOT NULL,
    priority              INT NOT NULL DEFAULT 100,
    reason                TEXT NOT NULL,
    start_time            TIMESTAMP WITH TIME ZONE NOT NULL,
    end_time              TIMESTAMP WITH TIME ZONE NOT NULL,
    created_by            VARCHAR(128) NOT NULL,
    created_at            TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
    CHECK (end_time > start_time),
    CHECK (target_user_id <> covering_user_id)
);

CREATE INDEX idx_schedule_overrides_lookup
    ON schedule_overrides(schedule_id, start_time, end_time);

-- The scheduling service validates overlapping windows inside a serializable transaction or per-schedule lock.

-- Maintenance silence rules are persisted so ingestion and routing workers share the same suppression policy.
CREATE TABLE silences (
    silence_id            UUID PRIMARY KEY,
    team_id               UUID NOT NULL,
    matchers              JSONB NOT NULL,
    version               BIGINT NOT NULL DEFAULT 1,
    starts_at             TIMESTAMP WITH TIME ZONE NOT NULL,
    ends_at               TIMESTAMP WITH TIME ZONE NOT NULL,
    created_by            VARCHAR(128) NOT NULL,
    comment               TEXT,
    created_at            TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
    CHECK (ends_at > starts_at)
);

CREATE INDEX idx_silences_active_window
    ON silences(team_id, starts_at, ends_at);

-- Transactional outbox. Alert state and this row commit atomically. CDC publishes the row asynchronously.
CREATE TABLE outbox_events (
    event_id              UUID PRIMARY KEY,
    aggregate_type        VARCHAR(64) NOT NULL,
    aggregate_id          UUID NOT NULL,
    topic                 VARCHAR(128) NOT NULL,
    partition_key         VARCHAR(256) NOT NULL,
    payload               JSONB NOT NULL,
    created_at            TIMESTAMP WITH TIME ZONE DEFAULT NOW()
);

CREATE INDEX idx_outbox_created
    ON outbox_events(created_at);

Fault Tolerance

Operational Failure Scenarios and Mitigations

The escalation platform maintains high availability across worker nodes, external telecom providers, persistent alert state, and clock synchronization. Redis acts as the timer execution index, while PostgreSQL remains the durable source for alert state and escalation due times:

Concern ScenarioSystem Solution Design
Timer Worker Node CrashesEscalation state and due times remain durable in PostgreSQL. Redis stores the execution index, workers claim timers with short leases, and reconciliation rebuilds missing timer entries after Redis loss or failover.
Redis Timer Index LossRebuild due timers from PostgreSQL next_action_at values, then resume execution with the same lease and idempotency controls.
Provider Request Timeout With Unknown Delivery StateKeep the notification attempt durable with its idempotency key, reconcile provider message status by provider message ID or callback, and retry only when the attempt remains unresolved.
Telecom Provider BlackoutsIntegrate redundant message gateways. Switch Twilio SMS to Vonage or MessageBird automatically on timeout or error.
PostgreSQL Primary FailurePromote the synchronous standby inside the failure domain, keep alert events durable in the replicated database and outbox, and allow the CDC pipeline to resume without losing accepted state.
Timer Resolution DriftAlign system clocks using NTP across all hosts, maintaining precision within 10 milliseconds.

1. Concurrent Acknowledgment Race Conditions ⭐

When primary and secondary responders acknowledge simultaneously, atomic Compare-And-Swap (CAS) updates guarantee deterministic state transitions:

SQL
-- Alice (primary) and Bob (secondary) both receive notifications
-- Alice clicks ACK at T=5:00.000 while Bob clicks ACK at T=5:00.050

-- Atomic Compare-And-Swap (CAS) update on PostgreSQL
UPDATE alerts 
SET status = 'acknowledged', 
    acknowledged_by = 'user-881', 
    acknowledged_at = NOW(),
    version = version + 1
WHERE alert_id = 'alert-9912' 
  AND status IN ('triggered', 'notified_primary', 'escalated_secondary', 'escalated_manager', 'escalated_vp', 'snoozed')
  AND version = 7;

-- Execution outcome evaluation:
-- Alice's transaction: expected state/version still matches -> rows_affected = 1. Success!
-- Bob's transaction: state or version has advanced -> rows_affected = 0.
--   Returns HTTP 409 Conflict: "Alert state or version changed by alice@company.com"
--   Escalation timers in Redis are immediately cancelled via ZREM

2. Cascading Alert Storm Defense ⭐

When a core database crashes, thousands of connection alerts can overwhelm responders. A multi-tier defense shields engineers from notification fatigue:

YAML
alert_storm_mitigation_hierarchy:
  stage_1_ingestion_deduplication:
    mechanism: "Deterministic SHA-256 fingerprint from normalized team_id + alert_name + service + severity + priority + sorted labels"
    action: "If an active alert exists with that key and priority, increment the occurrence counter and merge the payload instead of paging"
    storage: "Redis key alert:dedup:{team_id}:{fingerprint} with a 5 minute TTL as an advisory cache. A Redis hit only accelerates the lookup and never suppresses an alert without authoritative PostgreSQL validation. The PostgreSQL active-alert uniqueness constraint remains authoritative"

  stage_2_sliding_window_grouping:
    mechanism: "Sliding 5 minute window"
    action: "Aggregate related alarms across services into single consolidated incidents"

  stage_3_dependency_correlation:
    mechanism: "Upstream infrastructure dependency graph"
    action: "If the core database cluster is confirmed down, correlate downstream service symptoms into the same incident and suppress only alerts that are proven derivative signals"

  stage_4_personal_rate_limiting:
    mechanism: "Token bucket per responder device"
    limit: "Maximum 10 pages per hour per responder"
    overflow: "Excess non-P0 alerts queue into hourly email summaries"

3. Handling Schedule Gaps and Shift Transitions ⭐

If team transitions leave an unassigned calendar gap where no responder is active, automated background scans and fallback ladders ensure alarms are never dropped:

YAML
schedule_gap_detection_and_fallback:
  scenario:
    shift_end: "Alice shift ends Saturday 18:00 UTC"
    shift_start: "Bob shift starts Sunday 09:00 UTC"
    gap_duration: "15-hour unassigned coverage gap"
    alert_event: "Alert fires Saturday 23:00 UTC"

  preemptive_health_checks:
    schedule_audit_cron: "Daily background cron scans all 50,000 team schedules 7 days into the future"
    notification: "Alert team administrators immediately if unassigned slots are detected"

  runtime_fallback_ladder:
    tier_1: "Route alert to secondary on call engineer in the rotation layer"
    tier_2: "Escalate to team designated Engineering Manager"
    tier_3: "Escalate to VP of Engineering as root safety net"
    rule: "No accepted alert may be silently discarded. If no responder is available, the fallback ladder routes it to team management and records the condition"

Additional Considerations

1. Interactive Voice Response (IVR) Verification ⭐

Voice calls often provide stronger attention than app notifications during off hours incidents, but device and focus settings can still suppress or redirect calls. Twilio Voice workflows execute synthesized speech and DTMF tone capture:

YAML
interactive_voice_response_ivr_workflow:
  objective: "Voice calls provide an additional high attention channel during night shifts, subject to device and focus settings"

  call_flow_steps:
    step_1_initiation:
      action: "Escalation worker dispatches call request to Twilio REST API voice endpoint"
      destination: "+15550199"

    step_2_text_to_speech:
      twiml_response: |
        <Response>
          <Say voice="Polly.Joanna">Critical Alert: Payment service CPU usage exceeds 90 percent.</Say>
          <Gather numDigits="1" action="/api/v1/webhooks/twilio/gather" method="POST" timeout="10">
            <Say>Press 1 to acknowledge this alert. Press 2 to escalate to secondary.</Say>
          </Gather>
        </Response>

    step_3_dtmf_tone_capture:
      key_1: "Triggers Twilio webhook, executes atomic CAS update to acknowledge alert, and cancels Redis timer"
      key_2_or_hangup: "Triggers immediate escalation webhook, paging secondary on call responder"

    step_4_failure_recovery:
      retry_policy: "Busy signal or unanswered call retries within 60 seconds"
      max_attempts: "3 call attempts per escalation level before escalating"

2. Schedule Overrides and Shift Swaps

On-call engineers frequently need to cover slots for peers temporarily. The routing engine handles overrides dynamically with top priority over base rotations. Overlapping override windows for the same schedule are rejected during write validation so the effective responder remains deterministic:

YAML
schedule_override_resolution:
  scenario: "Alice is on call but requires coverage Tuesday from 14:00 to 16:00 UTC, swapping with Bob"

  data_model:
    base_rotations: "on_call_schedules table defines recurring weekly or daily participant rotations"
    overrides: "schedule_overrides table defines time bounded replacements with target and covering user IDs"

  resolution_hierarchy:
    step_1: "Check schedule_overrides for active records matching target timestamp (highest priority)"
    step_2: "If no active override exists, query on_call_schedules for regular rotation assignee"
    step_3: "If schedule is unassigned, trigger fallback escalation ladder to team management"

3. Achieving Five Nines (99.999% SLO) ⭐

Limiting downtime to under five minutes per year requires active active multi region deployment, multi-carrier redundancy, and continuous chaos testing:

YAML
five_nines_availability_architecture:
  target_slo: "99.999% availability (maximum 5.26 minutes of downtime per calendar year)"

  architectural_pillars:
    active_active_multi_region:
      topology: "Dual-region hot deployment across us-east-1 and eu-west-1 with active active stateless ingress and workers"
      database: "Each team schedule has a home write region with synchronous replication within its failure domain and tested cross-region replication for disaster recovery"
      fencing: "A monotonically increasing schedule ownership epoch fences stale regional writers after failover so only the current owner can commit schedule changes"

    provider_redundancy:
      push_gateways: "Dual delivery pipelines across Apple APNs and Google FCM"
      sms_gateways: "Primary dispatch via Twilio with automatic failover to MessageBird or Sinch on latency spike"
      voice_calls: "Primary routing via Twilio Voice with automatic failover to Vonage or Nexmo"

    chaos_engineering:
      validation: "Automated chaos drills periodically sever regional links, fail the active schedule region, and simulate telecom blackouts to verify alert continuity and failover behavior"

4. Maintenance Silences and Scheduled Suppression

Planned deployments and database migrations should not page on call engineers. A silence rule matches alert labels such as service, cluster, and alert name to suppress notifications during a scheduled window while continuing to record events for auditing. Active silence rules can be cached by team in Redis with a versioned policy key, while PostgreSQL remains authoritative. Pending escalation timers recheck the current silence state before dispatch, so creating a silence also protects against notifications already queued for execution. The alert remains stored and observable throughout the silence.

HTTP
POST /api/v1/silences
Content-Type: application/json
Authorization: Bearer <auth_token>

{
  "matchers": [
    {
      "label": "service",
      "value": "payment-service"
    }
  ],
  "starts_at": "2026-03-14T22:00:00Z",
  "ends_at": "2026-03-14T23:00:00Z",
  "comment": "Scheduled database migration"
}
JSON
{
  "silence_id": "sil-881",
  "status": "scheduled",
  "starts_at": "2026-03-14T22:00:00Z",
  "ends_at": "2026-03-14T23:00:00Z",
  "version": 1,
  "created_at": "2026-03-14T21:45:00Z"
}

5. Webhook Authentication and Replay Protection

Alert ingestion must assume that external monitoring systems can be delayed, duplicated, or attacked. Each integration uses a signed request or mTLS where supported, includes a source event identifier and timestamp, and rejects requests outside an allowed clock skew window. The ingestion service stores the source event identifier or idempotency key so retried webhooks return the original alert result without creating another incident. Provider delivery callbacks and responder channel webhooks are authenticated and matched using provider message IDs or channel specific request identifiers before changing notification state. Per source rate limits and payload size limits protect the ingestion path during provider malfunction or abuse.

Related Problems and Concepts

Explore related architectures and distributed systems foundations that connect directly to this design:

Interview Walkthrough

Follow this structured pacing guide during a 45 minute system design interview:

  • 25 minute core design pacing guide

    Focus on the core scheduling and escalation loops first, expanding into multi region disaster recovery if interviewing at staff level.

    • Walk the core loop: ingest alerts, deduplicate by fingerprint, match on call schedules, dispatch notifications, monitor acknowledgment timeouts, and escalate unacknowledged incidents (5 min)
    • Model on call schedules as time bounded rotations with override and shift handoff semantics (6 min)
    • Design an escalation ladder: push notifications, SMS texts, and interactive voice calls with per step timeouts (5 min)
    • Deduplicate alerts by fingerprint (service and error signature) to prevent notification fatigue (5 min)
    • Back timeouts with a durable timer engine utilizing Redis Sorted Sets or workflow schedulers (4 min)
  • Clarify the core loop: ingest alerts, match the active on call schedule, notify primary responders, and start escalation timers.
  • Model on call schedules as time bounded rotations with override support for temporary swaps and holiday coverage.
  • Design an escalation ladder: push notifications, SMS texts, and interactive voice calls with press-to-acknowledge DTMF confirmation and configurable per step timeouts.
  • Deduplicate alerts by fingerprint (combining service, cluster, and error signature) to prevent notification storms during cascading failures.
  • Use a durable timer engine utilizing Redis Sorted Sets so escalation steps fire reliably even if worker nodes restart.
  • Track delivery and acknowledgment state per channel because an unacknowledged push should trigger the next escalation tier automatically.
  • For 99.999% availability, describe an active active multi region deployment with explicit schedule ownership and external paging provider failover across multiple push, SMS, and voice providers.
  • Highlight the common pitfall where synchronous phone call initiation blocks the alert ingestion path, emphasizing that the hot ingestion path must enqueue and return HTTP 202 immediately.

Engineering Trade-offs

1. Timer Engine Implementations

This comparison evaluates timer precision, persistence durability, database load, and implementation complexity across alternative architectures.

ApproachPrecisionDurabilityDB LoadComplexity
Redis Sorted Set ⭐~100 ms target with 100 ms pollingVolatile execution index. PostgreSQL stores durable escalation state and timers can be rebuilt after Redis lossExtremely LowLow
Database PollingLow (bounded by polling frequency)HighHigh (frequent sequential queries)Low
Kafka Delayed MessagesLow (requires bucketed queues)HighNoneHigh

2. Notification Delivery Mechanisms: Push vs SMS vs Voice

Different communication channels exhibit distinct latency, reliability, battery, and financial trade offs:

PatternLatencyReliabilityBattery and Resource ImpactOperational Overhead
Push Notifications (APNs / FCM)Sub-second to 5 secondsHigh for awake devices, though standard alerts can be silenced by OS Do Not Disturb unless configured with critical alert entitlementsNegligible application battery impact because the mobile OS manages push deliveryRequires managing Apple APNs certificates and Google FCM credentials with automated token refresh
SMS Messaging (Twilio / MessageBird)5 to 15 secondsHigh when cellular service is available and does not require mobile data connectivity, but remains vulnerable to carrier outages and regional filteringZero mobile app overhead because cellular basebands receive standard SMS transmissionsPer-message carrier fees, regulatory 10DLC registration, and multi vendor fallback routing
Interactive Voice Calls (Twilio Voice)10 to 30 seconds to connectHigh attention value because a voice call can alert separately from app notifications, while device call and focus settings can still suppress or redirect itStandard incoming cellular call overheadHighest unit cost, voice trunk management, and synthesized text-to-speech rendering

3. Alert Grouping Latency vs Notification Promptness

This comparison balances alert aggregation delay against prompt notification during critical incidents.

Grouping WindowNotification LatencyAlert Fatigue SuppressionCascading Storm ResilienceRecommended Context
Zero Window (Immediate Dispatch)Zero additional delay (under 5 seconds)None because every threshold breach dispatches an isolated pageExtremely fragile because downstream engineers receive hundreds of duplicate pages during outagesP0 and emergency P1 infrastructure events where immediate human attention is paramount
30 to 60 Seconds (Microbatching)30 to 60 secondsModerate by consolidating bursts of identical alarms into single incident headersHigh because it absorbs initial flash floods from service restarts and rolling deploymentsDefault for non urgent application health alerts. Urgent P0 and P1 pages bypass the grouping delay
5 Minutes (Sliding Window)Up to 5 minutesMaximum by aggregating flapping signals and multi cluster dependency cascades into one ticketNear total suppression of noise across distributed service graphsOngoing correlation for P2 and P3 alerts after the initial notification

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...