Interview Setup
Interview Prompt
Design an on call escalation system like PagerDuty. Ingest alerts from monitoring systems, route to the right on call engineer via schedules and escalation policies, notify across channels, and escalate if unacknowledged.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Alert volume during steady state vs cascading outage peak? | A peak rate of 100,000 alerts per minute during site wide incidents necessitates aggressive grouping and deduplication because the system cannot page responders for every raw event. |
| Acknowledgment timeout and escalation chain depth? | A standard progression starts with a 5 minute timeout for the primary on call, advances to a 10 minute timeout for the secondary on call, and then escalates to the engineering manager, which drives the state machine timers. |
| Notification channels and delivery guarantees? | SMS, voice, and push notifications have varying latency characteristics, requiring a durable queue to buffer dispatches if downstream APIs experience outages. |
| Schedule complexity including weekly rotations, overrides, and follow the sun coverage? | Timezone conversions and daylight saving transitions invalidate naive cron-based schedule calculations. |
Scope
In scope
- Escalation policy state machine
- Schedule management
- Acknowledgment timeout
- Multi-channel alert
- Capacity estimation with shown math
Out of scope (state explicitly)
- Full incident management war-room UI
- Building a single monitoring tool vs integrating with upstream systems
- Log search and distributed trace analysis
Functional Requirements
System Scope and Core Capabilities
The system manages the end to end incident lifecycle. It ingests alarms from monitoring backends, evaluates on-call rotation schedules, enforces escalation policies, and dispatches multi channel notifications until an engineer acknowledges.
Interview perspective: Highlighting alert fatigue and deduplication early is essential because firing hundreds of pages for a single root cause represents a major system design flaw.
- Alert Ingestion: Accept incoming alarm webhooks from Prometheus Alertmanager, Datadog, CloudWatch, and custom monitoring backends with signed authentication, replay protection, idempotency, and per-source rate limits.
- On-Call Rotations: Define calendar schedules including weekly, daily, and follow the sun rotations with primary and secondary responders.
- Escalation Chains: Automatically escalate unacknowledged alerts to secondary engineers and engineering management after configurable timeouts.
- Multi Channel Alerts: Page responders across mobile push notifications, SMS text messages, interactive voice calls, email, and Slack.
- Acknowledge and Resolve: Expose explicit acknowledgment APIs that mark alerts acknowledged or resolved and terminate downstream escalation timers.
- Alert Aggregation: Correlate related alerts into unified incidents to prevent notification fatigue during cascading failures.
- Maintenance Silencing: Support scheduled maintenance windows that suppress active pages while preserving audit events.
- Incident Timeline: Record complete event histories from initial trigger to final resolution for mean time to acknowledge (MTTA) and mean time to resolve (MTTR) analytics.
Non-Functional Requirements
Reliability and Operational Objectives
Notification delivery within thirty seconds and no intentional loss after an accepted alert is durably persisted form the core service level objectives. The reliability of the paging path itself is paramount because failure in the alert infrastructure blinds engineering teams to production outages.
- Five Nines Availability: Deliver 99.999% availability for the paging path because monitoring failures should not prevent alert dispatch after an alert has been durably accepted.
- Low Propagation Latency: Route P0 and P1 alerts from initial webhook ingestion to the first responder notification attempt within 30 seconds at p99. Lower priority alerts can use the configured grouping window.
- Durable Notification Delivery: Provide at least once provider dispatch attempts after an alert is durably accepted, with provider delivery tracked separately because external channels can fail or delay indefinitely.
- Cascading Outage Scale: Ingest and process more than 100,000 alerts per minute during severe site wide infrastructure collapses.
- Alert Deduplication: Maintain a 5 minute deduplication window per incident fingerprint and priority to suppress redundant notifications without requiring distributed transport exactly once semantics. P0 and P1 alerts bypass the grouping delay but still use the same idempotency controls.
- Immutable Audit Logging: Persist an append only audit trail capturing alert timestamps, notification dispatches, and responder acknowledgment actions.
Capacity Estimations
Traffic Projections and Notification Fan Out
While steady state alert volume is modest compared with high throughput metric streams, cascading outages create massive transient bursts. Multi channel fan out further multiplies notifications across push, SMS, and voice channels.
| Metric | Calculation | Value |
|---|---|---|
| Alerts / day (normal) | Given (steady state operations) | 10,000 |
| Alerts / sec (avg) | 10,000 ÷ 86,400 | ~0.12 |
| Alerts / sec (cascade peak) | 100K alerts/min ÷ 60, before grouping | ~1,667 ungrouped alerts per second. The design target is well under 1 effective page per second for a single correlated storm after deduplication and grouping |
| Active On-Call Schedules | Derived | 50,000 teams |
| Initial notifications / day | 10K alerts x ~2.5 channels avg | ~25,000 |
| Escalations / day | ~5% of alerts reach L1+ | ~500 |
Scale Derivations and Peak Sizing
System components must handle both steady state processing and severe burst conditions. The peak example is a capacity target for a modeled correlated outage, not a guarantee that every 100,000 alerts per minute event collapses to fewer than one page per second:
# Capacity Math and Peak Volume Projections
steady_state_traffic:
daily_alerts: 10000 alerts per day
average_ingestion_qps: 10000 / 86400 ≈ 0.12 alerts per second
active_schedules: 50000 teams
daily_notifications: 10000 alerts * 2.5 channels average ≈ 25000 notifications per day
daily_escalations: ~5% of alerts reach Level 1 or higher ≈ 500 escalations per day
cascading_outage_peak:
peak_ingestion_rate: 100000 alerts per minute (≈ 1667 alerts per second) before grouping
deduplicated_effective_pages: Under 1 page per second for a modeled correlated storm after fingerprint deduplication and sliding window grouping
telecom_concurrency: High outbound concurrency required with Twilio and MessageBird to prevent queue backlogs during emergency voice dispatchArchitecture Diagram
End to End Alert Flow Architecture
The architecture receives alert webhooks from external monitoring providers. It authenticates and normalizes each payload, deduplicates repeated signals, evaluates the active on call schedule, and dispatches multi channel notifications with automated escalation when acknowledgments do not arrive.
Component Deep Dives
1. Stateful Escalation Policy State Machine
Escalating alarms through responder levels must occur reliably regardless of human intervention. The state machine transitions alerts through deterministic states with persisted audit timestamps:
team_alert_lifecycle:
team_id: "backend-infra"
escalation_policy:
level_0:
target: "Primary on call (Alice)"
channels: ["push", "sms"]
ack_timeout_seconds: 300 # 5 minutes
level_1:
target: "Secondary on call (Bob)"
channels: ["push", "sms", "voice_call"]
ack_timeout_seconds: 600 # 10 minutes
level_2:
target: "Engineering Manager (Carol)"
channels: ["voice_call"]
ack_timeout_seconds: 900 # 15 minutes
level_3:
target: "VP of Engineering (Dave)"
channels: ["voice_call", "sms"]
execution_trace:
t0_00: "Alert ingested and dispatched to Alice via Push and SMS"
t5_00: "No ACK received from Alice, escalating to Level 1 and dispatching to Bob via Push, SMS, and Voice"
t15_00: "No ACK received from Bob, escalating to Level 2 and dispatching to Carol via Voice"
t30_00: "No ACK received from Carol, escalating to Level 3 and dispatching to Dave via Voice and SMS"
termination_condition: "Immediate cancellation of pending timers upon receipt of valid responder acknowledgment or a resolved source event"2. Escalation Timer Implementation with Redis Sorted Sets ⭐
Scanning database tables on a periodic interval creates heavy query overhead and high contention. We use Redis Sorted Sets as a fast execution index with approximately 100 millisecond polling and short leases. PostgreSQL remains the durable source for escalation state and next action times:
# Redis Sorted Set Escalation Timer Implementation
# Key: escalation_timers
# Score: Unix epoch timestamp when the escalation action is due
# Member: {alert_id}:{escalation_level}
# PostgreSQL is the durable source of escalation state and next_action_at timestamps.
# Redis is the execution index and can be rebuilt after Redis loss.
def enqueue_escalation_timer(redis_client, alert_id: str, level: int, delay_seconds: int):
fire_timestamp = current_epoch_time() + delay_seconds
redis_client.zadd("escalation_timers", {f"{alert_id}:{level}": fire_timestamp})
def worker_poll_loop(redis_client):
while True:
now = current_epoch_time()
due_timers = redis_client.zrangebyscore("escalation_timers", 0, now, start=0, num=100)
for item in due_timers:
# Keep the timer until dispatch succeeds. A short lease prevents
# competing workers from processing the same timer concurrently.
lease_key = "timer:claim:" + item
owner_token = random_token()
claimed = redis_client.set(lease_key, owner_token, nx=True, px=5000)
if not claimed:
continue
try:
alert_id, level_str = item.split(":")
level = int(level_str)
# Revalidate the alert state in PostgreSQL before the provider call.
# If the alert was acknowledged, resolved, or rescheduled, do not page.
if not is_escalation_still_due(alert_id, level):
redis_client.zrem("escalation_timers", item)
continue
# The dispatch transaction creates or reuses a stable notification_attempt
# idempotency key and persists the next_action_at transition in PostgreSQL.
# Provider calls use that same logical idempotency identity where supported.
trigger_escalation_idempotently(alert_id, level, stable_notification_key(alert_id, level))
redis_client.zrem("escalation_timers", item)
if has_next_escalation_level(alert_id):
next_level = get_next_level(alert_id)
timeout = get_level_timeout(alert_id, next_level)
enqueue_escalation_timer(redis_client, alert_id, next_level, timeout)
finally:
release_lease_if_owner(redis_client, lease_key, owner_token)
sleep(0.1) # Approximately 100 ms scheduling precision target. Local elapsed-time checks use a monotonic clock.3. Alert Grouping and Noise Reduction
During cascading infrastructure failures, hundreds of downstream microservices fire alerts simultaneously. A multi layer grouping and deduplication pipeline merges related alerts into a single actionable incident:
alert_grouping_and_deduplication:
problem_context:
without_grouping: "Cascading outage causes 500 downstream microservices to fire alerts, overwhelming on call with alarm fatigue"
sliding_window_strategy:
window_duration_seconds: 300 # 5 minute sliding correlation window
emergency_bypass: "P0 and P1 notifications dispatch immediately"
correlation_after: "Correlation continues in parallel so later related alerts can be consolidated"
correlation_keys:
- "cluster"
- "environment"
- "alert_name"
- "dependency_id"
note: "Service remains an incident attribute, but dependency-aware correlation allows related alerts from multiple services to be grouped"
owner_lifecycle: "The canonical owner alert carries escalation state. If it resolves while related alerts remain active, the grouper elects another active alert as owner in the same transaction before the next escalation action is scheduled."
priority_rule: "The correlated incident inherits the highest active member priority. A transition to P0 or P1 uses the urgent path immediately while the correlation window continues."
outcome:
incident_title: "Database connection timeout (503 occurrences across 50 downstream microservices)"
action: "Dispatch one consolidated page to the on call responder. Acknowledging the incident silences grouped notifications while preserving all underlying alert events"4. Event Bus Architecture and Kafka Partitioning
Decoupling synchronous webhook ingestion from asynchronous schedule matching and notification dispatch prevents downstream telecom latencies from blocking incoming alarms:
topics:
alert-events:
partitions: 32
partition_key: dedup_key # preserves ordering for repeated events with the same fingerprint
retention_days: 7
grouped-alerts:
partitions: 32
partition_key: incident_correlation_key # brings alerts for the same incident onto one correlation stream
retention_days: 7
payload_fields: [incident_correlation_key, incident_owner_alert_id, grouped_alert_count]
urgent-alerts:
partitions: 32
partition_key: incident_correlation_key
retention_days: 7
escalation-actions:
partitions: 32
partition_key: alert_id
retention_days: 7
notification-events:
partitions: 32
partition_key: alert_id
retention_days: 7
producers: [Notification Router, Provider Callback Handler]
publication_rule: Notification state and each notification event are written through the transactional outbox before CDC publication
cluster_configuration:
replication_factor: 3
min_insync_replicas: 2
producer:
service: Alert Ingestion Service
execution_point: PostgreSQL transaction writes alert state and the outbox row. CDC publishes the event asynchronously
payload_schema:
alert_id: "8a219b1b-640a-4289-9812-42171542fca1"
dedup_key: "high_cpu_payment_us_east"
source_event_id: "evt_20260314_000001"
event_type: "firing"
idempotency_key: "prometheus:evt_20260314_000001"
severity: "critical"
service: "payment-service"
labels:
env: "production"
cluster: "us-east-1"
status: "triggered"
source: "prometheus"
alert_version: 7
priority: "P1"
timestamp: "2026-03-14T10:00:00Z"
consumer_groups:
alert_grouper:
consumes: alert-events
filter: "priority in {P2, P3}"
produces: grouped-alerts
state_store: Durable keyed correlation state with a changelog so a grouper restart does not lose an open five minute window
description: Coalesces exact duplicate source events, applies the initial 30 to 60 second notification grouping, continues broader incident correlation for five minutes, selects a canonical owner alert, and emits grouped incident state for downstream routing
schedule_router:
consumes: [grouped-alerts, urgent-alerts]
produces: escalation-actions
description: Resolves the effective team policy and on call schedule for the canonical incident owner, validates silence state, and creates or advances the escalation state machine
urgent_router:
consumes: alert-events
filter: "priority in {P0, P1}"
produces: urgent-alerts
description: Bypasses the grouping delay for urgent alerts, revalidates the active incident state, and emits only required state transitions to the immediate notification path
analytics_sink:
consumes: [alert-events, notification-events]
description: Streams alert state transitions and notification delivery events to ClickHouse for MTTA, MTTR, volume, and delivery analysis
notification_router:
consumes: escalation-actions
produces: notification-events
description: Creates idempotent notification attempts, dispatches to providers, and records provider callback state through the durable notification store
consumer_idempotency:
rule: Consumers key side effects by source_event_id and alert_version or by notification_attempt idempotency_key before committing a state transition
execution_paths:
synchronous_path: Webhook Ingest -> Authenticate and normalize -> Advisory dedup check in Redis -> Validate the active incident in PostgreSQL on cache hit or miss -> Persist PostgreSQL alert state and outbox row atomically -> HTTP 202 Accepted
asynchronous_path: PostgreSQL CDC/outbox -> alert-events -> Alert Grouper -> grouped-alerts -> Schedule Router -> escalation-actions -> Escalation Timers in Redis ZSET -> Notification Router -> provider callbacks -> notification-events
urgent_path: alert-events -> Urgent Router for P0/P1 -> urgent-alerts -> Schedule Router -> escalation-actions, while normal correlation continues independently
dead_letter_queue: "Failed events move to alert-events-dlq or grouped-alerts-dlq after 3 retries. Operators can replay a DLQ event after the root cause is fixed because the original source event ID remains the idempotency key. A separate lag alert pages if consumer lag exceeds 30 seconds."API Design
Core API Interface Signatures
Typed TypeScript interfaces define the programmatic contract between external monitoring systems, responder clients, and internal escalation workers:
// Core API Interface Contracts for On-Call Escalation Engine
interface IngestAlertRequest {
source: string;
alertName: string;
severity: "critical" | "warning" | "info";
priority: "P0" | "P1" | "P2" | "P3";
service: string;
labels: Record<string, string>;
description?: string;
dedupKey?: string; // Optional provider fingerprint. If absent, the service derives a canonical fingerprint from normalized alert fields.
sourceEventId: string;
eventType: "firing" | "resolved";
occurredAt: string;
}
interface IngestAlertResponse {
alertId: string;
status: "triggered" | "resolved" | "deduplicated";
acceptedAt: string;
}
interface AcknowledgeRequest {
responderId: string;
channel: "mobile" | "web" | "sms" | "voice";
expectedVersion: number;
}
interface AcknowledgeResponse {
alertId: string;
status: "acknowledged" | "already_acknowledged";
acknowledgedAt: string;
version: number;
}
interface ResolveRequest {
responderId: string;
resolutionCode: string;
expectedVersion: number;
}
interface ResolveResponse {
alertId: string;
status: "resolved";
resolvedAt: string;
version: number;
}
interface ActiveOnCallResponder {
userId: string;
email: string;
phone?: string;
layer: number;
rotationPriority: number;
scheduleVersion: number;
}
interface ActiveOnCallResponse {
teamId: string;
effectiveAt: string;
scheduleVersion: number;
coverageStatus: "covered" | "gap";
primary?: ActiveOnCallResponder;
secondary?: ActiveOnCallResponder;
}
interface ScheduleOverride {
targetUserId: string;
coveringUserId: string;
startTime: string;
endTime: string;
reason: string;
}
interface OverrideResponse {
overrideId: string;
status: "scheduled" | "active";
effectiveFrom: string;
effectiveUntil: string;
}
interface SnoozeRequest {
responderId: string;
durationSeconds: number;
expectedVersion: number;
}
interface SnoozeResponse {
alertId: string;
status: "snoozed";
snoozedUntil: string;
version: number;
}
interface SilenceMatcher {
label: string;
value: string;
}
interface CreateSilenceRequest {
teamId: string;
matchers: SilenceMatcher[];
startsAt: string;
endsAt: string;
comment?: string;
}
interface CreateSilenceResponse {
silenceId: string;
status: "scheduled" | "active";
startsAt: string;
endsAt: string;
version: number;
}
export interface OnCallEscalationService {
// Ingest incoming webhook alert from monitoring systems with deduplication
ingestAlert(payload: IngestAlertRequest): Promise<IngestAlertResponse>;
// Acknowledge an active alert, halting downstream escalation timers
acknowledgeAlert(alertId: string, request: AcknowledgeRequest): Promise<AcknowledgeResponse>;
// Resolve an active alert when operational health is restored
resolveAlert(alertId: string, request: ResolveRequest): Promise<ResolveResponse>;
// Snooze an active alert for a bounded period without resolving it
snoozeAlert(alertId: string, request: SnoozeRequest): Promise<SnoozeResponse>;
// Retrieve active on call responders for a team at a specific query timestamp
getActiveOnCall(teamId: string, queryTimestamp?: string): Promise<ActiveOnCallResponse>;
// Create a time bounded schedule override for shift coverage or swaps
createScheduleOverride(scheduleId: string, override: ScheduleOverride): Promise<OverrideResponse>;
// Create a time bounded maintenance silence for matching alert labels
createSilence(request: CreateSilenceRequest): Promise<CreateSilenceResponse>;
}1. Ingest Alert Webhook
External monitoring agents post alert payloads to the ingestion gateway:
POST /api/v1/alerts
Content-Type: application/json
Authorization: Bearer <integration_token>
Idempotency-Key: prometheus:evt_20260314_000001
X-Signature: sha256=<hmac>
{
"source": "prometheus",
"alert_name": "HighCPUUsage",
"severity": "critical",
"priority": "P1",
"service": "payment-service",
"labels": {
"env": "production",
"cluster": "us-east-1"
},
"description": "CPU usage > 90% for 5 consecutive minutes",
"dedup_key": "high_cpu_payment_us_east",
"source_event_id": "evt_20260314_000001",
"event_type": "firing",
"occurred_at": "2026-03-14T10:00:00Z"
}
Response: 202 Accepted
{
"alert_id": "8a219b1b-640a-4289-9812-42171542fca1",
"status": "triggered",
"accepted_at": "2026-03-14T10:00:00Z"
}2. Acknowledge Active Alert
Responders acknowledge alerts via mobile apps, SMS replies, or web dashboards to stop escalation timers:
POST /api/v1/alerts/8a219b1b-640a-4289-9812-42171542fca1/acknowledge
Content-Type: application/json
Authorization: Bearer <responder_token>
{
"responder_id": "user-881",
"channel": "mobile",
"expected_version": 7
}
Response: 200 OK
{
"alert_id": "8a219b1b-640a-4289-9812-42171542fca1",
"status": "acknowledged",
"acknowledged_at": "2026-03-14T10:05:12Z",
"version": 8
}3. Snooze Active Alert
Responders can temporarily suppress repeat notifications without resolving the incident. The operation uses optimistic version checking and reschedules the next escalation action after the snooze period.
POST /api/v1/alerts/8a219b1b-640a-4289-9812-42171542fca1/snooze
Content-Type: application/json
Authorization: Bearer <responder_token>
{
"responder_id": "user-881",
"duration_seconds": 3600,
"expected_version": 7
}
Response: 200 OK
{
"alert_id": "8a219b1b-640a-4289-9812-42171542fca1",
"status": "snoozed",
"snoozed_until": "2026-03-14T11:05:12Z",
"version": 8
}4. Retrieve Active On-Call Schedule
The routing service inspects active schedule rosters for a team at a specific point in time:
GET /api/v1/schedules/team-backend-infra/on-call?at=2026-03-14T10:00:00Z
Authorization: Bearer <responder_token>
Response: 200 OK
{
"team_id": "team-backend-infra",
"primary": {
"user_id": "user-881",
"email": "alice@company.com",
"phone": "+15550199",
"layer": 0,
"rotation_priority": 100,
"schedule_version": 42
},
"secondary": {
"user_id": "user-902",
"email": "bob@company.com",
"phone": "+15550212",
"layer": 1,
"rotation_priority": 90,
"schedule_version": 42
},
"effective_at": "2026-03-14T10:00:00Z",
"schedule_version": 42,
"coverage_status": "covered"
}Standard HTTP Error Responses
Standardized HTTP error responses communicate validation failures, authentication issues, and optimistic locking race conditions:
400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
422 Unprocessable Entity: password does not meet security requirements or MFA verification is required
423 Locked: account temporarily locked after repeated failed login attempts
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpoint
401 Unauthorized: webhook signature, integration token, or authentication credentials are invalid
404 Not Found: requested alert_id, team_id, or schedule_id does not exist
409 Conflict: alert state, expected version, or schedule override is stale because another operation changed it first
422 Unprocessable Entity: invalid escalation policy level, schedule interval, or unsupported notification channel configuration
429 Too Many Requests: source or responder has exceeded the configured ingestion or notification rate limit
503 Service Unavailable: a required scheduling, queue, or database dependency is temporarily unavailable. Previously accepted alerts remain durable while downstream provider failures are retried asynchronouslyData Model
PostgreSQL Database Schema
PostgreSQL provides durable transactional storage for alert states, pinned escalation policy versions, normalized schedule assignments, notification delivery state, incident event history, and temporary calendar overrides:
-- Authoritative alert and incident state
CREATE TABLE alerts (
alert_id UUID PRIMARY KEY,
source VARCHAR(64) NOT NULL,
latest_source_event_id VARCHAR(256),
source_state VARCHAR(16) NOT NULL DEFAULT 'firing', -- firing or resolved
last_source_event_at TIMESTAMP WITH TIME ZONE,
dedup_key VARCHAR(256) NOT NULL,
incident_correlation_key VARCHAR(256),
incident_owner BOOLEAN NOT NULL DEFAULT TRUE,
alert_name VARCHAR(256) NOT NULL,
severity VARCHAR(20) NOT NULL,
priority VARCHAR(4) NOT NULL,
service VARCHAR(128) NOT NULL,
labels JSONB,
description TEXT,
status VARCHAR(32) NOT NULL DEFAULT 'triggered', -- triggered, notified_primary, escalated_secondary, escalated_manager, escalated_vp, acknowledged, snoozed, resolved
escalation_level INT NOT NULL DEFAULT 0,
team_id UUID NOT NULL,
policy_id UUID,
policy_version BIGINT,
occurrence_count BIGINT NOT NULL DEFAULT 1,
last_seen_at TIMESTAMP WITH TIME ZONE NOT NULL DEFAULT NOW(),
acknowledged_by VARCHAR(256),
acknowledged_at TIMESTAMP WITH TIME ZONE,
resolved_by VARCHAR(256),
resolved_at TIMESTAMP WITH TIME ZONE,
snoozed_until TIMESTAMP WITH TIME ZONE,
next_action_at TIMESTAMP WITH TIME ZONE,
version BIGINT NOT NULL DEFAULT 0,
created_at TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
updated_at TIMESTAMP WITH TIME ZONE DEFAULT NOW()
);
-- The database is authoritative for active incident deduplication by team, fingerprint, and priority.
-- Redis is an advisory accelerator and may be rebuilt after cache loss.
CREATE UNIQUE INDEX idx_alerts_active_dedup
ON alerts(team_id, dedup_key, priority)
WHERE status <> 'resolved';
-- Related alerts can be correlated into one incident for paging while each source alert remains auditable.
-- The canonical owner alert carries the escalation state and next action for that correlated incident.
CREATE UNIQUE INDEX idx_alerts_active_incident_owner
ON alerts(team_id, incident_correlation_key)
WHERE status <> 'resolved' AND incident_owner = TRUE AND incident_correlation_key IS NOT NULL;
CREATE INDEX idx_alerts_active_team
ON alerts(team_id, status, updated_at)
WHERE status <> 'resolved';
CREATE INDEX idx_alerts_next_action
ON alerts(next_action_at)
WHERE next_action_at IS NOT NULL AND status NOT IN ('acknowledged', 'resolved');
CREATE INDEX idx_alerts_status
ON alerts(status);
-- Immutable state transition history for incident timelines and MTTA / MTTR reconstruction.
CREATE TABLE alert_events (
event_id UUID PRIMARY KEY,
alert_id UUID NOT NULL REFERENCES alerts(alert_id),
source VARCHAR(64),
source_event_id VARCHAR(256),
event_type VARCHAR(32) NOT NULL,
actor_id VARCHAR(256),
channel VARCHAR(32),
occurred_at TIMESTAMP WITH TIME ZONE,
state_version BIGINT,
payload JSONB,
created_at TIMESTAMP WITH TIME ZONE DEFAULT NOW()
);
CREATE UNIQUE INDEX idx_alert_events_source_event
ON alert_events(source, source_event_id)
WHERE source IS NOT NULL AND source_event_id IS NOT NULL;
CREATE INDEX idx_alert_events_alert_time
ON alert_events(alert_id, created_at);
-- Provider callbacks are normalized into notification-events so analytics can measure dispatch, delivery, and failure independently of alert ingestion.
-- Durable notification state and provider reconciliation data.
CREATE TABLE notification_attempts (
attempt_id UUID PRIMARY KEY,
alert_id UUID NOT NULL REFERENCES alerts(alert_id),
escalation_level INT NOT NULL,
recipient_user_id VARCHAR(128) NOT NULL,
channel VARCHAR(32) NOT NULL,
idempotency_key VARCHAR(256) NOT NULL UNIQUE,
provider VARCHAR(64),
provider_message_id VARCHAR(256),
status VARCHAR(32) NOT NULL,
attempt_count INT NOT NULL DEFAULT 1,
next_retry_at TIMESTAMP WITH TIME ZONE,
sent_at TIMESTAMP WITH TIME ZONE,
delivered_at TIMESTAMP WITH TIME ZONE,
acknowledged_at TIMESTAMP WITH TIME ZONE,
created_at TIMESTAMP WITH TIME ZONE DEFAULT NOW()
);
CREATE INDEX idx_notification_attempts_alert
ON notification_attempts(alert_id, escalation_level, channel);
CREATE INDEX idx_notification_attempts_provider_message
ON notification_attempts(provider, provider_message_id)
WHERE provider_message_id IS NOT NULL;
-- Provider callback handlers update notification_attempts and write the corresponding notification event to the outbox in the same transaction.
-- CDC publishes notification-events asynchronously after the state transition commits.
-- Versioned escalation policies. Active incidents retain the policy version selected when routed.
CREATE TABLE escalation_policies (
policy_id UUID PRIMARY KEY,
team_id UUID NOT NULL,
version BIGINT NOT NULL,
levels JSONB NOT NULL, -- [{level: 0, wait_min: 5, channels: ["push", "sms"]}, ...]
status VARCHAR(16) NOT NULL DEFAULT 'active',
created_at TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
UNIQUE (team_id, version)
);
-- Base recurring schedule definition.
CREATE TABLE on_call_schedules (
schedule_id UUID PRIMARY KEY,
team_id UUID NOT NULL,
rotation_type VARCHAR(20) NOT NULL,
recurrence_rule JSONB NOT NULL,
participants JSONB NOT NULL,
timezone VARCHAR(64) DEFAULT 'UTC',
schedule_version BIGINT NOT NULL DEFAULT 1,
updated_at TIMESTAMP WITH TIME ZONE DEFAULT NOW()
);
-- Materialized time bounded assignments support indexed point in time schedule lookup.
CREATE TABLE schedule_assignments (
assignment_id UUID PRIMARY KEY,
schedule_id UUID NOT NULL REFERENCES on_call_schedules(schedule_id),
team_id UUID NOT NULL,
user_id VARCHAR(128) NOT NULL,
layer INT NOT NULL,
priority INT NOT NULL,
start_time TIMESTAMP WITH TIME ZONE NOT NULL,
end_time TIMESTAMP WITH TIME ZONE NOT NULL,
schedule_version BIGINT NOT NULL,
CHECK (end_time > start_time)
);
CREATE INDEX idx_schedule_assignments_lookup
ON schedule_assignments(team_id, layer, priority, start_time DESC);
-- Schedule overrides define temporary coverage changes. Overlap validation is serialized per schedule so concurrent writes cannot create ambiguous active overrides.
CREATE TABLE schedule_overrides (
override_id UUID PRIMARY KEY,
schedule_id UUID NOT NULL REFERENCES on_call_schedules(schedule_id),
target_user_id VARCHAR(128) NOT NULL,
covering_user_id VARCHAR(128) NOT NULL,
priority INT NOT NULL DEFAULT 100,
reason TEXT NOT NULL,
start_time TIMESTAMP WITH TIME ZONE NOT NULL,
end_time TIMESTAMP WITH TIME ZONE NOT NULL,
created_by VARCHAR(128) NOT NULL,
created_at TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
CHECK (end_time > start_time),
CHECK (target_user_id <> covering_user_id)
);
CREATE INDEX idx_schedule_overrides_lookup
ON schedule_overrides(schedule_id, start_time, end_time);
-- The scheduling service validates overlapping windows inside a serializable transaction or per-schedule lock.
-- Maintenance silence rules are persisted so ingestion and routing workers share the same suppression policy.
CREATE TABLE silences (
silence_id UUID PRIMARY KEY,
team_id UUID NOT NULL,
matchers JSONB NOT NULL,
version BIGINT NOT NULL DEFAULT 1,
starts_at TIMESTAMP WITH TIME ZONE NOT NULL,
ends_at TIMESTAMP WITH TIME ZONE NOT NULL,
created_by VARCHAR(128) NOT NULL,
comment TEXT,
created_at TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
CHECK (ends_at > starts_at)
);
CREATE INDEX idx_silences_active_window
ON silences(team_id, starts_at, ends_at);
-- Transactional outbox. Alert state and this row commit atomically. CDC publishes the row asynchronously.
CREATE TABLE outbox_events (
event_id UUID PRIMARY KEY,
aggregate_type VARCHAR(64) NOT NULL,
aggregate_id UUID NOT NULL,
topic VARCHAR(128) NOT NULL,
partition_key VARCHAR(256) NOT NULL,
payload JSONB NOT NULL,
created_at TIMESTAMP WITH TIME ZONE DEFAULT NOW()
);
CREATE INDEX idx_outbox_created
ON outbox_events(created_at);
Fault Tolerance
Operational Failure Scenarios and Mitigations
The escalation platform maintains high availability across worker nodes, external telecom providers, persistent alert state, and clock synchronization. Redis acts as the timer execution index, while PostgreSQL remains the durable source for alert state and escalation due times:
| Concern Scenario | System Solution Design |
|---|---|
| Timer Worker Node Crashes | Escalation state and due times remain durable in PostgreSQL. Redis stores the execution index, workers claim timers with short leases, and reconciliation rebuilds missing timer entries after Redis loss or failover. |
| Redis Timer Index Loss | Rebuild due timers from PostgreSQL next_action_at values, then resume execution with the same lease and idempotency controls. |
| Provider Request Timeout With Unknown Delivery State | Keep the notification attempt durable with its idempotency key, reconcile provider message status by provider message ID or callback, and retry only when the attempt remains unresolved. |
| Telecom Provider Blackouts | Integrate redundant message gateways. Switch Twilio SMS to Vonage or MessageBird automatically on timeout or error. |
| PostgreSQL Primary Failure | Promote the synchronous standby inside the failure domain, keep alert events durable in the replicated database and outbox, and allow the CDC pipeline to resume without losing accepted state. |
| Timer Resolution Drift | Align system clocks using NTP across all hosts, maintaining precision within 10 milliseconds. |
1. Concurrent Acknowledgment Race Conditions ⭐
When primary and secondary responders acknowledge simultaneously, atomic Compare-And-Swap (CAS) updates guarantee deterministic state transitions:
-- Alice (primary) and Bob (secondary) both receive notifications
-- Alice clicks ACK at T=5:00.000 while Bob clicks ACK at T=5:00.050
-- Atomic Compare-And-Swap (CAS) update on PostgreSQL
UPDATE alerts
SET status = 'acknowledged',
acknowledged_by = 'user-881',
acknowledged_at = NOW(),
version = version + 1
WHERE alert_id = 'alert-9912'
AND status IN ('triggered', 'notified_primary', 'escalated_secondary', 'escalated_manager', 'escalated_vp', 'snoozed')
AND version = 7;
-- Execution outcome evaluation:
-- Alice's transaction: expected state/version still matches -> rows_affected = 1. Success!
-- Bob's transaction: state or version has advanced -> rows_affected = 0.
-- Returns HTTP 409 Conflict: "Alert state or version changed by alice@company.com"
-- Escalation timers in Redis are immediately cancelled via ZREM2. Cascading Alert Storm Defense ⭐
When a core database crashes, thousands of connection alerts can overwhelm responders. A multi-tier defense shields engineers from notification fatigue:
alert_storm_mitigation_hierarchy:
stage_1_ingestion_deduplication:
mechanism: "Deterministic SHA-256 fingerprint from normalized team_id + alert_name + service + severity + priority + sorted labels"
action: "If an active alert exists with that key and priority, increment the occurrence counter and merge the payload instead of paging"
storage: "Redis key alert:dedup:{team_id}:{fingerprint} with a 5 minute TTL as an advisory cache. A Redis hit only accelerates the lookup and never suppresses an alert without authoritative PostgreSQL validation. The PostgreSQL active-alert uniqueness constraint remains authoritative"
stage_2_sliding_window_grouping:
mechanism: "Sliding 5 minute window"
action: "Aggregate related alarms across services into single consolidated incidents"
stage_3_dependency_correlation:
mechanism: "Upstream infrastructure dependency graph"
action: "If the core database cluster is confirmed down, correlate downstream service symptoms into the same incident and suppress only alerts that are proven derivative signals"
stage_4_personal_rate_limiting:
mechanism: "Token bucket per responder device"
limit: "Maximum 10 pages per hour per responder"
overflow: "Excess non-P0 alerts queue into hourly email summaries"3. Handling Schedule Gaps and Shift Transitions ⭐
If team transitions leave an unassigned calendar gap where no responder is active, automated background scans and fallback ladders ensure alarms are never dropped:
schedule_gap_detection_and_fallback:
scenario:
shift_end: "Alice shift ends Saturday 18:00 UTC"
shift_start: "Bob shift starts Sunday 09:00 UTC"
gap_duration: "15-hour unassigned coverage gap"
alert_event: "Alert fires Saturday 23:00 UTC"
preemptive_health_checks:
schedule_audit_cron: "Daily background cron scans all 50,000 team schedules 7 days into the future"
notification: "Alert team administrators immediately if unassigned slots are detected"
runtime_fallback_ladder:
tier_1: "Route alert to secondary on call engineer in the rotation layer"
tier_2: "Escalate to team designated Engineering Manager"
tier_3: "Escalate to VP of Engineering as root safety net"
rule: "No accepted alert may be silently discarded. If no responder is available, the fallback ladder routes it to team management and records the condition"Additional Considerations
1. Interactive Voice Response (IVR) Verification ⭐
Voice calls often provide stronger attention than app notifications during off hours incidents, but device and focus settings can still suppress or redirect calls. Twilio Voice workflows execute synthesized speech and DTMF tone capture:
interactive_voice_response_ivr_workflow:
objective: "Voice calls provide an additional high attention channel during night shifts, subject to device and focus settings"
call_flow_steps:
step_1_initiation:
action: "Escalation worker dispatches call request to Twilio REST API voice endpoint"
destination: "+15550199"
step_2_text_to_speech:
twiml_response: |
<Response>
<Say voice="Polly.Joanna">Critical Alert: Payment service CPU usage exceeds 90 percent.</Say>
<Gather numDigits="1" action="/api/v1/webhooks/twilio/gather" method="POST" timeout="10">
<Say>Press 1 to acknowledge this alert. Press 2 to escalate to secondary.</Say>
</Gather>
</Response>
step_3_dtmf_tone_capture:
key_1: "Triggers Twilio webhook, executes atomic CAS update to acknowledge alert, and cancels Redis timer"
key_2_or_hangup: "Triggers immediate escalation webhook, paging secondary on call responder"
step_4_failure_recovery:
retry_policy: "Busy signal or unanswered call retries within 60 seconds"
max_attempts: "3 call attempts per escalation level before escalating"2. Schedule Overrides and Shift Swaps
On-call engineers frequently need to cover slots for peers temporarily. The routing engine handles overrides dynamically with top priority over base rotations. Overlapping override windows for the same schedule are rejected during write validation so the effective responder remains deterministic:
schedule_override_resolution:
scenario: "Alice is on call but requires coverage Tuesday from 14:00 to 16:00 UTC, swapping with Bob"
data_model:
base_rotations: "on_call_schedules table defines recurring weekly or daily participant rotations"
overrides: "schedule_overrides table defines time bounded replacements with target and covering user IDs"
resolution_hierarchy:
step_1: "Check schedule_overrides for active records matching target timestamp (highest priority)"
step_2: "If no active override exists, query on_call_schedules for regular rotation assignee"
step_3: "If schedule is unassigned, trigger fallback escalation ladder to team management"3. Achieving Five Nines (99.999% SLO) ⭐
Limiting downtime to under five minutes per year requires active active multi region deployment, multi-carrier redundancy, and continuous chaos testing:
five_nines_availability_architecture:
target_slo: "99.999% availability (maximum 5.26 minutes of downtime per calendar year)"
architectural_pillars:
active_active_multi_region:
topology: "Dual-region hot deployment across us-east-1 and eu-west-1 with active active stateless ingress and workers"
database: "Each team schedule has a home write region with synchronous replication within its failure domain and tested cross-region replication for disaster recovery"
fencing: "A monotonically increasing schedule ownership epoch fences stale regional writers after failover so only the current owner can commit schedule changes"
provider_redundancy:
push_gateways: "Dual delivery pipelines across Apple APNs and Google FCM"
sms_gateways: "Primary dispatch via Twilio with automatic failover to MessageBird or Sinch on latency spike"
voice_calls: "Primary routing via Twilio Voice with automatic failover to Vonage or Nexmo"
chaos_engineering:
validation: "Automated chaos drills periodically sever regional links, fail the active schedule region, and simulate telecom blackouts to verify alert continuity and failover behavior"4. Maintenance Silences and Scheduled Suppression
Planned deployments and database migrations should not page on call engineers. A silence rule matches alert labels such as service, cluster, and alert name to suppress notifications during a scheduled window while continuing to record events for auditing. Active silence rules can be cached by team in Redis with a versioned policy key, while PostgreSQL remains authoritative. Pending escalation timers recheck the current silence state before dispatch, so creating a silence also protects against notifications already queued for execution. The alert remains stored and observable throughout the silence.
POST /api/v1/silences
Content-Type: application/json
Authorization: Bearer <auth_token>
{
"matchers": [
{
"label": "service",
"value": "payment-service"
}
],
"starts_at": "2026-03-14T22:00:00Z",
"ends_at": "2026-03-14T23:00:00Z",
"comment": "Scheduled database migration"
}{
"silence_id": "sil-881",
"status": "scheduled",
"starts_at": "2026-03-14T22:00:00Z",
"ends_at": "2026-03-14T23:00:00Z",
"version": 1,
"created_at": "2026-03-14T21:45:00Z"
}5. Webhook Authentication and Replay Protection
Alert ingestion must assume that external monitoring systems can be delayed, duplicated, or attacked. Each integration uses a signed request or mTLS where supported, includes a source event identifier and timestamp, and rejects requests outside an allowed clock skew window. The ingestion service stores the source event identifier or idempotency key so retried webhooks return the original alert result without creating another incident. Provider delivery callbacks and responder channel webhooks are authenticated and matched using provider message IDs or channel specific request identifiers before changing notification state. Per source rate limits and payload size limits protect the ingestion path during provider malfunction or abuse.
Related Problems and Concepts
Explore related architectures and distributed systems foundations that connect directly to this design:
- Notification System: Multi channel message delivery, provider failover, priority queues, and delivery tracking.
- Distributed Job Scheduler: High precision delayed job execution, time-bucketed scheduling, and distributed worker execution.
- Log Aggregation & Search System: Ingestion pipelines, log indexing, and forensic search across distributed infrastructure.
- Distributed Tracing: Trace context propagation and causal dependency mapping to diagnose incident root causes.
- Message Queues Fundamentals: Partitioned commit logs, consumer group semantics, backpressure, and dead letter queues.
- Observability and Distributed Tracing: Telemetry metrics, monitoring dashboards, and alert generation fundamentals.
- Circuit Breaker, Retries, and Bulkheads: Fault isolation and resilience patterns when communicating with third party telecom APIs.
- Redis Patterns for Interview Systems: Sorted Set timer architectures, SETNX deduplication keys, and distributed locking.
- System Design Interview Patterns: Strategic frameworks for navigating staff level infrastructure and reliability interviews.
Interview Walkthrough
Follow this structured pacing guide during a 45 minute system design interview:
- 25 minute core design pacing guide
Focus on the core scheduling and escalation loops first, expanding into multi region disaster recovery if interviewing at staff level.
- Walk the core loop: ingest alerts, deduplicate by fingerprint, match on call schedules, dispatch notifications, monitor acknowledgment timeouts, and escalate unacknowledged incidents (5 min)
- Model on call schedules as time bounded rotations with override and shift handoff semantics (6 min)
- Design an escalation ladder: push notifications, SMS texts, and interactive voice calls with per step timeouts (5 min)
- Deduplicate alerts by fingerprint (service and error signature) to prevent notification fatigue (5 min)
- Back timeouts with a durable timer engine utilizing Redis Sorted Sets or workflow schedulers (4 min)
- Clarify the core loop: ingest alerts, match the active on call schedule, notify primary responders, and start escalation timers.
- Model on call schedules as time bounded rotations with override support for temporary swaps and holiday coverage.
- Design an escalation ladder: push notifications, SMS texts, and interactive voice calls with press-to-acknowledge DTMF confirmation and configurable per step timeouts.
- Deduplicate alerts by fingerprint (combining service, cluster, and error signature) to prevent notification storms during cascading failures.
- Use a durable timer engine utilizing Redis Sorted Sets so escalation steps fire reliably even if worker nodes restart.
- Track delivery and acknowledgment state per channel because an unacknowledged push should trigger the next escalation tier automatically.
- For 99.999% availability, describe an active active multi region deployment with explicit schedule ownership and external paging provider failover across multiple push, SMS, and voice providers.
- Highlight the common pitfall where synchronous phone call initiation blocks the alert ingestion path, emphasizing that the hot ingestion path must enqueue and return HTTP 202 immediately.
Engineering Trade-offs
1. Timer Engine Implementations
This comparison evaluates timer precision, persistence durability, database load, and implementation complexity across alternative architectures.
| Approach | Precision | Durability | DB Load | Complexity |
|---|---|---|---|---|
| Redis Sorted Set ⭐ | ~100 ms target with 100 ms polling | Volatile execution index. PostgreSQL stores durable escalation state and timers can be rebuilt after Redis loss | Extremely Low | Low |
| Database Polling | Low (bounded by polling frequency) | High | High (frequent sequential queries) | Low |
| Kafka Delayed Messages | Low (requires bucketed queues) | High | None | High |
2. Notification Delivery Mechanisms: Push vs SMS vs Voice
Different communication channels exhibit distinct latency, reliability, battery, and financial trade offs:
| Pattern | Latency | Reliability | Battery and Resource Impact | Operational Overhead |
|---|---|---|---|---|
| Push Notifications (APNs / FCM) | Sub-second to 5 seconds | High for awake devices, though standard alerts can be silenced by OS Do Not Disturb unless configured with critical alert entitlements | Negligible application battery impact because the mobile OS manages push delivery | Requires managing Apple APNs certificates and Google FCM credentials with automated token refresh |
| SMS Messaging (Twilio / MessageBird) | 5 to 15 seconds | High when cellular service is available and does not require mobile data connectivity, but remains vulnerable to carrier outages and regional filtering | Zero mobile app overhead because cellular basebands receive standard SMS transmissions | Per-message carrier fees, regulatory 10DLC registration, and multi vendor fallback routing |
| Interactive Voice Calls (Twilio Voice) | 10 to 30 seconds to connect | High attention value because a voice call can alert separately from app notifications, while device call and focus settings can still suppress or redirect it | Standard incoming cellular call overhead | Highest unit cost, voice trunk management, and synthesized text-to-speech rendering |
3. Alert Grouping Latency vs Notification Promptness
This comparison balances alert aggregation delay against prompt notification during critical incidents.
| Grouping Window | Notification Latency | Alert Fatigue Suppression | Cascading Storm Resilience | Recommended Context |
|---|---|---|---|---|
| Zero Window (Immediate Dispatch) | Zero additional delay (under 5 seconds) | None because every threshold breach dispatches an isolated page | Extremely fragile because downstream engineers receive hundreds of duplicate pages during outages | P0 and emergency P1 infrastructure events where immediate human attention is paramount |
| 30 to 60 Seconds (Microbatching) | 30 to 60 seconds | Moderate by consolidating bursts of identical alarms into single incident headers | High because it absorbs initial flash floods from service restarts and rolling deployments | Default for non urgent application health alerts. Urgent P0 and P1 pages bypass the grouping delay |
| 5 Minutes (Sliding Window) | Up to 5 minutes | Maximum by aggregating flapping signals and multi cluster dependency cascades into one ticket | Near total suppression of noise across distributed service graphs | Ongoing correlation for P2 and P3 alerts after the initial notification |
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.