System Design Problem

Design a Real-Time Dashboard and Metrics System

Commonly Asked By:DatadogUberGrafanaAWS

Interview Setup

Interview Prompt

Design a real time metrics dashboard system like Grafana or Datadog dashboards. Users create dashboards with multiple panels, query time series data sources produced by upstream telemetry infrastructure, auto-refresh, and get under one second panel updates during incidents.

Clarifying Questions (ask before designing)

QuestionWhy it matters
How many concurrent dashboard viewers and peak queries per second must the system support?Handling 50K queries per second at peak across 50K active dashboards requires an aggressive Query Proxy cache hit ratio above 80% to protect underlying time series storage.
Should panel auto-refresh rely on client side polling or server side WebSocket push?Client polling multiplies backend queries per viewer, whereas a server side WebSocket push distributes one backend query to all active viewers of that dashboard.
Are data sources confined to a single time series database or federated across multiple engines?The Query Proxy must route queries to Prometheus, ClickHouse, or Elasticsearch depending on metric cardinality and historical retention ranges.
Do alert rules evaluate over the same metric query path as dashboards?Alert evaluations are essentially scheduled time series queries, allowing the alerting engine to share the Query Proxy cache and execution pipelines.

Scope

In scope

  • WebSocket push for live dashboard auto-refresh
  • Query Proxy caching and singleflight deduplication
  • Time-range dynamic resolution and downsampling
  • Materialized views for high traffic incident dashboards
  • Capacity estimations with end-to-end mathematical derivations

Out of scope (state explicitly)

Functional Requirements

Clarify whether the real time target requires under one second streaming or sub-minute periodic refreshes. The platform combines live metric ingestion, rolling pre-aggregations, and WebSocket push updates to dashboard clients.

In an interview setting, distinguish this product-focused visualization system from the underlying Distributed Metrics Aggregation System. This design focuses on interactive querying, dashboard layout orchestration, cache coalescing, and client push delivery.

  • Create and organize dashboards: Allow users to construct, edit, and organize customizable dashboards composed of multiple visualization panels.
  • Support diverse widget types: Render line charts, bar charts, heatmaps, tabular grids, single-stat counters, and alert status panels.
  • Federate data sources: Query federated telemetry backends including Prometheus, ClickHouse, Elasticsearch, PostgreSQL, and custom API endpoints.
  • Configurable auto-refresh: Support configurable client refresh intervals and live delta streaming for continuous operational visibility.
  • Dynamic templating variables: Provide dynamic dashboard variables such as service name, cluster ID, and region to filter panel queries at runtime.
  • Integrated alerting rules: Define rule thresholds directly on metric expressions and dispatch notifications when states transition to firing.
  • Sharing and embedding: Share parameterized dashboard snapshots via scoped URLs and embed interactive widgets into external incident management tools. Shared views still enforce dashboard permissions and never expose unrestricted datasource credentials.
  • Event annotations: Overlay operational deployment markers, configuration updates, and incident alerts directly across time series charts.

Non-Functional Requirements

The system operates under a read heavy profile with sharp query bursts during active production incidents. Panel updates must render within seconds while maintaining graceful degradation when downstream data stores experience degradation.

  • Low latency rendering: Initial dashboard page loads and panel query refreshes complete in under 3 seconds at the 99th percentile across complex dashboards containing 20 panels.
  • High availability: Deliver 99.99% query proxy availability because engineers rely on these dashboards for mission critical incident diagnosis and operational mitigation.
  • Concurrent viewer capacity: Support at least 10,000 engineers viewing dashboards concurrently during severe widespread outages without degrading backend storage systems.
  • Scalability and query volume: Manage hundreds of thousands of registered dashboard configurations while handling millions of metric time series points per second across federated telemetry clusters.
  • Fault isolation: Individual panel query failures or slow telemetry backends must isolate cleanly without blocking sibling panels or hanging browser client threads.

Capacity Estimations

Concurrent dashboard viewers, active panel count, and metric cardinality dictate the sizing of the query proxy cache, WebSocket connection gateway, and backend singleflight concurrency limits.

MetricCalculationValue
Active dashboardsGiven50K
Panels per dashboard (avg)Given10
Dashboard views / dayGiven2M
Panel queries / sec (peak)50K active dashboards x 10 panels ÷ 10 sec refresh = 50K QPS planning ceiling50K
Panel query p99 targetGiven500 ms
Dashboard definitions storageGiven< 1 GB (JSON configs)

Architecture Diagram

The dashboard platform consumes telemetry produced by an upstream metrics ingestion and aggregation layer. That upstream layer provides rolling 1-minute and 5-minute rollups, while this system focuses on interactive querying, caching, panel isolation, and WebSocket delivery of live deltas. Isolating the Query Proxy from the underlying time series databases ensures that when hundreds of engineers view the same high severity incident dashboard, cache hits and request coalescing prevent load multiplication against upstream Prometheus and ClickHouse clusters.

In an interview, explicitly clarify the required delivery cadence. Sub-second streaming requires stateful stream processors such as Apache Flink or Kafka Streams alongside active WebSocket distribution gateways, whereas 30-second cadence allows stateless HTTP polling against cached query proxies.

Loading...

Component Deep Dives

Query Proxy: The Performance and Protection Layer

The Query Proxy mediates dashboard client requests before they reach backend telemetry stores. It validates tenant and authorization context, validates template variables against an allowlist, sanitizes query syntax, checks the distributed Redis cache, coalesces duplicate requests with in-process singleflight plus a short-lived distributed miss lock, enforces query cost limits, and formats downsampled payloads for frontend consumption. Custom API connectors are allowlisted by destination and protected by restricted outbound egress so user supplied queries cannot turn the proxy into an arbitrary network access path.

Query Proxy Execution Pipeline:
  1. Authenticate Request: Validate client identity, organization context, and access permissions.
  2. Template Variable Substitution: Replace dashboard template variables such as $service with 'user-service', after validating variable values against an allowlist and query safety policy.
  3. Query Policy Guard: Enforce read only datasource operations, maximum time ranges, query complexity, cardinality limits, and result-size limits before cache access or execution.
  4. Cache Key Derivation: Generate SHA-256 hash from (org_id, authz_scope_hash, datasource, normalized_query, aligned_time_range, step, max_data_points).
  5. Cache Inspection: Check Redis for an existing validated upstream result, returning immediately on a cache hit.
  6. Request Coalescing: Use in-process singleflight first, then a short-lived Redis inflight lock to coalesce identical cold-cache misses across proxy instances.
  7. Execute Upstream Query: Dispatch to Prometheus, ClickHouse, Elasticsearch, PostgreSQL, or an approved custom API with dynamic step resolution.
  8. Store Normalized Result: Save the compressed validated upstream series to Redis with TTL equal to the effective panel refresh interval. Panel-specific transformations remain outside the shared cache so equivalent queries with different visual transformations can still share the upstream result.
  9. Post-Processing Transformations: Apply mathematical formulas, unit conversions, renaming, and LTTB downsampling to the cached or freshly fetched normalized result.
 10. Deliver Response: Transmit the JSON payload to the client or broadcast an update delta only to subscribers sharing the same authorized query context.

Performance Multiplier:
  Without caching: 100 engineers opening the same incident dashboard can trigger 100 duplicate database queries.
  With cache hits and cross-instance coalescing: Normally 1 upstream query serves the identical cold-cache request group, while concurrent requests share that in-flight result and later requests resolve from the shared Redis result.

Dashboard Layout and Variable Definition Schema

Dashboard definitions are stored as versioned JSON documents. The schema encapsulates visual grid positions, multi tenant permission bindings, dynamic variable dropdowns, and panel specific data source expressions. Creation and updates write the current dashboard version and its version history in the same PostgreSQL transaction so readers never observe a version number without the corresponding history record. Alert rule definitions persist their evaluation interval and notification targets alongside the trigger, recovery, and sustained duration settings.

JSON
{
  "dashboard_id": "dash_7a3e8f1b",
  "title": "Production Overview",
  "refresh_interval_sec": 30,
  "variables": [
    {
      "name": "service",
      "type": "query",
      "query": "label_values(http_requests_total, service)"
    }
  ],
  "panels": [
    {
      "panel_id": 1,
      "title": "HTTP Request Rate",
      "type": "timeseries",
      "datasource": "prometheus",
      "query": "rate(http_requests_total{service=\"$service\"}[5m])",
      "interval": "1m",
      "position": {
        "x": 0,
        "y": 0,
        "w": 12,
        "h": 8
      }
    }
  ]
}

API Design

Core Domain Service Interface

The programmatic service contract defines strongly typed RPC interfaces for dashboard management, templated variable resolution, and high performance panel query execution.

TYPESCRIPT
// Real-time Dashboard Service Control Plane and Ingress Interfaces

export type DashboardId = string;
export type OrganizationId = string;
export type IsoTimestamp = string;

export type PanelType = "timeseries" | "stat" | "gauge" | "table" | "heatmap" | "alert_list";
export type DataSourceType = "prometheus" | "clickhouse" | "elasticsearch" | "postgresql" | "custom_api";

export interface PanelPosition {
  x: number;
  y: number;
  w: number;
  h: number;
}

export interface PanelDefinition {
  panelId: number;
  title: string;
  type: PanelType;
  datasource: DataSourceType;
  query: string;
  interval: string;
  position: PanelPosition;
  thresholds?: Array<{ value: number; color: string }>;
}

export interface QueryVariable {
  name: string;
  type: "query";
  query: string;
  currentValue?: string;
}

export interface CustomVariable {
  name: string;
  type: "custom";
  options: string[];
  currentValue?: string;
}

export interface ConstantVariable {
  name: string;
  type: "constant";
  value: string;
  currentValue?: string;
}

export type DashboardVariable = QueryVariable | CustomVariable | ConstantVariable;

export interface DashboardConfig {
  dashboardId: string;
  orgId: string;
  title: string;
  version: number;
  variables: DashboardVariable[];
  panels: PanelDefinition[];
  refreshIntervalSec: number;
}

export interface MetricDataPoint {
  timestamp: number; // Unix epoch seconds
  value: number;
}

export interface TimeSeriesResult {
  metricName: string;
  labels: Record<string, string>;
  points: MetricDataPoint[];
}

export interface QueryRequest {
  datasource: DataSourceType;
  query: string; // Read-only query expression validated by the datasource adapter
  from: IsoTimestamp;
  to: IsoTimestamp;
  interval: string;
  maxDataPoints?: number;
}

export interface QueryResponse {
  status: "success" | "error";
  series: TimeSeriesResult[];
  cached: boolean;
  executionTimeMs: number;
  error?: { code: string; message: string };
}

export interface PanelDataDelta {
  dashboardId: DashboardId;
  panelId: number;
  sequence: number;
  generatedAt: IsoTimestamp;
  mode: "append" | "replace";
  series: TimeSeriesResult[];
}

export type DashboardCreateInput = Omit<DashboardConfig, "dashboardId" | "version" | "orgId">;
export type DashboardUpdateInput = Partial<Omit<DashboardConfig, "dashboardId" | "orgId" | "version">>;

export interface DashboardService {
  // Create a new dashboard configuration record
  createDashboard(orgId: OrganizationId, config: DashboardCreateInput): Promise<DashboardConfig>;

  // Retrieve an existing dashboard configuration by unique identifier
  getDashboard(dashboardId: DashboardId): Promise<DashboardConfig>;

  // Update a dashboard definition using optimistic concurrency locking
  updateDashboard(dashboardId: DashboardId, expectedVersion: number, config: DashboardUpdateInput): Promise<DashboardConfig>;

  // Deregister a dashboard and its version history
  deleteDashboard(dashboardId: DashboardId): Promise<void>;

  // Execute a panel time series query through the caching and deduplication proxy
  queryPanel(request: QueryRequest): Promise<QueryResponse>;

  // Subscribe to live panel updates pushed over persistent WebSocket connections
  subscribeLiveUpdates(dashboardId: DashboardId, onDelta: (delta: PanelDataDelta) => void): () => void;
}

Dashboard Lifecycle REST Endpoints

Create, read, update, and delete dashboard configurations with optimistic concurrency control using version identifiers. Updates use the HTTP If-Match version as the concurrency token and return 409 Conflict when the supplied version is stale.

HTTP
POST /api/v1/dashboards HTTP/1.1
Host: monitoring.internal.net
Authorization: Bearer <jwt>
Content-Type: application/json

{
  "title": "Production Overview",
  "refresh_interval_sec": 30,
  "variables": [
    {
      "name": "service",
      "type": "query",
      "query": "label_values(http_requests_total, service)"
    }
  ],
  "panels": [
    {
      "panel_id": 1,
      "title": "HTTP Request Rate",
      "type": "timeseries",
      "datasource": "prometheus",
      "query": "rate(http_requests_total[5m])",
      "interval": "1m",
      "position": { "x": 0, "y": 0, "w": 12, "h": 8 }
    }
  ]
}

HTTP/1.1 201 Created
Location: /api/v1/dashboards/dash_7a3e8f1b

---

GET /api/v1/dashboards/dash_7a3e8f1b HTTP/1.1
Host: monitoring.internal.net
Authorization: Bearer <jwt>

HTTP/1.1 200 OK
Content-Type: application/json

{
  "dashboard_id": "dash_7a3e8f1b",
  "org_id": "org_8f9c1b2d",
  "title": "Production Overview",
  "version": 1,
  "variables": [
    {
      "name": "service",
      "type": "query",
      "query": "label_values(http_requests_total, service)"
    }
  ],
  "panels": [
    {
      "panel_id": 1,
      "title": "HTTP Request Rate",
      "type": "timeseries",
      "datasource": "prometheus",
      "query": "rate(http_requests_total[5m])",
      "interval": "1m",
      "position": { "x": 0, "y": 0, "w": 12, "h": 8 }
    }
  ],
  "refresh_interval_sec": 30
}

---

PUT /api/v1/dashboards/dash_7a3e8f1b HTTP/1.1
Host: monitoring.internal.net
Authorization: Bearer <jwt>
If-Match: "1"
Content-Type: application/json

{
  "title": "Production Overview (Updated)",
  "refresh_interval_sec": 15
}

HTTP/1.1 200 OK
Content-Type: application/json

{
  "dashboard_id": "dash_7a3e8f1b",
  "org_id": "org_8f9c1b2d",
  "title": "Production Overview (Updated)",
  "version": 2,
  "variables": [
    {
      "name": "service",
      "type": "query",
      "query": "label_values(http_requests_total, service)"
    }
  ],
  "panels": [
    {
      "panel_id": 1,
      "title": "HTTP Request Rate",
      "type": "timeseries",
      "datasource": "prometheus",
      "query": "rate(http_requests_total[5m])",
      "interval": "1m",
      "position": { "x": 0, "y": 0, "w": 12, "h": 8 }
    }
  ],
  "refresh_interval_sec": 15
}

---

DELETE /api/v1/dashboards/dash_7a3e8f1b HTTP/1.1
Host: monitoring.internal.net
Authorization: Bearer <jwt>

HTTP/1.1 204 No Content

Panel Query Execution Endpoint

Execute parallel panel queries against federated data sources with validated variable interpolation, dynamic resolution, query cost controls, and downsampling. Read only PostgreSQL access and approved custom API connectors are supported alongside Prometheus, ClickHouse, and Elasticsearch.

HTTP
POST /api/v1/query HTTP/1.1
Host: monitoring.internal.net
Authorization: Bearer <jwt>
Content-Type: application/json

{
  "datasource": "prometheus",
  "query": "rate(http_requests_total{service="user-service"}[5m])",
  "from": "2026-03-14T04:00:00Z",
  "to": "2026-03-14T10:00:00Z",
  "interval": "1m"
}

HTTP/1.1 200 OK
Content-Type: application/json

{
  "status": "success",
  "cached": true,
  "execution_time_ms": 12,
  "series": [
    {
      "metric_name": "http_requests_total",
      "labels": { "service": "user-service" },
      "points": [
        { "timestamp": 1773460800, "value": 342.5 },
        { "timestamp": 1773460860, "value": 358.2 },
        { "timestamp": 1773460920, "value": 349.8 }
      ]
    }
  ]
}

Common Error Responses

Standardized error responses handle rate limits, query validation failures, and upstream telemetry timeouts.

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
504 Gateway Timeout: search index shard responded slowly, narrow query parameters or retry

Data Model

PostgreSQL: Relational Dashboard Configuration Schema

PostgreSQL stores user accounts, dashboard definitions, and dynamic variable definitions. An integer version field enforces optimistic concurrency during simultaneous edits.

SQL
CREATE TABLE dashboards (
    dashboard_id    UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    org_id          UUID NOT NULL,
    title           VARCHAR(256) NOT NULL,
    config          JSONB NOT NULL,
    version         INT NOT NULL DEFAULT 1,
    created_by      UUID NOT NULL,
    updated_at      TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP
);

CREATE INDEX idx_dashboards_org ON dashboards (org_id);

CREATE TABLE dashboard_versions (
    dashboard_id    UUID NOT NULL REFERENCES dashboards(dashboard_id) ON DELETE CASCADE,
    version         INT NOT NULL,
    config          JSONB NOT NULL,
    updated_by      UUID NOT NULL,
    updated_at      TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
    PRIMARY KEY (dashboard_id, version)
);

CREATE TABLE alert_rules (
    alert_rule_id          UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    org_id                 UUID NOT NULL,
    dashboard_id           UUID REFERENCES dashboards(dashboard_id) ON DELETE CASCADE,
    panel_id               INT,
    expression             TEXT NOT NULL,
    trigger_threshold      DOUBLE PRECISION NOT NULL,
    recovery_threshold     DOUBLE PRECISION NOT NULL,
    for_seconds            INT NOT NULL DEFAULT 300,
    eval_interval_seconds  INT NOT NULL DEFAULT 30,
    notification_targets   JSONB NOT NULL DEFAULT '[]',
    enabled                BOOLEAN NOT NULL DEFAULT TRUE,
    version                INT NOT NULL DEFAULT 1,
    updated_at             TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);

CREATE INDEX idx_alert_rules_org ON alert_rules (org_id, enabled);

CREATE TABLE alert_evaluation_state (
    alert_rule_id      UUID PRIMARY KEY REFERENCES alert_rules(alert_rule_id) ON DELETE CASCADE,
    state              VARCHAR(16) NOT NULL,
    state_started_at   TIMESTAMPTZ NOT NULL,
    last_evaluated_at  TIMESTAMPTZ NOT NULL,
    updated_at         TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);

Redis: Distributed Query Cache Key and Value Schema

Redis caches executed query results using tenant identity, effective authorization scope, normalized data source and query state, aligned time window boundaries, step, and output point limits. Cached time series payloads are compressed to conserve memory, and the cache is never shared across incompatible authorization scopes. WebSocket subscription membership is refreshed by heartbeats, and expired connections are pruned so a stale connection does not remain a broadcast target indefinitely.

YAML
# Redis key-value layout for time series query result caching
query_cache:
  key: "query_cache:{org_id}:{authz_scope_hash}:{sha256(datasource + normalized_query + aligned_time_range + step + max_data_points)}"
  value: "compressed_json_data_points"
  ttl_seconds: "refresh_interval_sec"
  purpose: "Tenant and authorization scoped cache. TTL follows the effective panel refresh interval"

# Cross-instance request coalescing for cold-cache misses
query_inflight_lock:
  key: "query_inflight:{org_id}:{authz_scope_hash}:{cache_hash}"
  type: "string"
  ttl_seconds: 35
  value: "{owner_token}"
  purpose: "Lock lifetime slightly exceeds the 30-second hard query deadline, releasing only when the owner token matches"

# Active WebSocket dashboard viewer registration for push distribution
dashboard_subscribers:
  key: "dashboard_viewers:{dashboard_id}:{view_scope_hash}"
  type: "set"
  members: ["conn_uuid_01", "conn_uuid_02"]
  ttl_seconds: 60
  purpose: "Refresh membership on heartbeat. view_scope_hash groups viewers with the same authorized variable and query context"

Fault Tolerance

The dashboard platform must remain responsive during downstream database degradation, network partitions, and mass concurrent query bursts:

ConcernSolution
Data source failure or unreachable backendServe stale cached data from Redis accompanied by an explicit warning banner while routing queries to backup read replicas.
Runaway panel query timeoutEnforce an isolated 30-second hard execution deadline per panel. The normal p99 target remains under 500 milliseconds, while the hard cap protects resources from pathological queries.
PostgreSQL dashboard metadata store outageFailover read traffic to PostgreSQL read replicas while serving cached dashboard configurations from Redis memory.
Concurrent dashboard configuration mutationsEnforce optimistic concurrency control using an incremental version number field on the dashboard record.
Excessive query load or abusive user queriesApply token-bucket rate limiting per user and organization, rejecting queries that exceed complexity and time range thresholds.

Additional Considerations

Dashboard Stampede: Incident Causes Mass Concurrent Query Load

When a critical outage occurs, hundreds of engineers open identical dashboards within seconds. Without coordinated defenses, this sudden wave of concurrent panel executions can overwhelm upstream metric databases.

Incident Stampede Scenario:
  During a major operational outage, 500 engineers open the identical production dashboard simultaneously, triggering up to 5,000 panel queries at the exact same moment.

Defensive Architecture:
  1. Query Result Caching: The first request triggers an upstream database query, while subsequent requests within the effective cache TTL resolve directly from Redis.
  2. Singleflight Request Coalescing: Concurrent cache misses for the same query key share a single in-flight backend execution, preventing duplicate queries.
  3. Auto-Refresh Jitter: The client adds a plus-or-minus 10% randomized timing offset to auto-refresh timers, smoothing request spikes across time.
  4. Pre-Computed Materialized Views: Critical P0 dashboards are precalculated by background workers every 30 seconds so queries read static tables instead of scanning raw time series.

Data Source Timeout: Independent Asynchronous Panel Loading

A single slow telemetry backend must never prevent responsive panels from rendering. Querying data sources independently in parallel guarantees localized failure isolation.

Independent Panel Loading Architecture:
  Flawed Synchronous Approach:
    The user interface waits for the slowest panel to finish querying before rendering the dashboard, delaying page visibility by up to 30 seconds.

  Resilient Asynchronous Pattern:
    1. Independent Panel Queries: Every panel initiates an asynchronous query request in parallel without blocking sibling widgets.
    2. Progressive Rendering: Panels querying fast caches or local metric stores render immediately in under 100 milliseconds.
    3. Isolated Timeout Handling: Slow or unresponsive panels display localized loading indicators and fail gracefully after a 30-second timeout while the remainder of the dashboard functions normally. The proxy cancels the upstream query when the backend supports cancellation so timed-out work does not continue consuming resources.

Alert Evaluation: State Transitions and Flapping Prevention

Alert rules evaluate time series conditions across sliding evaluation windows using the same validated query execution path as dashboards. Persisted rule definitions provide the trigger threshold, recovery threshold, evaluation interval, notification targets, and sustained duration, while the evaluator maintains the OK, PENDING, FIRING, and RESOLVED states independently from dashboard rendering. Notification events use stable deduplication identities so evaluator retries remain idempotent. Incorporating pending states and hysteresis thresholds prevents rapid alert flapping caused by transient metric spikes.

Alert Evaluation State Machine:
  Evaluation Rule: Trigger an alert if service error rate exceeds 5% continuously for 5 minutes.
  The evaluator reads the persisted alert rule definition, including evaluation interval and notification targets, executes its normalized read only query through the shared query path, and maintains evaluation state independently from dashboard rendering. Alert evaluation can use cached results only when they satisfy the rule's freshness requirement, while the sustained PENDING interval is tracked durably enough to survive evaluator restarts without losing the rule's evaluation window. Notification delivery uses a stable alert event identity so evaluator retries do not create duplicate notifications.

  State Transition Lifecycle:
    1. OK: Metric remains below the configured threshold.
    2. PENDING: Metric breaches the 5% threshold and a timer initializes for the 5-minute sustained duration window.
    3. FIRING: Metric has remained continuously breached for 5 minutes, prompting webhook notifications to PagerDuty or Slack.
    4. RESOLVED: Metric recovers below the recovery threshold, returning state to OK.

  Hysteresis Configuration:
    Trigger threshold: Fire when error rate exceeds 5.0%.
    Recovery threshold: Resolve only when error rate drops below 3.0%.
    A 2.0% hysteresis margin eliminates alert flapping on noisy boundary metrics.

Multi-Region Federated Queries

Production dashboards often aggregate metrics across regional time series database clusters. The Query Proxy fans out sub-queries in parallel to each regional cluster, aligns and merges returned time series along unified timestamp boundaries, and caches the merged result in Redis. Operators accept 30 to 60 seconds of cache staleness on global rollups during incidents to protect cross-region network links. Enforcing strict panel level deadlines ensures that an outage in one region never blanks the entire operational dashboard.

Related Problems and Core Concepts

Explore complementary architectural designs spanning telemetry ingestion, time series storage, stream processing, and on call alerting workflows:

  • Distributed Metrics Aggregation System: Scalable telemetry ingestion pipeline that collects raw samples, computes rolling window aggregations, and compresses long term time series.
  • Distributed Tracing System: Trace collection and span graph indexing engine enabling drill-downs from anomalous dashboard latency spikes directly into distributed request paths.
  • On-Call Escalation System: Incident paging and escalation platform that routes alert firing events evaluated by dashboard rules to responsible engineering rotations.
  • Distributed Stream Processing: Stateful stream computing architecture handling high throughput tumbling and sliding window rollups for under one second telemetry dashboards.
  • Distributed Cache: Memory caching architecture providing singleflight coalescing and low latency response delivery for high-concurrency query patterns.
  • Time-Series and Metrics Storage: In-depth analysis of Gorilla delta-of-delta compression, inverted label index structures, and tiered storage lifecycles.
  • Stream Processing Basics: Foundations of event-time processing, watermark semantics, and sliding window aggregations powering live dashboards.
  • Network Protocols: HTTP, gRPC, WebSocket, DNS: Protocol trade-offs governing persistent WebSocket push channels vs stateless HTTP polling for real time visualization.
  • Caching Patterns and Invalidation: Cache-aside, read-through, and singleflight coalescing strategies designed to prevent stampedes during incident traffic surges.

Interview Walkthrough

When presenting this real time dashboard design in a system design interview, maintain focus on query proxy caching, asynchronous panel isolation, and pre-aggregation mechanics:

  • 25-minute interview pacing guide

    Balance dashboard visualization concerns against backend query performance and stampede mitigation.

    • Separate dashboard rendering from metric storage and clarify real time latency targets (5 min)
    • Design preaggregated rollups at ingest and define the versioned dashboard schema (6 min)
    • Architect the Query Proxy with Redis caching and singleflight request coalescing (5 min)
    • Implement independent panel queries with localized timeout and degradation boundaries (5 min)
    • Formulate the alert evaluation state machine with hysteresis to eliminate alert flapping (4 min)
  • Separate dashboard rendering from underlying time series storage so the visualization platform and telemetry storage tiers scale independently.
  • Pre-aggregate common queries at ingest into 1-minute and 5-minute rollups so dashboard widgets query precomputed buckets rather than scanning raw metric samples.
  • Push live updates via WebSocket subscriptions keyed by dashboard identifier so connected clients receive delta refreshes instead of issuing full re-queries.
  • Cache query results with time-to-live settings aligned with dashboard refresh intervals and apply singleflight coalescing so hundreds of engineers opening an incident dashboard share a single backend query.
  • Load each panel independently with per panel timeouts so fast data sources render immediately while slow or degraded backends display localized loading states.
  • Implement alert evaluation as a deterministic state machine moving from OK to PENDING to FIRING, incorporating hysteresis to prevent flapping on noisy telemetry thresholds.
  • Add random jitter to auto-refresh intervals so thousands of active client browser tabs do not synchronize their requests on the same second.
  • Highlight the common operational trap of permitting every browser refresh to trigger raw database queries, which causes catastrophic query stampedes during critical outages.

Engineering Trade-offs

Architectural decisions in real time dashboards balance client rendering interactivity against server computational load, and persistent WebSocket streams against cacheable HTTP polling.

Server-Side vs Client-Side Dashboard Rendering

Evaluating client side canvas and SVG graph rendering against server-rendered static image panels shapes browser memory consumption and interactive drill-down capabilities.

ApproachInteractivityServer Resource LoadLarge Dataset Handling
Client-side rendering (e.g., Grafana, WebGL canvas)High interactivity with client side zoom, pan, hover tooltips, and time range adjustments.Low server load because servers only transmit downsampled JSON vector points.High browser memory usage if series downsampling is omitted, risking client frame drops.
Server-side rendering (e.g., headless browser or PNG generation)Static image representation with zero client interactivity, where tooltips and zoom require fresh requests.Substantial server CPU and memory overhead required to execute image rendering libraries.Predictable browser performance because clients load static images regardless of data point volume.

WebSocket Push Streaming vs Periodic HTTP Polling

Choosing between persistent bidirectional WebSocket connections and decoupled periodic HTTP polling impacts gateway connection scale, cache hit rates, and incident survivability.

DimensionWebSocket Push StreamingPeriodic HTTP Polling
Connection LifecycleLong-lived persistent TCP and TLS state maintained per active browser tab.Short-lived stateless HTTP requests opened and closed per refresh cycle.
Update LatencySub-second delivery of new metric data points as soon as aggregation completes.Update latency bounded by the client polling interval, typically 10 to 60 seconds.
Edge CacheabilityBypasses standard HTTP caching layers and requires application-level distribution brokers.Highly cacheable via CDN edges, API gateways, and Redis reverse proxy caches.
Incident SurvivabilityMass disconnects cause reconnection storms that can exhaust gateway file descriptors.Standard HTTP retry backoff and client jitter naturally distribute network demand.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...