Interview Setup
Interview Prompt
Design a real time metrics dashboard system like Grafana or Datadog dashboards. Users create dashboards with multiple panels, query time series data sources produced by upstream telemetry infrastructure, auto-refresh, and get under one second panel updates during incidents.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| How many concurrent dashboard viewers and peak queries per second must the system support? | Handling 50K queries per second at peak across 50K active dashboards requires an aggressive Query Proxy cache hit ratio above 80% to protect underlying time series storage. |
| Should panel auto-refresh rely on client side polling or server side WebSocket push? | Client polling multiplies backend queries per viewer, whereas a server side WebSocket push distributes one backend query to all active viewers of that dashboard. |
| Are data sources confined to a single time series database or federated across multiple engines? | The Query Proxy must route queries to Prometheus, ClickHouse, or Elasticsearch depending on metric cardinality and historical retention ranges. |
| Do alert rules evaluate over the same metric query path as dashboards? | Alert evaluations are essentially scheduled time series queries, allowing the alerting engine to share the Query Proxy cache and execution pipelines. |
Scope
In scope
- WebSocket push for live dashboard auto-refresh
- Query Proxy caching and singleflight deduplication
- Time-range dynamic resolution and downsampling
- Materialized views for high traffic incident dashboards
- Capacity estimations with end-to-end mathematical derivations
Out of scope (state explicitly)
- Application client instrumentation SDKs
- The underlying raw metrics ingestion and storage pipeline, which is provided by upstream telemetry infrastructure and addressed in Distributed Metrics Aggregation System
- Full distributed tracing ingestion pipelines, which are addressed in Distributed Tracing System
- On call paging and phone escalation policies, which are addressed in On Call Escalation System
Functional Requirements
Clarify whether the real time target requires under one second streaming or sub-minute periodic refreshes. The platform combines live metric ingestion, rolling pre-aggregations, and WebSocket push updates to dashboard clients.
In an interview setting, distinguish this product-focused visualization system from the underlying Distributed Metrics Aggregation System. This design focuses on interactive querying, dashboard layout orchestration, cache coalescing, and client push delivery.
- Create and organize dashboards: Allow users to construct, edit, and organize customizable dashboards composed of multiple visualization panels.
- Support diverse widget types: Render line charts, bar charts, heatmaps, tabular grids, single-stat counters, and alert status panels.
- Federate data sources: Query federated telemetry backends including Prometheus, ClickHouse, Elasticsearch, PostgreSQL, and custom API endpoints.
- Configurable auto-refresh: Support configurable client refresh intervals and live delta streaming for continuous operational visibility.
- Dynamic templating variables: Provide dynamic dashboard variables such as service name, cluster ID, and region to filter panel queries at runtime.
- Integrated alerting rules: Define rule thresholds directly on metric expressions and dispatch notifications when states transition to firing.
- Sharing and embedding: Share parameterized dashboard snapshots via scoped URLs and embed interactive widgets into external incident management tools. Shared views still enforce dashboard permissions and never expose unrestricted datasource credentials.
- Event annotations: Overlay operational deployment markers, configuration updates, and incident alerts directly across time series charts.
Non-Functional Requirements
The system operates under a read heavy profile with sharp query bursts during active production incidents. Panel updates must render within seconds while maintaining graceful degradation when downstream data stores experience degradation.
- Low latency rendering: Initial dashboard page loads and panel query refreshes complete in under 3 seconds at the 99th percentile across complex dashboards containing 20 panels.
- High availability: Deliver 99.99% query proxy availability because engineers rely on these dashboards for mission critical incident diagnosis and operational mitigation.
- Concurrent viewer capacity: Support at least 10,000 engineers viewing dashboards concurrently during severe widespread outages without degrading backend storage systems.
- Scalability and query volume: Manage hundreds of thousands of registered dashboard configurations while handling millions of metric time series points per second across federated telemetry clusters.
- Fault isolation: Individual panel query failures or slow telemetry backends must isolate cleanly without blocking sibling panels or hanging browser client threads.
Capacity Estimations
Concurrent dashboard viewers, active panel count, and metric cardinality dictate the sizing of the query proxy cache, WebSocket connection gateway, and backend singleflight concurrency limits.
| Metric | Calculation | Value |
|---|---|---|
| Active dashboards | Given | 50K |
| Panels per dashboard (avg) | Given | 10 |
| Dashboard views / day | Given | 2M |
| Panel queries / sec (peak) | 50K active dashboards x 10 panels ÷ 10 sec refresh = 50K QPS planning ceiling | 50K |
| Panel query p99 target | Given | 500 ms |
| Dashboard definitions storage | Given | < 1 GB (JSON configs) |
Architecture Diagram
The dashboard platform consumes telemetry produced by an upstream metrics ingestion and aggregation layer. That upstream layer provides rolling 1-minute and 5-minute rollups, while this system focuses on interactive querying, caching, panel isolation, and WebSocket delivery of live deltas. Isolating the Query Proxy from the underlying time series databases ensures that when hundreds of engineers view the same high severity incident dashboard, cache hits and request coalescing prevent load multiplication against upstream Prometheus and ClickHouse clusters.
In an interview, explicitly clarify the required delivery cadence. Sub-second streaming requires stateful stream processors such as Apache Flink or Kafka Streams alongside active WebSocket distribution gateways, whereas 30-second cadence allows stateless HTTP polling against cached query proxies.
Component Deep Dives
Query Proxy: The Performance and Protection Layer
The Query Proxy mediates dashboard client requests before they reach backend telemetry stores. It validates tenant and authorization context, validates template variables against an allowlist, sanitizes query syntax, checks the distributed Redis cache, coalesces duplicate requests with in-process singleflight plus a short-lived distributed miss lock, enforces query cost limits, and formats downsampled payloads for frontend consumption. Custom API connectors are allowlisted by destination and protected by restricted outbound egress so user supplied queries cannot turn the proxy into an arbitrary network access path.
Query Proxy Execution Pipeline: 1. Authenticate Request: Validate client identity, organization context, and access permissions. 2. Template Variable Substitution: Replace dashboard template variables such as $service with 'user-service', after validating variable values against an allowlist and query safety policy. 3. Query Policy Guard: Enforce read only datasource operations, maximum time ranges, query complexity, cardinality limits, and result-size limits before cache access or execution. 4. Cache Key Derivation: Generate SHA-256 hash from (org_id, authz_scope_hash, datasource, normalized_query, aligned_time_range, step, max_data_points). 5. Cache Inspection: Check Redis for an existing validated upstream result, returning immediately on a cache hit. 6. Request Coalescing: Use in-process singleflight first, then a short-lived Redis inflight lock to coalesce identical cold-cache misses across proxy instances. 7. Execute Upstream Query: Dispatch to Prometheus, ClickHouse, Elasticsearch, PostgreSQL, or an approved custom API with dynamic step resolution. 8. Store Normalized Result: Save the compressed validated upstream series to Redis with TTL equal to the effective panel refresh interval. Panel-specific transformations remain outside the shared cache so equivalent queries with different visual transformations can still share the upstream result. 9. Post-Processing Transformations: Apply mathematical formulas, unit conversions, renaming, and LTTB downsampling to the cached or freshly fetched normalized result. 10. Deliver Response: Transmit the JSON payload to the client or broadcast an update delta only to subscribers sharing the same authorized query context. Performance Multiplier: Without caching: 100 engineers opening the same incident dashboard can trigger 100 duplicate database queries. With cache hits and cross-instance coalescing: Normally 1 upstream query serves the identical cold-cache request group, while concurrent requests share that in-flight result and later requests resolve from the shared Redis result.
Dashboard Layout and Variable Definition Schema
Dashboard definitions are stored as versioned JSON documents. The schema encapsulates visual grid positions, multi tenant permission bindings, dynamic variable dropdowns, and panel specific data source expressions. Creation and updates write the current dashboard version and its version history in the same PostgreSQL transaction so readers never observe a version number without the corresponding history record. Alert rule definitions persist their evaluation interval and notification targets alongside the trigger, recovery, and sustained duration settings.
{
"dashboard_id": "dash_7a3e8f1b",
"title": "Production Overview",
"refresh_interval_sec": 30,
"variables": [
{
"name": "service",
"type": "query",
"query": "label_values(http_requests_total, service)"
}
],
"panels": [
{
"panel_id": 1,
"title": "HTTP Request Rate",
"type": "timeseries",
"datasource": "prometheus",
"query": "rate(http_requests_total{service=\"$service\"}[5m])",
"interval": "1m",
"position": {
"x": 0,
"y": 0,
"w": 12,
"h": 8
}
}
]
}API Design
Core Domain Service Interface
The programmatic service contract defines strongly typed RPC interfaces for dashboard management, templated variable resolution, and high performance panel query execution.
// Real-time Dashboard Service Control Plane and Ingress Interfaces
export type DashboardId = string;
export type OrganizationId = string;
export type IsoTimestamp = string;
export type PanelType = "timeseries" | "stat" | "gauge" | "table" | "heatmap" | "alert_list";
export type DataSourceType = "prometheus" | "clickhouse" | "elasticsearch" | "postgresql" | "custom_api";
export interface PanelPosition {
x: number;
y: number;
w: number;
h: number;
}
export interface PanelDefinition {
panelId: number;
title: string;
type: PanelType;
datasource: DataSourceType;
query: string;
interval: string;
position: PanelPosition;
thresholds?: Array<{ value: number; color: string }>;
}
export interface QueryVariable {
name: string;
type: "query";
query: string;
currentValue?: string;
}
export interface CustomVariable {
name: string;
type: "custom";
options: string[];
currentValue?: string;
}
export interface ConstantVariable {
name: string;
type: "constant";
value: string;
currentValue?: string;
}
export type DashboardVariable = QueryVariable | CustomVariable | ConstantVariable;
export interface DashboardConfig {
dashboardId: string;
orgId: string;
title: string;
version: number;
variables: DashboardVariable[];
panels: PanelDefinition[];
refreshIntervalSec: number;
}
export interface MetricDataPoint {
timestamp: number; // Unix epoch seconds
value: number;
}
export interface TimeSeriesResult {
metricName: string;
labels: Record<string, string>;
points: MetricDataPoint[];
}
export interface QueryRequest {
datasource: DataSourceType;
query: string; // Read-only query expression validated by the datasource adapter
from: IsoTimestamp;
to: IsoTimestamp;
interval: string;
maxDataPoints?: number;
}
export interface QueryResponse {
status: "success" | "error";
series: TimeSeriesResult[];
cached: boolean;
executionTimeMs: number;
error?: { code: string; message: string };
}
export interface PanelDataDelta {
dashboardId: DashboardId;
panelId: number;
sequence: number;
generatedAt: IsoTimestamp;
mode: "append" | "replace";
series: TimeSeriesResult[];
}
export type DashboardCreateInput = Omit<DashboardConfig, "dashboardId" | "version" | "orgId">;
export type DashboardUpdateInput = Partial<Omit<DashboardConfig, "dashboardId" | "orgId" | "version">>;
export interface DashboardService {
// Create a new dashboard configuration record
createDashboard(orgId: OrganizationId, config: DashboardCreateInput): Promise<DashboardConfig>;
// Retrieve an existing dashboard configuration by unique identifier
getDashboard(dashboardId: DashboardId): Promise<DashboardConfig>;
// Update a dashboard definition using optimistic concurrency locking
updateDashboard(dashboardId: DashboardId, expectedVersion: number, config: DashboardUpdateInput): Promise<DashboardConfig>;
// Deregister a dashboard and its version history
deleteDashboard(dashboardId: DashboardId): Promise<void>;
// Execute a panel time series query through the caching and deduplication proxy
queryPanel(request: QueryRequest): Promise<QueryResponse>;
// Subscribe to live panel updates pushed over persistent WebSocket connections
subscribeLiveUpdates(dashboardId: DashboardId, onDelta: (delta: PanelDataDelta) => void): () => void;
}Dashboard Lifecycle REST Endpoints
Create, read, update, and delete dashboard configurations with optimistic concurrency control using version identifiers. Updates use the HTTP If-Match version as the concurrency token and return 409 Conflict when the supplied version is stale.
POST /api/v1/dashboards HTTP/1.1
Host: monitoring.internal.net
Authorization: Bearer <jwt>
Content-Type: application/json
{
"title": "Production Overview",
"refresh_interval_sec": 30,
"variables": [
{
"name": "service",
"type": "query",
"query": "label_values(http_requests_total, service)"
}
],
"panels": [
{
"panel_id": 1,
"title": "HTTP Request Rate",
"type": "timeseries",
"datasource": "prometheus",
"query": "rate(http_requests_total[5m])",
"interval": "1m",
"position": { "x": 0, "y": 0, "w": 12, "h": 8 }
}
]
}
HTTP/1.1 201 Created
Location: /api/v1/dashboards/dash_7a3e8f1b
---
GET /api/v1/dashboards/dash_7a3e8f1b HTTP/1.1
Host: monitoring.internal.net
Authorization: Bearer <jwt>
HTTP/1.1 200 OK
Content-Type: application/json
{
"dashboard_id": "dash_7a3e8f1b",
"org_id": "org_8f9c1b2d",
"title": "Production Overview",
"version": 1,
"variables": [
{
"name": "service",
"type": "query",
"query": "label_values(http_requests_total, service)"
}
],
"panels": [
{
"panel_id": 1,
"title": "HTTP Request Rate",
"type": "timeseries",
"datasource": "prometheus",
"query": "rate(http_requests_total[5m])",
"interval": "1m",
"position": { "x": 0, "y": 0, "w": 12, "h": 8 }
}
],
"refresh_interval_sec": 30
}
---
PUT /api/v1/dashboards/dash_7a3e8f1b HTTP/1.1
Host: monitoring.internal.net
Authorization: Bearer <jwt>
If-Match: "1"
Content-Type: application/json
{
"title": "Production Overview (Updated)",
"refresh_interval_sec": 15
}
HTTP/1.1 200 OK
Content-Type: application/json
{
"dashboard_id": "dash_7a3e8f1b",
"org_id": "org_8f9c1b2d",
"title": "Production Overview (Updated)",
"version": 2,
"variables": [
{
"name": "service",
"type": "query",
"query": "label_values(http_requests_total, service)"
}
],
"panels": [
{
"panel_id": 1,
"title": "HTTP Request Rate",
"type": "timeseries",
"datasource": "prometheus",
"query": "rate(http_requests_total[5m])",
"interval": "1m",
"position": { "x": 0, "y": 0, "w": 12, "h": 8 }
}
],
"refresh_interval_sec": 15
}
---
DELETE /api/v1/dashboards/dash_7a3e8f1b HTTP/1.1
Host: monitoring.internal.net
Authorization: Bearer <jwt>
HTTP/1.1 204 No ContentPanel Query Execution Endpoint
Execute parallel panel queries against federated data sources with validated variable interpolation, dynamic resolution, query cost controls, and downsampling. Read only PostgreSQL access and approved custom API connectors are supported alongside Prometheus, ClickHouse, and Elasticsearch.
POST /api/v1/query HTTP/1.1
Host: monitoring.internal.net
Authorization: Bearer <jwt>
Content-Type: application/json
{
"datasource": "prometheus",
"query": "rate(http_requests_total{service="user-service"}[5m])",
"from": "2026-03-14T04:00:00Z",
"to": "2026-03-14T10:00:00Z",
"interval": "1m"
}
HTTP/1.1 200 OK
Content-Type: application/json
{
"status": "success",
"cached": true,
"execution_time_ms": 12,
"series": [
{
"metric_name": "http_requests_total",
"labels": { "service": "user-service" },
"points": [
{ "timestamp": 1773460800, "value": 342.5 },
{ "timestamp": 1773460860, "value": 358.2 },
{ "timestamp": 1773460920, "value": 349.8 }
]
}
]
}Common Error Responses
Standardized error responses handle rate limits, query validation failures, and upstream telemetry timeouts.
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff 504 Gateway Timeout: search index shard responded slowly, narrow query parameters or retry
Data Model
PostgreSQL: Relational Dashboard Configuration Schema
PostgreSQL stores user accounts, dashboard definitions, and dynamic variable definitions. An integer version field enforces optimistic concurrency during simultaneous edits.
CREATE TABLE dashboards (
dashboard_id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
org_id UUID NOT NULL,
title VARCHAR(256) NOT NULL,
config JSONB NOT NULL,
version INT NOT NULL DEFAULT 1,
created_by UUID NOT NULL,
updated_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX idx_dashboards_org ON dashboards (org_id);
CREATE TABLE dashboard_versions (
dashboard_id UUID NOT NULL REFERENCES dashboards(dashboard_id) ON DELETE CASCADE,
version INT NOT NULL,
config JSONB NOT NULL,
updated_by UUID NOT NULL,
updated_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
PRIMARY KEY (dashboard_id, version)
);
CREATE TABLE alert_rules (
alert_rule_id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
org_id UUID NOT NULL,
dashboard_id UUID REFERENCES dashboards(dashboard_id) ON DELETE CASCADE,
panel_id INT,
expression TEXT NOT NULL,
trigger_threshold DOUBLE PRECISION NOT NULL,
recovery_threshold DOUBLE PRECISION NOT NULL,
for_seconds INT NOT NULL DEFAULT 300,
eval_interval_seconds INT NOT NULL DEFAULT 30,
notification_targets JSONB NOT NULL DEFAULT '[]',
enabled BOOLEAN NOT NULL DEFAULT TRUE,
version INT NOT NULL DEFAULT 1,
updated_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX idx_alert_rules_org ON alert_rules (org_id, enabled);
CREATE TABLE alert_evaluation_state (
alert_rule_id UUID PRIMARY KEY REFERENCES alert_rules(alert_rule_id) ON DELETE CASCADE,
state VARCHAR(16) NOT NULL,
state_started_at TIMESTAMPTZ NOT NULL,
last_evaluated_at TIMESTAMPTZ NOT NULL,
updated_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);Redis: Distributed Query Cache Key and Value Schema
Redis caches executed query results using tenant identity, effective authorization scope, normalized data source and query state, aligned time window boundaries, step, and output point limits. Cached time series payloads are compressed to conserve memory, and the cache is never shared across incompatible authorization scopes. WebSocket subscription membership is refreshed by heartbeats, and expired connections are pruned so a stale connection does not remain a broadcast target indefinitely.
# Redis key-value layout for time series query result caching
query_cache:
key: "query_cache:{org_id}:{authz_scope_hash}:{sha256(datasource + normalized_query + aligned_time_range + step + max_data_points)}"
value: "compressed_json_data_points"
ttl_seconds: "refresh_interval_sec"
purpose: "Tenant and authorization scoped cache. TTL follows the effective panel refresh interval"
# Cross-instance request coalescing for cold-cache misses
query_inflight_lock:
key: "query_inflight:{org_id}:{authz_scope_hash}:{cache_hash}"
type: "string"
ttl_seconds: 35
value: "{owner_token}"
purpose: "Lock lifetime slightly exceeds the 30-second hard query deadline, releasing only when the owner token matches"
# Active WebSocket dashboard viewer registration for push distribution
dashboard_subscribers:
key: "dashboard_viewers:{dashboard_id}:{view_scope_hash}"
type: "set"
members: ["conn_uuid_01", "conn_uuid_02"]
ttl_seconds: 60
purpose: "Refresh membership on heartbeat. view_scope_hash groups viewers with the same authorized variable and query context"Fault Tolerance
The dashboard platform must remain responsive during downstream database degradation, network partitions, and mass concurrent query bursts:
| Concern | Solution |
|---|---|
| Data source failure or unreachable backend | Serve stale cached data from Redis accompanied by an explicit warning banner while routing queries to backup read replicas. |
| Runaway panel query timeout | Enforce an isolated 30-second hard execution deadline per panel. The normal p99 target remains under 500 milliseconds, while the hard cap protects resources from pathological queries. |
| PostgreSQL dashboard metadata store outage | Failover read traffic to PostgreSQL read replicas while serving cached dashboard configurations from Redis memory. |
| Concurrent dashboard configuration mutations | Enforce optimistic concurrency control using an incremental version number field on the dashboard record. |
| Excessive query load or abusive user queries | Apply token-bucket rate limiting per user and organization, rejecting queries that exceed complexity and time range thresholds. |
Additional Considerations
Dashboard Stampede: Incident Causes Mass Concurrent Query Load
When a critical outage occurs, hundreds of engineers open identical dashboards within seconds. Without coordinated defenses, this sudden wave of concurrent panel executions can overwhelm upstream metric databases.
Incident Stampede Scenario: During a major operational outage, 500 engineers open the identical production dashboard simultaneously, triggering up to 5,000 panel queries at the exact same moment. Defensive Architecture: 1. Query Result Caching: The first request triggers an upstream database query, while subsequent requests within the effective cache TTL resolve directly from Redis. 2. Singleflight Request Coalescing: Concurrent cache misses for the same query key share a single in-flight backend execution, preventing duplicate queries. 3. Auto-Refresh Jitter: The client adds a plus-or-minus 10% randomized timing offset to auto-refresh timers, smoothing request spikes across time. 4. Pre-Computed Materialized Views: Critical P0 dashboards are precalculated by background workers every 30 seconds so queries read static tables instead of scanning raw time series.
Data Source Timeout: Independent Asynchronous Panel Loading
A single slow telemetry backend must never prevent responsive panels from rendering. Querying data sources independently in parallel guarantees localized failure isolation.
Independent Panel Loading Architecture:
Flawed Synchronous Approach:
The user interface waits for the slowest panel to finish querying before rendering the dashboard, delaying page visibility by up to 30 seconds.
Resilient Asynchronous Pattern:
1. Independent Panel Queries: Every panel initiates an asynchronous query request in parallel without blocking sibling widgets.
2. Progressive Rendering: Panels querying fast caches or local metric stores render immediately in under 100 milliseconds.
3. Isolated Timeout Handling: Slow or unresponsive panels display localized loading indicators and fail gracefully after a 30-second timeout while the remainder of the dashboard functions normally. The proxy cancels the upstream query when the backend supports cancellation so timed-out work does not continue consuming resources.Alert Evaluation: State Transitions and Flapping Prevention
Alert rules evaluate time series conditions across sliding evaluation windows using the same validated query execution path as dashboards. Persisted rule definitions provide the trigger threshold, recovery threshold, evaluation interval, notification targets, and sustained duration, while the evaluator maintains the OK, PENDING, FIRING, and RESOLVED states independently from dashboard rendering. Notification events use stable deduplication identities so evaluator retries remain idempotent. Incorporating pending states and hysteresis thresholds prevents rapid alert flapping caused by transient metric spikes.
Alert Evaluation State Machine:
Evaluation Rule: Trigger an alert if service error rate exceeds 5% continuously for 5 minutes.
The evaluator reads the persisted alert rule definition, including evaluation interval and notification targets, executes its normalized read only query through the shared query path, and maintains evaluation state independently from dashboard rendering. Alert evaluation can use cached results only when they satisfy the rule's freshness requirement, while the sustained PENDING interval is tracked durably enough to survive evaluator restarts without losing the rule's evaluation window. Notification delivery uses a stable alert event identity so evaluator retries do not create duplicate notifications.
State Transition Lifecycle:
1. OK: Metric remains below the configured threshold.
2. PENDING: Metric breaches the 5% threshold and a timer initializes for the 5-minute sustained duration window.
3. FIRING: Metric has remained continuously breached for 5 minutes, prompting webhook notifications to PagerDuty or Slack.
4. RESOLVED: Metric recovers below the recovery threshold, returning state to OK.
Hysteresis Configuration:
Trigger threshold: Fire when error rate exceeds 5.0%.
Recovery threshold: Resolve only when error rate drops below 3.0%.
A 2.0% hysteresis margin eliminates alert flapping on noisy boundary metrics.Multi-Region Federated Queries
Production dashboards often aggregate metrics across regional time series database clusters. The Query Proxy fans out sub-queries in parallel to each regional cluster, aligns and merges returned time series along unified timestamp boundaries, and caches the merged result in Redis. Operators accept 30 to 60 seconds of cache staleness on global rollups during incidents to protect cross-region network links. Enforcing strict panel level deadlines ensures that an outage in one region never blanks the entire operational dashboard.
Related Problems and Core Concepts
Explore complementary architectural designs spanning telemetry ingestion, time series storage, stream processing, and on call alerting workflows:
- Distributed Metrics Aggregation System: Scalable telemetry ingestion pipeline that collects raw samples, computes rolling window aggregations, and compresses long term time series.
- Distributed Tracing System: Trace collection and span graph indexing engine enabling drill-downs from anomalous dashboard latency spikes directly into distributed request paths.
- On-Call Escalation System: Incident paging and escalation platform that routes alert firing events evaluated by dashboard rules to responsible engineering rotations.
- Distributed Stream Processing: Stateful stream computing architecture handling high throughput tumbling and sliding window rollups for under one second telemetry dashboards.
- Distributed Cache: Memory caching architecture providing singleflight coalescing and low latency response delivery for high-concurrency query patterns.
- Time-Series and Metrics Storage: In-depth analysis of Gorilla delta-of-delta compression, inverted label index structures, and tiered storage lifecycles.
- Stream Processing Basics: Foundations of event-time processing, watermark semantics, and sliding window aggregations powering live dashboards.
- Network Protocols: HTTP, gRPC, WebSocket, DNS: Protocol trade-offs governing persistent WebSocket push channels vs stateless HTTP polling for real time visualization.
- Caching Patterns and Invalidation: Cache-aside, read-through, and singleflight coalescing strategies designed to prevent stampedes during incident traffic surges.
Interview Walkthrough
When presenting this real time dashboard design in a system design interview, maintain focus on query proxy caching, asynchronous panel isolation, and pre-aggregation mechanics:
- 25-minute interview pacing guide
Balance dashboard visualization concerns against backend query performance and stampede mitigation.
- Separate dashboard rendering from metric storage and clarify real time latency targets (5 min)
- Design preaggregated rollups at ingest and define the versioned dashboard schema (6 min)
- Architect the Query Proxy with Redis caching and singleflight request coalescing (5 min)
- Implement independent panel queries with localized timeout and degradation boundaries (5 min)
- Formulate the alert evaluation state machine with hysteresis to eliminate alert flapping (4 min)
- Separate dashboard rendering from underlying time series storage so the visualization platform and telemetry storage tiers scale independently.
- Pre-aggregate common queries at ingest into 1-minute and 5-minute rollups so dashboard widgets query precomputed buckets rather than scanning raw metric samples.
- Push live updates via WebSocket subscriptions keyed by dashboard identifier so connected clients receive delta refreshes instead of issuing full re-queries.
- Cache query results with time-to-live settings aligned with dashboard refresh intervals and apply singleflight coalescing so hundreds of engineers opening an incident dashboard share a single backend query.
- Load each panel independently with per panel timeouts so fast data sources render immediately while slow or degraded backends display localized loading states.
- Implement alert evaluation as a deterministic state machine moving from OK to PENDING to FIRING, incorporating hysteresis to prevent flapping on noisy telemetry thresholds.
- Add random jitter to auto-refresh intervals so thousands of active client browser tabs do not synchronize their requests on the same second.
- Highlight the common operational trap of permitting every browser refresh to trigger raw database queries, which causes catastrophic query stampedes during critical outages.
Engineering Trade-offs
Architectural decisions in real time dashboards balance client rendering interactivity against server computational load, and persistent WebSocket streams against cacheable HTTP polling.
Server-Side vs Client-Side Dashboard Rendering
Evaluating client side canvas and SVG graph rendering against server-rendered static image panels shapes browser memory consumption and interactive drill-down capabilities.
| Approach | Interactivity | Server Resource Load | Large Dataset Handling |
|---|---|---|---|
| Client-side rendering (e.g., Grafana, WebGL canvas) | High interactivity with client side zoom, pan, hover tooltips, and time range adjustments. | Low server load because servers only transmit downsampled JSON vector points. | High browser memory usage if series downsampling is omitted, risking client frame drops. |
| Server-side rendering (e.g., headless browser or PNG generation) | Static image representation with zero client interactivity, where tooltips and zoom require fresh requests. | Substantial server CPU and memory overhead required to execute image rendering libraries. | Predictable browser performance because clients load static images regardless of data point volume. |
WebSocket Push Streaming vs Periodic HTTP Polling
Choosing between persistent bidirectional WebSocket connections and decoupled periodic HTTP polling impacts gateway connection scale, cache hit rates, and incident survivability.
| Dimension | WebSocket Push Streaming | Periodic HTTP Polling |
|---|---|---|
| Connection Lifecycle | Long-lived persistent TCP and TLS state maintained per active browser tab. | Short-lived stateless HTTP requests opened and closed per refresh cycle. |
| Update Latency | Sub-second delivery of new metric data points as soon as aggregation completes. | Update latency bounded by the client polling interval, typically 10 to 60 seconds. |
| Edge Cacheability | Bypasses standard HTTP caching layers and requires application-level distribution brokers. | Highly cacheable via CDN edges, API gateways, and Redis reverse proxy caches. |
| Incident Survivability | Mass disconnects cause reconnection storms that can exhaust gateway file descriptors. | Standard HTTP retry backoff and client jitter naturally distribute network demand. |
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.