Interview Setup
Interview Prompt
Design ChatGPT for 100M DAU, ~8K messages/sec peak, ~20M tokens/sec inference throughput, TTFT p99 < 1.5s, and 500+ tokens/sec effective GPU generation throughput at the inference tier as a planning assumption. Assume self-hosted GPU inference is in scope and chat history operates at billion-message scale.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Are we hosting our own LLM models or using a third party API? | Hosting own models requires deep GPU cluster management and inference routing. Using APIs shifts the focus entirely to state management and UI streaming. |
| Do we need to retain chat history indefinitely? | Impacts database choice (e.g., Cassandra for massive write heavy logs). |
| What is the expected latency for the first token? | Time To First Token (TTFT) is the most critical UX metric for LLM apps. |
Scope
In scope
- Real time streaming of tokens to the client
- Chat session and history management
- Context window truncation and summarization
- Orchestration of inference requests (if hosting models)
Out of scope (state explicitly)
- Training or fine tuning the foundational models
- Detailed implementation of the transformer architecture
Functional Requirements
Start by asking your interviewer whether you host your own models or call a third party API, because that single answer shifts the design from GPU fleet management to session storage and streaming UX. Confirm context window size, chat history retention, and whether multi turn memory is in scope.
- Conversational UI: Users can send text prompts and receive text responses from an AI.
- Streaming Responses: Responses must be streamed back to the user token by token in real time.
- Context/History: The AI must remember the context of the current conversation session.
- Chat History: Users can view, resume, and delete past conversation threads.
- Identity & Access Control: Each request is authenticated and authorized against the conversation owner, with strict isolation between users and conversations.
- Generation Control: Users can stop an in flight generation, and the service cancels unnecessary inference work when safe while preserving partial response state.
Non-Functional Requirements
Your interviewer will stress-test TTFT (time to first token) vs TPOT (time per output token) as separate tuning knobs, because users feel latency until the first token arrives and then care about ongoing generation speed. They will also probe context window overflow and how history is truncated or summarized without losing coherence.
- Low TTFT (Time To First Token): Treat 500ms as the user experience stretch target for the first visible token, while the engineering SLO targets TTFT p99 below 1.5s.
- High Throughput (Tokens per Second): Generation speed should match or exceed human reading speed (~20-30 tokens/sec).
- Scalability: GPU resources are extremely expensive and limited. The system must maximize GPU utilization via batching.
- Context Window Limits: The system must gracefully handle truncating or summarizing history when the context window is exceeded.
- Safety and Abuse Controls: Required input and output policy checks, token based rate limits, and abuse controls must protect the service without making normal streaming unnecessarily slow.
Capacity Estimations
Size the GPU fleet from separate prefill and decode workloads rather than raw messages per second. Each request carries approximately 2.5K total tokens on average, producing ~20M total tokens/sec at peak. The stated ~500 tokens/sec/GPU figure is an effective generation rate, so it sizes the ~4M output tokens/sec decode workload at roughly 8K GPUs before headroom. Prefill capacity requires its own model and hardware benchmark.
| Metric | Calculation | Value |
|---|---|---|
| Daily Active Users (DAU) | Given | 100M |
| Peak concurrent chat sessions | DAU x 5% | 5M |
| Messages / sec (peak) | 5M x 0.1 msg/min | ~8K msg/s |
| Avg prompt tokens | Given (history + system) | ~2K tokens |
| Avg completion tokens | Given | ~500 tokens |
| Peak input tokens/sec | 8K x 2K tokens | ~16M input tokens/s |
| Peak output tokens/sec | 8K x 500 tokens | ~4M output tokens/s |
| Total tokens/sec (peak) | 8K x 2.5K tokens | ~20M tokens/s |
| GPU decode capacity (A100 80GB equivalents) | 4M output tokens/s ÷ ~500 tokens/s/GPU effective generation rate | ~8K GPUs for decode |
| Naive all-token sizing check | 20M total tokens/s ÷ ~500 tokens/s/GPU = 40K | 40K is a naive upper bound, not a valid fleet count because 500 tokens/s is a decode assumption. Prefill requires a separate benchmark. |
| Chat history storage | 100M DAU x 50 msgs x 1KB | ~5 TB/day logical payload before replication and index overhead |
Latency Budget: TTFT vs TPOT
| Phase | Component | Budget | Notes |
|---|---|---|---|
| Request | Auth + rate limits + required input safety checks | 50ms | Runs before model scheduling. Keep checks bounded so they do not consume the full TTFT budget. |
| Prefill | Context fetch (DynamoDB / Cassandra) | 50ms | Partition lookup by conversation_id plus bounded range read for the required context slice. Cache hot sessions in Redis. |
| Prefill | Prompt tokenization | 20ms | 2K tokens on CPU. Tokenizer choice and language mix affect this budget |
| Prefill | GPU prefill (attention over prompt) | 400 to 800ms | Dominates TTFT. KV cache is written here. |
| Prefill | Scheduler queue wait | 0 to 500ms | Grows with batch size. Tune separately for chat and batch jobs. |
| Decode | TPOT per token (batched decode) | 30 to 50ms | ≈20 to 30 tokens/sec perceived speed |
| Decode | SSE flush to client | 5ms | Per chunk target. Transport and client buffering can add jitter. |
TTFT target: context fetch + tokenization + GPU prefill + scheduler queue wait should fit the stated 1.5s p99 target. TPOT target: 30 to 50ms keeps generation ahead of reading speed (~250 WPM). Prefix aware routing to a GPU with a warm reusable prefix can reduce prefill work by up to 60% in favorable follow up cases.
Architecture Diagram
Walk your interviewer through a split that mirrors a production chat system. The application side handles authentication, rate limiting, session state, and streaming. The inference side handles model routing, continuous batching, and KV cache locality. The API keeps a long lived SSE stream, the Context Manager assembles only the context needed for the model window, and prefix aware routing reuses cached prompt prefixes when they remain available.
With capacity requirements of approximately 8K messages/sec peak and 20M tokens/sec aggregate throughput, prefix aware routing to GPUs with warm KV caches can reduce prefill work by up to 60% in favorable multi turn cases. Chat history lands in a write optimized NoSQL store partitioned by conversation_id, with bounded range reads for the context slice needed on each turn.
In the room
Budget TTFT (prefill over the full context window) separately from TPOT (decode per token), as they rely on different tuning knobs. Larger continuous-batching batch sizes raise throughput but add queue wait times that inflate TTFT. Ask how context window overflow is handled before exploring summarization, sliding window policies, or retrieval augmented generation.
Component Deep Dives
The synchronous chat path authenticates the request, enforces request and token limits, applies required input safety checks, fetches the needed conversation state, builds the prompt within the context budget, routes to a GPU with prefix affinity when useful, and streams tokens over SSE. The assistant turn is committed idempotently using conversation_id and message_seq, and the final response state is marked complete only after generation terminates normally or is cancelled. Summarization, analytics, model experimentation, and other noncritical work stay asynchronous. Streaming output safety checks that are required by policy remain on the generation path.
1. Streaming via Server-Sent Events (SSE)
Never wait for a full completion before responding to the user. Server-Sent Events stream tokens incrementally as the GPU generates them, reducing perceived latency even when full generation takes 10 seconds or longer. For transport details, see Network Protocols (HTTP, gRPC, WebSocket).
Waiting for a 500-word response to fully generate on the GPU before sending it to the client takes 10+ seconds, creating an unacceptable user experience. We use Server-Sent Events (SSE) (or WebSockets) to stream chunks to the client as soon as the GPU produces them. SSE fits this server to client stream because it uses ordinary HTTP and does not require a bidirectional protocol. The client should detect disconnects, and the server can support resumable streams with event IDs plus a short replay buffer.
HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache
Connection: keep-alive
X-Accel-Buffering: no
id: chunk_1
data: {"id": "msg_1", "object": "chat.completion.chunk", "choices": [{"delta": {"content": "The"}}]}
id: chunk_2
data: {"id": "msg_1", "object": "chat.completion.chunk", "choices": [{"delta": {"content": " capital"}}]}
id: chunk_3
data: {"id": "msg_1", "object": "chat.completion.chunk", "choices": [{"delta": {"content": " of"}}]}
id: chunk_4
data: [DONE]2. Context Manager & Prompt Builder
The inference model is stateless between requests, so the Context Manager reconstructs the prompt from persisted conversation state on each turn and enforces the model context window budget. It should fetch the required recent messages plus any compact summary or retrieved older context rather than blindly replaying the entire history.
The model computes each request from the tokens supplied in that request and does not persist conversation memory by itself. To provide conversational memory, the server fetches the recent message slice, system instructions, and any retained summary or retrieved context, then assembles that state into the prompt before sending it to the model.
[
{"role": "system", "content": "You are a helpful AI assistant. Limit answers to 1 sentence."},
{"role": "user", "content": "Hi"},
{"role": "assistant", "content": "Hello! How can I help?"},
{"role": "user", "content": "What's the capital of France?"}
]If the history exceeds the model context window (such as 8K or 128K tokens), the Context Manager must intervene. It can apply truncation with a sliding window, summarize older blocks into a dense memory block, or use retrieval augmented generation to fetch only semantically relevant past messages using techniques covered in Vector Embeddings and Approximate Nearest Neighbor Search. A strong production design can combine these techniques by retaining a compact summary while retrieving specific older turns only when needed. Summarization can run asynchronously, with a synchronous fallback to recent history when a fresh summary is unavailable. The invariant is prompt_tokens + max_output_tokens ≤ model_context_limit, with explicit headroom for system instructions, tool metadata, and other request overhead.
3. The Inference Engine & GPU Optimization ⭐
Production inference requires specialized runtimes such as vLLM or TensorRT-LLM rather than raw PyTorch scripts. Continuous batching and paged KV cache management improve GPU utilization and reduce memory fragmentation.
Continuous Batching (Iteration-Level Scheduling)
In traditional request level batching, the GPU waits for the longest request in the batch to finish before accepting new requests.Continuous Batching ejects finished requests at the exact token iteration they finish and inserts new requests instantly. This generally improves GPU utilization and throughput by keeping the scheduler supplied with runnable work.
KV Cache & PagedAttention
The KV Cache stores the internal self attention state tensors, Keys and Values, for previously generated tokens so they do not have to be recomputed for every new token. However, KV cache memory is highly dynamic and can fragment physical VRAM.PagedAttention, as implemented by vLLM, manages KV cache blocks in pages so logical blocks need not occupy one contiguous physical VRAM region. This reduces fragmentation and can increase effective batching capacity, with the actual gain depending on model architecture, sequence lengths, memory budget, and workload.
4. Model Router & Prefix Caching
Prefix aware routing directs follow up messages to a GPU worker that still holds the reusable prompt prefix in VRAM, reducing prefill work by up to 60% in favorable multi turn cases. Similar local context caching patterns power advanced code agents such as our AI Coding Assistant (Cursor).
A specialized L7 router sends requests to inference workers. Advanced routers implement prefix aware routing: when a large reusable prefix, such as a 10,000 word document, is known to be cached, the router prefers the worker holding that prefix KV state. The unchanged prefix can avoid repeated prefill work, but newly appended user turns still require their own prefill and decode steps. Load aware fallback prevents a hot prefix from concentrating traffic on one worker.
API Design
A non-streaming request would force the client to wait for the full model response. We therefore use Server-Sent Events (SSE) for incremental token delivery. The API authenticates the request, applies token and request limits, assigns an idempotent generation identifier, and keeps the conversation state durable.
export type ChatRole = "system" | "user" | "assistant";
export interface ChatCompletionRequest {
generation_id: string;
conversation_id: string;
message: {
role: "user";
content: string;
};
stream: true;
max_output_tokens: number;
}
export interface ChatCompletionStreamEvent {
id: string;
choices: Array<{
delta: { content?: string };
finish_reason?: "stop" | "length" | "cancelled" | "error" | null;
}>;
}
export interface ChatCompletionCancelRequest {
generation_id: string;
}Chat Completion Stream
HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache
Connection: keep-alive
X-Accel-Buffering: no
id: chunk_1
data: {"id": "msg_1", "object": "chat.completion.chunk", "choices": [{"delta": {"content": "The"}}]}
id: chunk_2
data: {"id": "msg_1", "object": "chat.completion.chunk", "choices": [{"delta": {"content": " capital"}}]}
id: chunk_3
data: {"id": "msg_1", "object": "chat.completion.chunk", "choices": [{"delta": {"content": " of"}}]}
id: chunk_4
data: [DONE]Prompt Construction
The API server authenticates and authorizes the request, applies request and token limits, assigns or validates generation_id, constructs the required context window, and then queries the LLM Engine.
[
{"role": "system", "content": "You are a helpful AI assistant. Limit answers to 1 sentence."},
{"role": "user", "content": "Hi"},
{"role": "assistant", "content": "Hello! How can I help?"},
{"role": "user", "content": "What's the capital of France?"}
]Stream Resume
Each streamed chunk carries a stable SSE event ID. A reconnecting EventSource client can send the last received event ID through the Last-Event-ID header so the service can replay missing buffered chunks when available. If the replay window has expired, the service reconciles the reconnect against the same generation_id and persisted partial response state rather than creating a duplicate assistant turn. It should not silently start a new model generation unless the client explicitly requests a retry.
Generation Control
Clients can stop an in flight generation through a cancellation request keyed by generation_id. The inference scheduler should cancel queued work immediately and stop active decoding at a safe token boundary, while the conversation store preserves the partial assistant response with a cancellation finish reason. Repeated cancellation requests are idempotent.
Data Model
Because the inference model is stateless between requests, the Context Manager must fetch the required conversation state for each turn and construct the prompt within the model context budget. We use a highly scalable NoSQL database such as DynamoDB or Cassandra, partitioned by conversation_id with ordered message ranges for efficient bounded reads. The server assigns a monotonic message_seq per conversation for stable ordering. Idempotent writes keyed by generation_id protect against client retries.
{
"conversation_id": "conv_12345",
"message_seq": 1842,
"message_id": "msg_001",
"generation_id": "gen_001",
"user_id": "usr_987",
"role": "user",
"content": "What's the capital?",
"created_at": "2023-10-01T12:00:00Z",
"tokens": 6,
"status": "complete"
}Partition by conversation_id and order by the monotonic message_seq. Avoid a global user index on the hot chat path. If user wide history listing is required, maintain a separate user_conversations access pattern keyed by user_id and ordered by last_activity_at rather than a global message index.
Fault Tolerance
| Failure Case | System Solution Design |
|---|---|
| GPU Node OOM (Out of Memory) | Inference engine monitors KV cache allocation, enforces admission control before VRAM exhaustion, and queues or rejects requests when capacity is constrained. Failed nodes are removed from routing and restarted or replaced. |
| High Traffic Spikes | Scale-out can take minutes because model weights are large (10-100GB). Use dynamic queues, admission control, prewarmed standby capacity, and user visible wait states when demand exceeds warm capacity. |
| Database Slowdown | Use a highly scalable NoSQL database (DynamoDB / Cassandra) partitioned by conversation_id with bounded range reads for the required context slice. |
| Client Disconnect During Generation | Detect stream disconnects, cancel unnecessary GPU work when safe, and retain a short replay buffer keyed by stream event ID so reconnecting clients can recover missing chunks. |
| Inference Timeout or Stalled Stream | Apply per generation deadlines and heartbeat monitoring. Cancel stalled work, release KV cache state, and retry only when the generation is safe to repeat idempotently. |
Additional Considerations
Observability and Capacity Signals
Track TTFT p50 and p99, TPOT p50 and p99, scheduler queue wait, active generation count, GPU utilization, KV cache occupancy, prefix cache hit rate, cancellation rate, stream reconnect rate, context truncation frequency, safety decision latency, and tokens generated per dollar. Alert separately on user visible latency and infrastructure saturation so queue growth is detected before it becomes an outage.
Interview Walkthrough
- 25 minute cut
Skip deep architectural variants unless targeting staff level.
- Server-Sent Events streaming protocol (8 min)
- Context manager and prompt assembly from database history (9 min)
- GPU continuous batching and KV cache affinity routing (8 min)
- The inference model is stateless between requests. The server reconstructs the required prompt from persisted conversation state on every turn.
- Stream tokens via SSE because the response is primarily server to client and fits ordinary HTTP. Never wait for full completion before sending output.
- Split TTFT into queue wait plus prefill and track TPOT separately because they represent distinct tuning knobs.
- GPU optimization: use continuous batching and paged KV cache management, combined with prefix routing for reusable follow up prefixes.
- Handle context window overflow with sliding window truncation, periodic summarization, or RAG over past messages.
- Enforce rate limits based on token counts rather than request counts, and maintain a strict maximum token cap per response.
- Common pitfall: building a synchronous full response API where a 10 second wait degrades user experience. Always stream output incrementally.
Engineering Trade-offs
Open Weights vs Managed API
Using managed model APIs is easy and requires no GPU fleet operations, but token costs scale with usage and enterprise requirements may impose additional privacy, residency, and governance constraints. Hosting open weight models such as Llama 3 or Mistral on AWS or GCP GPUs requires substantial MLOps engineering and upfront capacity investment, but it provides more control over data handling, model placement, and cost at sufficiently high utilization.
Latency (TTFT) vs. Throughput (Batch Size)
To process requests efficiently, the inference engine batches decode steps and can combine prefill work where the runtime permits. Larger batches generally improve throughput but can increase queueing and memory pressure, which raises TTFT. Tune the maximum batch size and scheduler policy based on whether the workload prioritizes interactive chat latency or offline throughput.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.