Interview Setup
Interview Prompt
Design a workflow orchestration engine that coordinates multi step, long running business processes across distributed services. Workflows must survive crashes, restarts, and deployments.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| How long do workflows run: minutes, days, or months? | Timers and human approval signals require durable state, not just a job queue. |
| Do we need compensation (saga rollback) or is retry enough? | Use sagas for transactions that span multiple services. Use simple retries for idempotent single-step failures. |
| Build from scratch or use Temporal/Cadence? | Staff candidates should discuss both approaches and explain why most product teams buy unless workflow orchestration itself is a core platform capability. |
| What are the activity failure semantics: at least once or exactly-once side effects? | Activities run at least once, so side effects need stable idempotency keys. Workflow decisions are deterministically reconstructed from history during replay. |
Scope
In scope
- Durable execution via event sourced replay
- Activity dispatch with retry policies
- Saga compensation pattern
- Workflow versioning for safe deploys
Out of scope (state explicitly)
- Full Temporal server implementation details (discuss architecture)
- UI/workflow designer (assume code based workflows)
- High-throughput periodic job scheduling
Functional Requirements
Start by asking your interviewer how long workflows run (whether seconds, days, or months with human approval signals) and whether saga compensation is required or simple retry suffices. Contrast with a job queue early so they know you understand durable execution is a fundamentally different problem.
- Long-running, multi step workflows spanning seconds to months
- Durable execution: workflow state survives crashes, restarts, deployments
- Activity tasks: individual units of work executed by workers
- Timers and sleep: workflow.sleep(Duration.ofDays(30))
- Retry policies: automatic retry with configurable backoff
- Saga pattern: compensating transactions for distributed rollback
- Child workflows: compose workflows hierarchically
- Signals and queries: send events TO or read state FROM running workflows
- Versioning: update workflow code without breaking in flight executions
- Visibility: search, filter, monitor workflows by custom attributes
Non-Functional Requirements
Your interviewer will stress-test the distinction between deterministic workflow replay and at least once activity side effects. They will also probe what happens when you deploy new workflow code while 10,000 in flight executions are still running, which is precisely what getVersion() resolves.
- Durability: 100% of acknowledged workflow state must survive infrastructure failures through the configured durable history store.
- Scalability: 100K+ concurrent workflows and 10K peak starts/sec.
- Low Latency: Activity dispatch targets < 50ms p99, and workflow decisions target < 10ms p99.
- Effectively Once Side Effects: Workflow decisions are deterministically reconstructed from history during replay, while activities run at least once and require idempotent handlers.
- High Availability: Target 99.99% availability because prolonged workflow service unavailability can block critical business processes.
- Observability: Full audit trail of every state transition
Capacity Estimations
Event history storage grows with every workflow step. If the 10,000 workflow starts per second peak rate were sustained, approximately 86 TB/day of history would be generated and would become the dominant cost driver. The 100,000 concurrent workflow figure is a separate capacity target, while the 5-minute duration assumption implies about 333 average starts per second under steady state. Size Cassandra or PostgreSQL for append-only writes and plan retention policies early.
| Metric | Calculation | Value |
|---|---|---|
| Concurrent workflows | Given capacity target | 100K |
| Peak workflow starts / sec | Peak workload capacity | 10K |
| Peak activity dispatch capacity | 10K workflow starts/sec x 10 activities/workflow if all activities dispatch immediately | 100K/sec |
| Concurrency sanity check | 100K concurrent workflows ÷ 300 seconds | ~333 average starts/sec at 5-minute average duration |
| Avg activities per workflow | Given (typical workload assumption) | 10 |
| Avg workflow duration | Given (typical workload assumption) | 5 minutes (some: months) |
| Event history per workflow | avg 50 events x 2 KB | 100 KB |
| Peak history write throughput | 10K starts/sec x 100 KB/workflow | ~1 GB/sec if the peak start rate is sustained |
| History storage if peak rate is sustained | 10K starts/sec x 50 events/workflow x 2 KB/event x 86400 | ~86 TB/day |
| History retention at sustained peak (30 days) | ~86 TB/day x 30 | ~2.5 PB |
Architecture Diagram
In the interview, draw three distinct roles: Temporal Server components for Frontend, History, Matching, and Visibility, workflow workers for deterministic replay, and activity workers for side effects. The server never executes your business logic. It schedules tasks, persists history, and exposes workflow state to the visibility layer.
Every workflow decision is recorded as an append-only event in the durable history store. The canonical design uses Cassandra, while PostgreSQL is a valid alternative. On worker crash, a new worker replays history from event 1, skips completed activities using the recorded results, and resumes at the first incomplete step. Separate execution metadata may still exist for task management and indexing, but it is not the authoritative workflow state.
Activities call external services such as payments, shipping, and email, and run at least once. Side effects must be idempotent through stable keys. Workflow decisions are deterministically reconstructed from history during replay, while Kafka only fans out visibility events to external systems and is not the source of truth. Use a stable workflow ID and explicit reuse policy so client retries do not create duplicate workflows.
In the room
Explain that workflow code must be deterministic. It should not use random numbers, direct wall-clock reads, or direct I/O. Temporal APIs provide deterministic time and signal handling, while Activities own external side effects and run at least once. This split is the first thing staff interviewers probe.
Component Deep Dives
How Durable Execution Works: The Core Innovation
Durable execution via event replay is Temporal's core innovation. Walk through replay first, then explain saga compensation, human approval through Signals, and workflow versioning. Draw the server components after the replay model is clear.
// Traditional approach (stateless service):
function processOrder(order) {
payment = chargeCard(order) // Process crashes here
shipment = createShipment(order) // Never reached
}
// On crash: partial state remains, requiring manual DB recovery
// Temporal approach (durable execution):
function processOrder(order) {
payment = workflow.executeActivity(ChargeCard, order) // Step 1
shipment = workflow.executeActivity(CreateShipment, order) // Step 2
}
// On crash: Temporal replays history, returns the recorded ChargeCard result,
// and resumes execution at step 2 without re-running the completed activity}HOW THIS WORKS:
1. Worker replays event history:
Event 1: WorkflowExecutionStarted
Event 2: ActivityTaskScheduled (ChargeCard)
Event 3: ActivityTaskCompleted (result: {payment_id: "pay_123"})
Workflow execution paused, with history recorded up to here
2. Worker re-executes deterministic code:
- the recorded ChargeCard result is consumed from history
- createShipment emits the NEW command
3. Worker returns command: ScheduleActivityTask(CreateShipment)
CRITICAL RULE: Workflow code must be DETERMINISTIC
No random(), no direct wall-clock reads, and no direct I/O calls
Use Temporal workflow APIs for time and signals. All external side effects execute through registered activitiesEvent History (Event-Sourced State)
The event history is the authoritative workflow state. Separate execution metadata can still support indexing, task management, and leases. Show a concrete event sequence so your interviewer sees replay mechanics rather than an abstract concept.
Every workflow execution is an append-only event log:
Event 1: WorkflowExecutionStarted {input: {orderId: "42"}}
Event 2: WorkflowTaskScheduled
Event 3: WorkflowTaskCompleted {commands: [ScheduleActivity(ChargeCard)]}
Event 4: ActivityTaskScheduled {activityType: "ChargeCard"}
Event 5: ActivityTaskStarted {attempt: 1}
Event 6: ActivityTaskCompleted {result: {paymentId: "pay_123"}}
Event 7: WorkflowTaskScheduled
Event 8: WorkflowTaskCompleted {commands: [ScheduleActivity(CreateShipment)]}
Event 9: TimerStarted {duration: 30 days}
...
Event 42: TimerFired
Event 43: WorkflowTaskCompleted {commands: [CompleteWorkflow]}
Event 44: WorkflowExecutionCompleted {result: {status: "success"}}
This history is the authoritative logical workflow state. Execution metadata may still live in separate tables for indexing, leases, and task management.2.5 PB History Retention and Archival
If the 10,000 workflow starts per second peak rate were sustained, history storage would become the dominant cost. Continue-As-New, archival, and retention controls are mandatory operational requirements.
If the 10,000 workflow starts/sec peak rate were sustained, event history grows ~86 TB/day, reaching ~2.5 PB over a 30-day retention window. Continue-As-New starts a fresh run with the same Workflow ID when active history becomes too large, carrying forward the required state and keeping replay bounded. Archival moves completed histories older than 7 days to S3 or GCS object storage, while the visibility index retains search metadata. Cassandra tuning requires increasing compaction throughput with ConcurrentCompactors, separating history and visibility keyspaces, and monitoring SSTable count per shard. Without proactive history bounding, archival, and retention management, history storage cost dominates the entire platform budget.
Saga Pattern (Distributed Compensation)
Saga compensation belongs directly in workflow code as standard try/catch blocks rather than being scattered across message queues, because orchestration wins when rollback paths must remain auditable. See Distributed Transactions: 2PC vs Saga and Microservices Patterns: Saga and Outbox.
// Problem: Book a trip = Flight + Hotel + Car Rental (3 services)
Temporal Saga:
function bookTrip(trip) {
flight = workflow.executeActivity(BookFlight, trip)
try {
hotel = workflow.executeActivity(BookHotel, trip)
} catch (e) {
workflow.executeActivity(CancelFlight, flight) // compensate
throw e
}
try {
car = workflow.executeActivity(RentCar, trip)
} catch (e) {
workflow.executeActivity(CancelHotel, hotel) // compensate
workflow.executeActivity(CancelFlight, flight) // compensate
throw e
}
return {flight, hotel, car}
}
Why better than choreography (event driven):
- Compensation logic stays in one workflow definition instead of being scattered across event consumers
- Temporal can retry compensation activities even when a worker crashes
- Compensation failures remain visible and can be retried or escalated without losing workflow history
Activity Idempotency (At-Least-Once Delivery)
Activities execute at least once by design. Idempotency keys must remain stable across retries rather than incorporating attempt numbers, which would otherwise cause duplicate operations.
Workflow start idempotency:
1. Choose a stable workflow_id for the business operation (for example order-42).
2. Start the workflow with an explicit workflow-ID reuse/conflict policy.
3. If the client retries the same start request, reuse the same workflow_id so the server does not create duplicate business workflows.
Activity idempotency:
Problem: Temporal guarantees at least once activity execution. If a worker crashes after the downstream side effect commits but before activity completion is recorded, Temporal may retry the activity.
Idempotency Key Pattern:
1. Generate a stable key before calling the activity. NEVER include attempt_number because every retry must reuse the same key:
key = sha256(workflow_id + run_id + activity_id)
2. Pass key to downstream: POST /charge { idempotency_key: "abc123" }
3. Downstream stores (key -> result) for TTL=24h or for the maximum retry/recovery window, whichever is longer
4. Duplicate request returns the prior result with no duplicate side effect
Non-retryable errors:
HTTP 400: non_retryable (retrying will not resolve the client error)
HTTP 503: retryable with exponential backoff (transient failure)Visibility Events (Kafka Fan-out)
Kafka is used for visibility fan-out rather than workflow durability. The append-only workflow history is the authoritative source of truth, so be careful not to conflate messaging queues with orchestration state. Learn more about queuing mechanics in Message Queue Fundamentals.
Topic: workflow-visibility-events
Partitions: 64 (partition by workflow_id)
Events: WorkflowExecutionStarted, Completed, Failed, TimedOut, Cancelled
Retention: 30 days (ops dashboards + billing meter)
Producers: Temporal server components after the corresponding history append
Consumers: observability stack, SLA alerting, cost attribution
Topic: activity-task-events
Purpose: External operational telemetry only, because worker task dispatch still uses Temporal task queues
Partitions: 128 (partition by task_queue)
Events: ActivityTaskScheduled, Started, Completed, Failed, TimedOut
Consumers: worker autoscaling, retry alerting, retry policy tuning
Note: Temporal's source of truth is append-only event history in Cassandra or PostgreSQL,
not Kafka. Kafka fans out visibility signals to external systems only.
Start workflow: Client -> Frontend -> history write -> return run_id (< 100ms target)
Workers poll task queues. Activity side effects run outside workflow code and are isolated from deterministic workflow replayAPI Design
Temporal Client SDK (TypeScript example)
Temporal uses an SDK-first workflow programming model rather than a REST workflow interface. Demonstrate workflow start, query, signal, and result retrieval with stable workflow IDs and explicit reuse behavior for client retries. Workflow code must remain deterministic.
interface OrderInput {
orderId: string;
customerId: string;
}
interface OrderStatus {
state: "RUNNING" | "COMPLETED" | "FAILED" | "CANCELLED";
}
interface OrderResult {
status: "success" | "failed";
}
type WorkflowId = string;
interface StartWorkflowOptions {
workflowId: WorkflowId;
taskQueue: string;
workflowIdReusePolicy?: "ALLOW_DUPLICATE" | "REJECT_DUPLICATE";
}
interface WorkflowHandle<TResult> {
query<T>(name: string): Promise<T>;
signal(name: string, payload: unknown): Promise<void>;
result(): Promise<TResult>;
}
interface WorkflowClient {
start<TInput, TResult>(
workflow: string,
options: StartWorkflowOptions,
input: TInput,
): Promise<WorkflowHandle<TResult>>;
}
const handle = await client.start<OrderInput, OrderResult>("ProcessOrder", {
workflowId: "order-42",
taskQueue: "order-processing",
workflowIdReusePolicy: "REJECT_DUPLICATE",
}, {
orderId: "order-42",
customerId: "cust-123",
});
const status = await handle.query<OrderStatus>("getStatus");
await handle.signal("cancelOrder", { reason: "customer_requested" });
const result = await handle.result();Temporal Server gRPC API
The server exposes workflow lifecycle and task polling operations over gRPC. Signals become durable workflow events, while Queries are read-only state reads that do not change workflow history. Application developers normally use the language SDK rather than calling these APIs directly.
StartWorkflowExecution Start new workflow SignalWorkflowExecution Send signal QueryWorkflow Query state TerminateWorkflowExecution Force-terminate RequestCancelWorkflowExecution Request graceful cancellation ListWorkflowExecutions Search via visibility GetWorkflowExecutionHistory Full event history PollWorkflowTaskQueue Worker long-poll for workflow PollActivityTaskQueue Worker long-poll for activities
Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
Cassandra (History Store)
-- Execution state
CREATE TABLE executions (
namespace_id UUID, workflow_id TEXT, run_id UUID,
state INT, next_event_id BIGINT,
PRIMARY KEY ((namespace_id, workflow_id), run_id)
);
-- Event history (append-only, event sourced)
CREATE TABLE history_events (
namespace_id UUID, workflow_id TEXT, run_id UUID,
event_id BIGINT, event_type INT, data BLOB,
PRIMARY KEY ((namespace_id, workflow_id, run_id), event_id)
) WITH CLUSTERING ORDER BY (event_id ASC);
-- Pending tasks with lease state for worker crash recovery
CREATE TABLE tasks (
namespace_id UUID, task_queue TEXT, task_type INT,
task_id BIGINT, scheduled_time TIMESTAMP,
lease_owner TEXT, lease_expiry TIMESTAMP,
PRIMARY KEY ((namespace_id, task_queue, task_type), task_id)
);
-- Timers, bucketed by time to avoid an unbounded namespace partition
CREATE TABLE timers (
namespace_id UUID, bucket TEXT, fire_time TIMESTAMP,
workflow_id TEXT, timer_id TEXT,
PRIMARY KEY ((namespace_id, bucket), fire_time, workflow_id, timer_id)
) WITH CLUSTERING ORDER BY (fire_time ASC);Elasticsearch (Visibility)
{
"namespace": "production",
"workflow_id": "order-42",
"workflow_type": "ProcessOrder",
"status": "Running",
"start_time": "2026-03-14T10:00:00Z",
"custom_attributes": {
"customer_id": "cust-789",
"order_total": 299.99
}
}Fault Tolerance
| Concern | Solution |
|---|---|
| Workflow worker crash | Task times out on the server, gets re-dispatched, and a new worker replays history to resume |
| Activity worker crash | Task times out and the server automatically schedules retries per the activity retry policy |
| Server node crash | History shards rebalance across surviving nodes while all state remains durable in Cassandra |
| DB unavailable | The server backpressures or rejects new durable writes until persistence recovers, while clients retry with stable request identity. It does not rely on unbounded in-memory queues for durability. |
| Long activity timeout | Workers send periodic heartbeats every N seconds so the server can detect liveness and avoid treating healthy long running work as abandoned |
| Poison pill workflow | Retries back off until the configured policy is exhausted. The workflow then fails or is terminated, and a diagnostic event can be routed to an operator queue for inspection. |
| Versioning during deploy | Use getVersion() to route in flight workflows through legacy paths while new workflows use updated code |
Additional Considerations
Payload Size, Tenant Isolation, and Fairness
Keep workflow inputs, activity results, and signals small because large payloads inflate history size and replay cost. Store large documents in object storage and keep only durable references in workflow state. Enforce namespace quotas for workflow starts, concurrent executions, and task queue backlog so one tenant cannot starve others. Apply authentication, namespace scoped authorization, encryption in transit and at rest, and audit logging because workflow inputs and history can contain sensitive business data.
Production Patterns
In order processing pipelines, the workflow sequentially validates the order, reserves inventory, charges payment, schedules shipment, and dispatches customer notifications, executing a compensating saga if any intermediate step fails.
For user onboarding journeys, the workflow creates the account, sends a welcome email, sleeps for 3 days, sends an interactive tutorial, sleeps for 7 days, and delivers a promotional offer.
In subscription billing loops, the engine executes monthly recurring charges, retrying up to 3 times over 7 days upon failure, and automatically cancels the subscription if payments remain uncollectible.
For human approval, an expense report workflow submits the request, pauses indefinitely until an external approval signal arrives via webhook, and then either reimburses or rejects the transaction.
Temporal vs Saga with Events (Choreography vs Orchestration)
Choreography scatters workflow logic and error compensations across disparate microservices, making end-to-end status invisible and distributed rollbacks notoriously complex to trace. In contrast, orchestration via Temporal centralizes control flow into readable sequential code, wraps compensation logic inside standard try/catch blocks, preserves immutable audit history, guarantees transition ordering, and remains fully unit-testable. In practice, production platforms use orchestration for critical business transactions, while reserving decoupled event choreography for asynchronous notifications and analytical fan-out.
Interview Walkthrough
- 25-minute cut
Use the deeper storage, versioning, and scaling material when the interviewer probes staff-level concerns.
- Define durable execution and why ordinary queues do not provide it (5 min)
- Contrast Temporal orchestration with Kafka choreography (6 min)
- Model each business step as an idempotent Activity with a retry policy (5 min)
- Show compensation explicitly with a concrete saga rollback path (5 min)
- Explain human approval through Signals without holding worker threads (4 min)
- Define the problem as durable execution: workflows that survive process crashes, network failures, and rollouts without losing state.
- Contrast orchestration (centralized flow via Temporal) with choreography (decentralized Kafka events) using Distributed Transactions: 2PC vs Saga framing.
- Model each step as an idempotent Activity with configurable retry policies. Because Temporal replays history, side effects must be safe to repeat.
- Show compensation explicitly:
try { charge } catch { refund }, because interviewers want to see concrete rollback paths rather than just the happy path. - Cover human approval via Signals, where a workflow waits days for an approval event without holding worker threads or polling a database.
- Address workflow versioning for zero downtime deploys by pinning running instances to old code paths while new executions use updated logic.
- Quantify timeout budgets: activity start-to-close timeout, schedule-to-start timeout, and heartbeat timeout for long running tasks.
- Common pitfall: hand-rolling saga logic with Kafka consumers and cron reconciliation, because it quickly becomes unmaintainable without durable state and replay guarantees. Contrast with a dedicated Distributed Job Scheduler when only periodic execution is required.
Engineering Trade-offs
Temporal vs Airflow vs AWS Step Functions
| Dimension | Temporal | Airflow | Step Functions | |---|---|---|---| | Execution model | Durable functions | DAG scheduler | State machine (JSON) | | Replay | Full event sourced replay | Re-run whole task | Re-drive from failed state | | Long-running | Timers up to years | Task timeout limits | Standard: up to 1 year | | Saga/compensation | Built-in via code | Manual compensation | Catch + rollback | | Developer model | Code (Go/Java/TS) | Python DAGs | Amazon States Language (JSON) | Temporal shines when: 1. The flow has more than three distinct steps across services 2. Steps require granular retry and compensation logic 3. The execution spans minutes to months with durable timers or human approval 4. You require complete audit visibility into every state transition 5. You want workflow logic written in real code rather than YAML configurations
Workflow Versioning for Zero Downtime Deploys
Problem: 10,000 workflows running processOrder v1.
Deploy v2 (added new step). How to handle in flight v1 workflows?
Solution: Temporal's getVersion()
function processOrder(order) {
payment = executeActivity(ChargeCard, order)
version = workflow.getVersion("add-fraud-check", 1, 2)
if (version >= 2) {
executeActivity(FraudCheck, order) // NEW step in v2
}
shipment = executeActivity(CreateShipment, order)
}
In flight v1 workflows: getVersion returns 1 → skip FraudCheck
New workflows: getVersion returns 2 → run FraudCheck
Eventually all v1 workflows complete, so you can remove the version check in v3When to Use Temporal vs Other Approaches
| Approach | Use When | Don't Use When |
|---|---|---|
| Temporal | Multi-step, long running, retry/compensation | Simple request-response |
| Kafka + Consumers | Event-driven, fan-out, decoupled | Request-response or saga |
| Step Functions | AWS native, managed state machines, simple DAGs | Very high workflow scale or complex code first logic |
| Choreography (events) | Loose coupling, simple flows | Complex compensation, visibility |
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.