Interview Setup
Interview Prompt
Design a fraud detection system that scores every payment transaction in under 100ms, combining rules and ML, with a human review queue for borderline cases.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Fail-open or fail-close when fraud service is down? |
|
| Precision vs recall trade-off? | 95% fraud recall target with under 5% false positives, because blocking legitimate users costs more than small fraud losses. |
| Real-time features vs batch features? |
|
| Graph-based ring detection in sync path? |
|
Scope
In scope
- Real-time feature computation
- Rule engine + ML ensemble
- Graph-based fraud ring detection
- Case management and analyst feedback
- Model retraining pipeline
Out of scope (state explicitly)
- Payment gateway design
- SAR regulatory filing workflow
- GNN training infrastructure
Functional Requirements
Start by asking your interviewer which fraud channels are in scope: payments only, or also account creation, login, and promo abuse. Confirm whether ML scoring or rule-based velocity checks are expected for this round.
- Real-time scoring: Score every transaction/event for fraud risk in < 100 ms
- Rule engine: Configurable rules (velocity checks, amount limits, geo-anomalies)
- ML models: Machine learning models for pattern detection (supervised + unsupervised)
- Case management: Queue suspicious events for human analyst review
- Block/allow decisions: Auto-block high-risk, auto-allow low-risk, manual review for medium
- Feature store: Real-time and historical features (user behavior, device fingerprint, transaction patterns)
- Feedback loop: Analyst decisions feed back into ML model training
- Multi-channel: Detect fraud across payments, account creation, login, promo abuse
Non-Functional Requirements
Your interviewer will care most about the recall vs precision tradeoff on the payment critical path. Call out the 100 ms budget and what happens when the fraud service is down, where choosing fail-open or fail-closed is a key product decision.
- Low Latency: Fraud decision in < 100 ms (in transaction critical path)
- High Recall: Catch > 95% of fraud (false negatives are costly)
- Acceptable Precision: False positive rate < 5% (blocking legitimate users is costly too)
- Scale: 50K+ transactions/sec scoring
- Availability: 99.99%: fraud system down means either blocking all or allowing all
- Adaptability: New fraud patterns detected and rules deployed within hours
Capacity Estimations
Run this math before you size the feature store. Transactions per second and features per score tell you Redis and ClickHouse pressure; analyst queue depth drives case-management staffing.
| Metric | Calculation | Value |
|---|---|---|
| Transactions scored / sec | Derived from daily volume ÷ 86400 (+ peak factor) | 50K |
| ML features per transaction | Given | ~200 |
| Feature computation latency | Given | < 30 ms |
| Model inference latency | Given | < 20 ms |
| Rule evaluation latency | Given | < 10 ms |
| Total fraud decision latency | Given | < 100 ms |
| Fraud rate | Given | ~0.5% of transactions |
| Manual review queue | Given | ~50K cases/day |
Architecture Diagram
In the room: draw sync scoring (<50ms on auth path) separately from nightly batch graph analysis, because mule rings need both lanes.
Walk your interviewer through the synchronous scoring path first. Fraud scoring sits on the payment authorization path: every transaction gets 200 features assembled from Redis (real-time velocity) and ClickHouse (historical behavior) before an ensemble model returns approve, review, or decline in under 50ms. Batch graph analysis runs nightly to catch mule rings the real-time tier cannot see; I draw these as two separate lanes. This fraud pipeline protects balance transactions like those designed in Digital Wallet System and platform ratings in Review and Rating System, utilizing real-time event aggregation from Stream Processing Basics.
Component Deep Dives
Feature Computation: Real-Time and Historical
Next we walk through each component on the diagram. Starting with feature computation, approximately 200 features must assemble in under 30 ms before any model executes.
Split real-time velocity features (Redis) from historical behavior (ClickHouse cache), because interviewers will ask where each lives and why.
200 features per transaction, computed in < 30 ms:
- Real-time features (Redis, < 5 ms): transaction_count_last_1h, transaction_amount_last_24h, unique_merchants_last_7d, device_fingerprint_seen_before, ip_address_country
- Historical features (pre-computed, ClickHouse → Redis cache): avg_transaction_amount_30d, typical_transaction_hour, account_age_days
- Derived features (computed at scoring time): amount_deviation, geo_velocity (distance_from_last_txn / time_since_last_txn), is_new_device, is_new_merchant
Feature store architecture: Flink consumes transaction events → updates real-time features in Redis. Spark (nightly) computes historical features. At scoring time: Feature Service reads from Redis → constructs 200-dim feature vector.
ML Model Architecture: Ensemble
1. XGBoost (primary, supervised): Trained on labeled data with all 200 features Output: P(fraud) 0.0-1.0, fast inference (< 5 ms) 2. Autoencoder (anomaly detection, unsupervised): Trained on legitimate transactions only High reconstruction error = anomaly = potential fraud Catches NEW fraud patterns not in labeled training data 3. Graph Neural Network (network analysis): Detect fraud rings: cluster of accounts sharing devices/IPs Run offline, flag suspicious clusters for enhanced scrutiny Scoring: final_score = 0.5 * xgboost_score + 0.3 * autoencoder_anomaly + 0.2 * graph_risk score < 0.3: ALLOW, score 0.3-0.7: REVIEW, score > 0.7: BLOCK
Rule Engine: Fast, Configurable
{
"rule_id": "R001", "name": "high_amount_new_account",
"condition": "transaction.amount > 500 AND user.account_age_days < 7",
"action": "review", "priority": 10
},
{
"rule_id": "R002", "name": "impossible_travel",
"condition": "geo_velocity_kmh > 1000",
"action": "block", "priority": 1
},
{
"rule_id": "R003", "name": "velocity_breach",
"condition": "transaction_count_last_1h > 20",
"action": "block", "priority": 2
}API Design
Score Transaction API
Synchronous scoring endpoint evaluated by payment authorization services within the 100 ms latency budget.
POST /api/v1/fraud/score
{
"event_type": "payment",
"transaction_id": "txn-uuid",
"user_id": "user-uuid",
"amount": 599.99,
"merchant_id": "m-uuid",
"device_fingerprint": "fp-abc",
"ip_address": "203.0.113.42",
"timestamp": "2026-03-14T11:00:00Z"
}
Response: 200 OK (< 100 ms)
{
"decision": "allow",
"score": 0.15,
"risk_factors": ["new_device"],
"rule_triggers": []
}Analyst Feedback API
Endpoint used by fraud investigation teams to confirm fraudulent behavior or mark false positives for retraining.
POST /api/v1/fraud/feedback
{ "case_id": "case-uuid", "decision": "fraud_confirmed", "analyst_id": "..." }Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
Redis: Real-Time Features
txn_count_1h:{user_id} -> INT, TTL 3600
txn_amount_24h:{user_id} -> Sorted Set, TTL 86400
device_history:{user_id} -> SET of fingerprints
last_location:{user_id} -> Hash { lat, lng, timestamp }
ip_reputation:{ip} -> FLOAT (risk score)PostgreSQL: Rules and Cases
CREATE TABLE fraud_rules (
rule_id VARCHAR(20) PRIMARY KEY, name VARCHAR(100),
condition TEXT NOT NULL, action ENUM('allow','review','block'),
priority INT, active BOOLEAN DEFAULT TRUE
);
CREATE TABLE fraud_cases (
case_id UUID PRIMARY KEY, transaction_id UUID,
user_id UUID, score DECIMAL(4,3),
decision ENUM('pending','fraud_confirmed','legitimate','escalated'),
auto_decision VARCHAR(10), analyst_id UUID,
created_at TIMESTAMPTZ DEFAULT NOW()
);Event Bus Design (Kafka)
Topic: fraud-scoring-events
Partitions: 128
Partition key: transaction_id
Retention: 90 days (model training + compliance)
Producer: Fraud Scoring Service after sync decision (ALLOW/REVIEW/BLOCK)
Event: { txn_id, user_id, score, decision, features_hash, rule_hits[], model_version }
Consumer groups:
1. case-management: REVIEW decisions to analyst dashboard queue
2. feature-updater: Flink updates velocity and device history in Redis
3. analytics: ClickHouse precision and recall per category
4. feedback-pipeline: analyst decisions to labeled training data for weekly XGBoost retrain
Sync path: score + decide < 100ms (rules + XGBoost in-process)
Async path: case mgmt, feature updates, model training never block payment
DLQ: fraud-scoring-events-dlqFault Tolerance
| Concern | Solution |
|---|---|
| Fraud service down |
|
| Feature store lag |
|
| ML model error |
|
| False positive spike |
|
| Feedback loop delay |
|
Fail-Open vs Fail-Close
Hybrid (recommended):
if transaction.amount < 50 AND user.account_age > 90 days:
allow (low risk, fail-open)
elif transaction.amount > 500:
decline (high risk, fail-close)
else:
allow + queue for async review (medium risk)Additional Considerations
Interview Walkthrough
- 25-minute cut
Skip arch50/arch75 depth unless staff.
- Latency budget: scoring must complete in <50 ms (5 min)
- Tiered pipeline: rules first, ML second (6 min)
- Flink-updated velocity counters in Redis feature store (5 min)
- Hybrid fail-open/fail-close by transaction amount (5 min)
- Offline GNN batch pipeline for fraud rings (4 min)
- Frame as a latency-budget problem: scoring must complete in <50 ms, prioritizing rules first (deterministic, explainable) and ML second (novel patterns).
- Walk through the tiered pipeline: rule engine blocks known patterns instantly → feature store lookup → ML model score → decision (allow/review/decline).
- Explain feature freshness: Flink-updated velocity counters in Redis with staleness detection that biases toward REVIEW when features lag.
- Cover hybrid fail-open/fail-close: auto-allow low-amount trusted users, auto-decline high-amount when scorer is down, queue medium for async review.
- Mention offline GNN batch pipeline for fraud rings, where nightly community detection writes cluster risk scores back to Redis for O(1) lookup at scoring time.
- Discuss champion/challenger model deployment with shadow scoring logged to ClickHouse before promoting a new model version.
- Common pitfall: pure fail-close when the ML service hiccups, because blocking every transaction costs more revenue than the fraud it prevents.
Why Both Rules AND ML
Rules: deterministic, explainable, instant deployment for known patterns. "Block all transactions from sanctioned countries" -> rule, not ML. ML: catches novel patterns, handles complex feature interactions. "User's spending pattern changed subtly over 3 weeks" -> ML, not rules. Together: Rules catch known fraud immediately. ML catches new fraud patterns. Defense in depth: if ML misses it, rules may catch it (and vice versa).
Feature Freshness & Staleness Handling
Detection: Every feature hash includes _updated_at timestamp At scoring time: feature_age = now() - features._updated_at Strategy 1: Conservative bias If velocity features are stale → add 0.15 to fraud score Rationale: "I can't see recent activity → assume higher risk" Strategy 2: Default substitution Replace stale feature with population median Strategy 3: Feature importance gating If top-5 most important features are stale → route to REVIEW (skip ML)
Weekly vs Daily Model Retraining
Hybrid approach (recommended): Base model: retrained weekly (stable foundation) Rule engine: updated within HOURS for known new patterns → Analyst spots new pattern → creates rule → deployed in 30 min → Rule catches new fraud immediately while model takes a week to learn
Engineering Trade-offs
Model Version Management During Canary Deploys
Fraud systems balance false positives against false negatives, weighing auto-decline precision against manual review queue capacity.
Two models always loaded in memory: Champion: current production model (v12) Challenger: candidate model (v13) Dual scoring: every transaction scored by BOTH models Only champion's decision is enforced Both scores logged to ClickHouse for offline comparison Canary deployment stages: Week 1: Shadow mode (100% champion enforced, challenger logged only) Week 2: 5% canary (challenger enforced on 5% traffic) Week 3: 50% split Week 4: 100% promotion Rollback: Config flag in Redis, flip from "v13" to "v12" → 1 second
Graph Neural Network (GNN): Production Architecture
GNN runs as OFFLINE batch pipeline (not in real-time scoring path):
Nightly pipeline (Spark):
1. Build transaction graph: 500M nodes, 2B edges
2. Community detection (Louvain algorithm): detect fraud rings
3. Score clusters: if cluster_fraud_rate > 10% → all members HIGH risk
4. Write to Redis: HSET graph:risk:{user_id} score 0.85
At scoring time: graph_risk = HGET graph:risk:{user_id} → <1ms lookup
Cost: ~4 hours nightly compute (Spark cluster)
Value: catches fraud rings that individual-transaction models missFail-Open vs Fail-Close: Threshold-Based Decision
Revenue impact analysis (5 min downtime):
Average fraud service downtime: 5 min/month
During 5 min downtime with threshold-based failover:
~1,500 transactions processed
~750 low-risk auto-allowed → 0 expected fraud
~250 high-risk auto-declined → $12,500 lost revenue
~500 medium-risk auto-allowed → ~$250 fraud loss
Compare: fail-close (block all) → $75,000 lost revenue
Compare: fail-open (allow all) → $3,750 fraud loss
Threshold-based: best of both worldsReview
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.