System Design Problem

Design a Fraud Detection System

Commonly Asked By:StripePayPalCoinbaseRobinhoodBlock

Interview Setup

Interview Prompt

Design a fraud detection system that scores every payment transaction in under 100ms, combining rules and ML, with a human review queue for borderline cases.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Fail-open or fail-close when fraud service is down?
  • Fail-open loses money
  • fail-close blocks revenue. Tier by txn amount: allow < $20, decline > $500.
Precision vs recall trade-off?95% fraud recall target with under 5% false positives, because blocking legitimate users costs more than small fraud losses.
Real-time features vs batch features?
  • 200 features per txn
  • velocity and device fingerprint need < 30ms from Redis
  • historical aggregates can be stale.
Graph-based ring detection in sync path?
  • Connected-component analysis is async
  • sync path uses precomputed graph features from nightly batch.

Scope

In scope

  • Real-time feature computation
  • Rule engine + ML ensemble
  • Graph-based fraud ring detection
  • Case management and analyst feedback
  • Model retraining pipeline

Out of scope (state explicitly)

  • Payment gateway design
  • SAR regulatory filing workflow
  • GNN training infrastructure

Functional Requirements

Start by asking your interviewer which fraud channels are in scope: payments only, or also account creation, login, and promo abuse. Confirm whether ML scoring or rule-based velocity checks are expected for this round.

  • Real-time scoring: Score every transaction/event for fraud risk in < 100 ms
  • Rule engine: Configurable rules (velocity checks, amount limits, geo-anomalies)
  • ML models: Machine learning models for pattern detection (supervised + unsupervised)
  • Case management: Queue suspicious events for human analyst review
  • Block/allow decisions: Auto-block high-risk, auto-allow low-risk, manual review for medium
  • Feature store: Real-time and historical features (user behavior, device fingerprint, transaction patterns)
  • Feedback loop: Analyst decisions feed back into ML model training
  • Multi-channel: Detect fraud across payments, account creation, login, promo abuse

Non-Functional Requirements

Your interviewer will care most about the recall vs precision tradeoff on the payment critical path. Call out the 100 ms budget and what happens when the fraud service is down, where choosing fail-open or fail-closed is a key product decision.

  • Low Latency: Fraud decision in < 100 ms (in transaction critical path)
  • High Recall: Catch > 95% of fraud (false negatives are costly)
  • Acceptable Precision: False positive rate < 5% (blocking legitimate users is costly too)
  • Scale: 50K+ transactions/sec scoring
  • Availability: 99.99%: fraud system down means either blocking all or allowing all
  • Adaptability: New fraud patterns detected and rules deployed within hours

Capacity Estimations

Run this math before you size the feature store. Transactions per second and features per score tell you Redis and ClickHouse pressure; analyst queue depth drives case-management staffing.

MetricCalculationValue
Transactions scored / secDerived from daily volume ÷ 86400 (+ peak factor)50K
ML features per transactionGiven~200
Feature computation latencyGiven< 30 ms
Model inference latencyGiven< 20 ms
Rule evaluation latencyGiven< 10 ms
Total fraud decision latencyGiven< 100 ms
Fraud rateGiven~0.5% of transactions
Manual review queueGiven~50K cases/day

Architecture Diagram

In the room: draw sync scoring (<50ms on auth path) separately from nightly batch graph analysis, because mule rings need both lanes.

Walk your interviewer through the synchronous scoring path first. Fraud scoring sits on the payment authorization path: every transaction gets 200 features assembled from Redis (real-time velocity) and ClickHouse (historical behavior) before an ensemble model returns approve, review, or decline in under 50ms. Batch graph analysis runs nightly to catch mule rings the real-time tier cannot see; I draw these as two separate lanes. This fraud pipeline protects balance transactions like those designed in Digital Wallet System and platform ratings in Review and Rating System, utilizing real-time event aggregation from Stream Processing Basics.

Loading...

Component Deep Dives

Feature Computation: Real-Time and Historical

Next we walk through each component on the diagram. Starting with feature computation, approximately 200 features must assemble in under 30 ms before any model executes.

Split real-time velocity features (Redis) from historical behavior (ClickHouse cache), because interviewers will ask where each lives and why.

200 features per transaction, computed in < 30 ms:

  • Real-time features (Redis, < 5 ms): transaction_count_last_1h, transaction_amount_last_24h, unique_merchants_last_7d, device_fingerprint_seen_before, ip_address_country
  • Historical features (pre-computed, ClickHouse → Redis cache): avg_transaction_amount_30d, typical_transaction_hour, account_age_days
  • Derived features (computed at scoring time): amount_deviation, geo_velocity (distance_from_last_txn / time_since_last_txn), is_new_device, is_new_merchant

Feature store architecture: Flink consumes transaction events → updates real-time features in Redis. Spark (nightly) computes historical features. At scoring time: Feature Service reads from Redis → constructs 200-dim feature vector.

ML Model Architecture: Ensemble

1. XGBoost (primary, supervised):
   Trained on labeled data with all 200 features
   Output: P(fraud) 0.0-1.0, fast inference (< 5 ms)

2. Autoencoder (anomaly detection, unsupervised):
   Trained on legitimate transactions only
   High reconstruction error = anomaly = potential fraud
   Catches NEW fraud patterns not in labeled training data

3. Graph Neural Network (network analysis):
   Detect fraud rings: cluster of accounts sharing devices/IPs
   Run offline, flag suspicious clusters for enhanced scrutiny

Scoring:
  final_score = 0.5 * xgboost_score + 0.3 * autoencoder_anomaly + 0.2 * graph_risk
  score < 0.3: ALLOW,  score 0.3-0.7: REVIEW,  score > 0.7: BLOCK

Rule Engine: Fast, Configurable

JSON
{
  "rule_id": "R001", "name": "high_amount_new_account",
  "condition": "transaction.amount > 500 AND user.account_age_days < 7",
  "action": "review", "priority": 10
},
{
  "rule_id": "R002", "name": "impossible_travel",
  "condition": "geo_velocity_kmh > 1000",
  "action": "block", "priority": 1
},
{
  "rule_id": "R003", "name": "velocity_breach",
  "condition": "transaction_count_last_1h > 20",
  "action": "block", "priority": 2
}

API Design

Score Transaction API

Synchronous scoring endpoint evaluated by payment authorization services within the 100 ms latency budget.

HTTP
POST /api/v1/fraud/score
{
  "event_type": "payment",
  "transaction_id": "txn-uuid",
  "user_id": "user-uuid",
  "amount": 599.99,
  "merchant_id": "m-uuid",
  "device_fingerprint": "fp-abc",
  "ip_address": "203.0.113.42",
  "timestamp": "2026-03-14T11:00:00Z"
}
Response: 200 OK (< 100 ms)
{
  "decision": "allow",
  "score": 0.15,
  "risk_factors": ["new_device"],
  "rule_triggers": []
}

Analyst Feedback API

Endpoint used by fraud investigation teams to confirm fraudulent behavior or mark false positives for retraining.

HTTP
POST /api/v1/fraud/feedback
{ "case_id": "case-uuid", "decision": "fraud_confirmed", "analyst_id": "..." }

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

Redis: Real-Time Features

txn_count_1h:{user_id}          -> INT, TTL 3600
txn_amount_24h:{user_id}        -> Sorted Set, TTL 86400
device_history:{user_id}        -> SET of fingerprints
last_location:{user_id}         -> Hash { lat, lng, timestamp }
ip_reputation:{ip}              -> FLOAT (risk score)

PostgreSQL: Rules and Cases

SQL
CREATE TABLE fraud_rules (
    rule_id VARCHAR(20) PRIMARY KEY, name VARCHAR(100),
    condition TEXT NOT NULL, action ENUM('allow','review','block'),
    priority INT, active BOOLEAN DEFAULT TRUE
);

CREATE TABLE fraud_cases (
    case_id UUID PRIMARY KEY, transaction_id UUID,
    user_id UUID, score DECIMAL(4,3),
    decision ENUM('pending','fraud_confirmed','legitimate','escalated'),
    auto_decision VARCHAR(10), analyst_id UUID,
    created_at TIMESTAMPTZ DEFAULT NOW()
);

Event Bus Design (Kafka)

Topic: fraud-scoring-events
  Partitions: 128
  Partition key: transaction_id
  Retention: 90 days (model training + compliance)

Producer: Fraud Scoring Service after sync decision (ALLOW/REVIEW/BLOCK)
  Event: { txn_id, user_id, score, decision, features_hash, rule_hits[], model_version }

Consumer groups:
  1. case-management: REVIEW decisions to analyst dashboard queue
  2. feature-updater: Flink updates velocity and device history in Redis
  3. analytics: ClickHouse precision and recall per category
  4. feedback-pipeline: analyst decisions to labeled training data for weekly XGBoost retrain

Sync path: score + decide < 100ms (rules + XGBoost in-process)
Async path: case mgmt, feature updates, model training never block payment
DLQ: fraud-scoring-events-dlq

Fault Tolerance

ConcernSolution
Fraud service down
  • Fail-open for low-value txns (allow with async review)
  • fail-close for high-value (decline)
Feature store lag
  • Use stale features with reduced confidence
  • increase review threshold
ML model error
False positive spike
  • Monitor auto-block rate
  • alert if > 2x normal
  • auto-switch to review mode
Feedback loop delay
  • Weekly model retrain
  • interim: update rules for new patterns within hours

Fail-Open vs Fail-Close

Hybrid (recommended):
  if transaction.amount < 50 AND user.account_age > 90 days:
    allow (low risk, fail-open)
  elif transaction.amount > 500:
    decline (high risk, fail-close)
  else:
    allow + queue for async review (medium risk)

Additional Considerations

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • Latency budget: scoring must complete in <50 ms (5 min)
    • Tiered pipeline: rules first, ML second (6 min)
    • Flink-updated velocity counters in Redis feature store (5 min)
    • Hybrid fail-open/fail-close by transaction amount (5 min)
    • Offline GNN batch pipeline for fraud rings (4 min)
  • Frame as a latency-budget problem: scoring must complete in <50 ms, prioritizing rules first (deterministic, explainable) and ML second (novel patterns).
  • Walk through the tiered pipeline: rule engine blocks known patterns instantly → feature store lookup → ML model score → decision (allow/review/decline).
  • Explain feature freshness: Flink-updated velocity counters in Redis with staleness detection that biases toward REVIEW when features lag.
  • Cover hybrid fail-open/fail-close: auto-allow low-amount trusted users, auto-decline high-amount when scorer is down, queue medium for async review.
  • Mention offline GNN batch pipeline for fraud rings, where nightly community detection writes cluster risk scores back to Redis for O(1) lookup at scoring time.
  • Discuss champion/challenger model deployment with shadow scoring logged to ClickHouse before promoting a new model version.
  • Common pitfall: pure fail-close when the ML service hiccups, because blocking every transaction costs more revenue than the fraud it prevents.

Why Both Rules AND ML

Rules: deterministic, explainable, instant deployment for known patterns.
  "Block all transactions from sanctioned countries" -> rule, not ML.

ML: catches novel patterns, handles complex feature interactions.
  "User's spending pattern changed subtly over 3 weeks" -> ML, not rules.

Together: Rules catch known fraud immediately. ML catches new fraud patterns.
  Defense in depth: if ML misses it, rules may catch it (and vice versa).

Feature Freshness & Staleness Handling

Detection: Every feature hash includes _updated_at timestamp
  At scoring time: feature_age = now() - features._updated_at
  
Strategy 1: Conservative bias
  If velocity features are stale → add 0.15 to fraud score
  Rationale: "I can't see recent activity → assume higher risk"
  
Strategy 2: Default substitution
  Replace stale feature with population median
  
Strategy 3: Feature importance gating
  If top-5 most important features are stale → route to REVIEW (skip ML)

Weekly vs Daily Model Retraining

Hybrid approach (recommended):
  Base model: retrained weekly (stable foundation)
  Rule engine: updated within HOURS for known new patterns
  → Analyst spots new pattern → creates rule → deployed in 30 min
  → Rule catches new fraud immediately while model takes a week to learn

Engineering Trade-offs

Model Version Management During Canary Deploys

Fraud systems balance false positives against false negatives, weighing auto-decline precision against manual review queue capacity.

Two models always loaded in memory:
  Champion: current production model (v12)
  Challenger: candidate model (v13)

Dual scoring: every transaction scored by BOTH models
  Only champion's decision is enforced
  Both scores logged to ClickHouse for offline comparison

Canary deployment stages:
  Week 1: Shadow mode (100% champion enforced, challenger logged only)
  Week 2: 5% canary (challenger enforced on 5% traffic)
  Week 3: 50% split
  Week 4: 100% promotion

Rollback: Config flag in Redis, flip from "v13" to "v12" → 1 second

Graph Neural Network (GNN): Production Architecture

GNN runs as OFFLINE batch pipeline (not in real-time scoring path):

Nightly pipeline (Spark):
  1. Build transaction graph: 500M nodes, 2B edges
  2. Community detection (Louvain algorithm): detect fraud rings
  3. Score clusters: if cluster_fraud_rate > 10% → all members HIGH risk
  4. Write to Redis: HSET graph:risk:{user_id} score 0.85

At scoring time: graph_risk = HGET graph:risk:{user_id} → <1ms lookup
  Cost: ~4 hours nightly compute (Spark cluster)
  Value: catches fraud rings that individual-transaction models miss

Fail-Open vs Fail-Close: Threshold-Based Decision

Revenue impact analysis (5 min downtime):
  Average fraud service downtime: 5 min/month
  During 5 min downtime with threshold-based failover:
    ~1,500 transactions processed
    ~750 low-risk auto-allowed → 0 expected fraud
    ~250 high-risk auto-declined → $12,500 lost revenue
    ~500 medium-risk auto-allowed → ~$250 fraud loss
    
  Compare: fail-close (block all) → $75,000 lost revenue
  Compare: fail-open (allow all) → $3,750 fraud loss
  Threshold-based: best of both worlds

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...