Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
โฑ๏ธ 27 min read

System Design Cheatsheet โ€” Patterns, When to Use What, How to Scale


1. Decision Framework

Use these 5 questions for any system design problem:

Step 1: What is the write path? (sync or async?)
Step 2: What is the read pattern? (point, range, aggregation, search?)
Step 3: Where does hot data live? (memory, disk, CDN?)
Step 4: What breaks under load? (hot partition, write contention, fan-out storm?)
Step 5: What is the consistency requirement? (strong, eventual, causal?)

2. 60-Minute Interview Structure

[0-5]   Requirements (3 FR, 3 NFR, scale numbers)
[5-8]   Core entities + 2-3 API endpoints
[8-12]  Naive diagram (4 boxes) - explain why it breaks
[12-30] Evolve into full architecture (add components one by one with justification)
[30-45] Deep dives (2-3: the hardest parts - scaling, consistency, real-time)
[45-55] Trade-offs summary (state them even if not asked)
[55-60] What you would add with more time (monitoring, failover, ML ranking)

Golden rule: Never spend more than 5 minutes on one component without checking in with the interviewer. Say: โ€œI can go deeper on the real-time tracking or move on to the allocation logic โ€” which is more interesting to you?โ€


3. Data Store Selection โ€” When to Use What

Need Use Why NOT this
Transactions + joins + ACID Postgres / MySQL Row-level locking, MVCC, mature tooling DynamoDB (no joins), Redis (no durability)
High write throughput + no joins Cassandra / DynamoDB LSM-tree, horizontal scale, eventual consistency Postgres (write contention at scale)
Real-time geo proximity Redis Geo In-memory, GEOSEARCH in sub-ms PostGIS (disk I/O too slow for 100K+ writes/sec)
Full-text + multi-attribute search Elasticsearch Inverted index + geo + filters + ranking Postgres LIKE (full scan), Redis (no text search)
Time-series data (metrics, logs) ClickHouse / TimescaleDB Columnar compression, fast range aggregation Postgres (too slow for billions of rows)
Session / cache / counters Redis In-memory, sub-ms, TTL, atomic operations Postgres (too slow for 100K reads/sec)
Large objects (images, files) S3 / GCS Cheap, unlimited, CDN-friendly Postgres BLOB (expensive, hard to scale)
Graph relationships Neo4j / DGraph Traversal in O(depth), not O(joins) Postgres (recursive CTEs too slow at depth > 3)

4. Caching Strategies

When to cache

Cache if Do not cache if
Read-heavy (read:write > 10:1) Write-heavy (cache invalidation hell)
Computation is expensive Data changes every second
Data is tolerant of staleness (10s-5min) Must always be fresh (account balance)
Same data requested by many users Each request is unique (personalized)

Cache patterns

Pattern How it works Use when
Cache-Aside App checks cache - miss - read DB - populate cache Most common. Read-heavy. Can tolerate brief staleness
Write-Through App writes DB + cache simultaneously Need immediate read-after-write consistency
Write-Behind App writes cache only, async flush to DB Ultra-fast writes, can tolerate data loss window
Read-Through Cache itself fetches from DB on miss Simplifies app logic, cache owns data loading

Cache invalidation

Strategy How Trade-off
TTL (Time-To-Live) Key expires after N seconds Simple. Stale for up to TTL duration
Event-based invalidation On DB write, delete cache key Fresh. But requires pub/sub or CDC
Version/ETag Cache stores version, client checks Client-side freshness check

Stampede protection

When TTL expires on a hot key, 1000 requests simultaneously hit DB:


5. Message Queues โ€” When to Use Kafka

USE Kafka when

Scenario Why Kafka
Producer and consumer run at different speeds Kafka buffers. Producer does not wait for consumer
You need replay on failure Consumer crashes? Restart from last committed offset
Multiple consumers need the same event Consumer groups โ€” each group gets every message
Write path is bursty (spikes) Kafka absorbs 100K+/sec burst, consumers process at steady rate
You need an audit trail Kafka retains messages for days/weeks โ€” replay for debugging
Event-driven architecture Services communicate through events, not direct calls
Order matters within an entity Partition by entityId โ€” ordering guaranteed within partition

DO NOT use Kafka when

Scenario Use instead
Request-response (need answer now) Direct HTTP/gRPC call
Simple job queue (one consumer, no replay) SQS / RabbitMQ (simpler)
Real-time push to browser WebSocket + Redis Pub/Sub (lower latency)
Less than 100 messages/sec, no durability need In-memory queue or direct processing

Kafka partitioning strategy

Partition by When Example
userId Events must be ordered per user Chat messages, user activity
entityId (orderId, truckId) State machine transitions must be ordered Shipment lifecycle, order status
region/geo Locality of processing Location pings by city
Random (round-robin) Max throughput, order does not matter Logs, metrics, analytics events

6. Real-Time Patterns

Push vs Pull decision

Use Push (WebSocket/SSE) when Use Pull (Polling) when
User expects instant updates (less than 1s) Acceptable delay of 5-30s
Data changes frequently Data changes rarely
Few active viewers per entity Many viewers but they check infrequently
Examples: live tracking, chat, bidding, stock ticker Examples: email inbox, dashboard refresh, report generation

WebSocket vs SSE vs gRPC Streaming

Protocol Direction Use when
WebSocket Bidirectional Browser/mobile client needs to both send and receive (chat, bidding)
SSE (Server-Sent Events) Server to Client only Simple push, client only listens (notifications, live scores)
gRPC Streaming Both (but typically service-to-service) Backend microservices, binary protocol, typed contracts
Long Polling Simulated push Legacy systems, cannot use WebSocket, need broad compatibility

Real-time architecture patterns

Pattern A: Direct Push (low scale)
  Event happens - Service - WebSocket - Client

Pattern B: Pub/Sub Fan-out (medium scale)
  Event happens - Service - Redis Pub/Sub - WebSocket instances - Clients

Pattern C: Kafka + Pub/Sub (high scale, durability needed)
  Event happens - Kafka (durable) - Consumer - Redis Pub/Sub - WebSocket - Clients

7. Database Internals

B-Tree vs LSM-Tree

ย  B-Tree (Postgres, MySQL) LSM-Tree (Cassandra, RocksDB, LevelDB)
Write O(log N) โ€” random I/O to update pages O(1) amortized โ€” sequential append to memtable + WAL
Read O(log N) โ€” traverse tree directly O(log N) but may check multiple SSTables (amplification)
Best for Read-heavy, transactions, range scans Write-heavy, append-only workloads (logs, metrics, events)
Space Moderate (pages can have fragmentation) Compact after compaction, but write amplification during

One-liner: โ€œPostgres uses B-tree which is optimized for reads. Cassandra uses LSM-tree which is optimized for writes because it converts random writes to sequential appends.โ€

WAL (Write-Ahead Log)

Every database uses this. The rule: write to log BEFORE modifying data.

Flow: Client write - Append to WAL (sequential I/O, fast) - Update in-memory - Ack client
      On crash: replay WAL to recover state

Why sequential: disk sequential write = 200MB/s. Random write = 2MB/s. 100x difference.

MVCC (Multi-Version Concurrency Control)

How Postgres handles concurrent reads + writes without locking:
  - Each row has a version (xmin, xmax transaction IDs)
  - Readers see a SNAPSHOT โ€” consistent view at their transaction start time
  - Writers create NEW versions of rows, do not modify existing
  - Old versions cleaned up by VACUUM

Result: Readers never block writers. Writers never block readers.
Only conflict: two writers updating the same row - one waits for row-level lock.

Indexing strategies

Index type Use when Example
B-tree (default) Range queries, equality, sorting WHERE created_at > X, ORDER BY price
Hash Equality only (faster than B-tree for =) WHERE id = X (only in some DBs)
GIN (Generalized Inverted) Full-text search, array contains, JSONB WHERE tags contains java
GiST (Generalized Search Tree) Geo queries, range types, nearest-neighbor PostGIS ST_DWithin, IP range lookups
Composite (multi-column) Queries filter on multiple columns together (user_id, created_at) for users recent orders
Partial Only index a subset of rows WHERE status = ACTIVE โ€” skip inactive rows (smaller index)

Interview tip: When you design a schema, always state your indexes explicitly and explain which query they serve.


8. Sharding and Partitioning

When to shard

Sharding strategies

Strategy How Best for Watch out for
Hash-based shard = hash(key) % N Uniform distribution, point lookups Range queries become scatter-gather. Resharding is painful
Range-based shard = based on key range (A-M, N-Z) Range scans, time-series Hot partitions if distribution is skewed
Geo-based shard = region (US-East, US-West, EU) Location-aware apps, compliance Cross-region queries are expensive
Directory-based Lookup table maps key to shard Flexible, can rebalance Directory is single point of failure

Hot partition problem

Problem Solution
Celebrity user (10M followers) writes to one shard Split writes across sub-partitions, merge on read
Viral content (one videoId gets all traffic) Add random suffix: videoId_0, videoId_1โ€ฆ spread across shards
Time-based hot partition (all writes go to today) Add random prefix to partition key

9. Consistency Patterns

Level Meaning Use when Technology
Strong (linearizable) Read always sees latest write Payments, booking, inventory Postgres, Spanner, single-leader
Eventual Read may see stale data, converges over time Search index, timeline, counters Cassandra, DynamoDB, cache
Causal If A caused B, everyone sees A before B Chat (messages in order), comments Kafka (per-partition ordering)
Read-your-writes Writer sees their own write immediately User profile edit - immediate reflection Write to primary, read from primary for same user

Achieving strong consistency for bookings and payments

Option 1: DB row lock (simplest, works at less than 1K TPS)
  UPDATE seats SET booked=true WHERE id=? AND booked=false
  Check rowsAffected โ€” if 0, already taken

Option 2: Optimistic locking (medium scale)
  UPDATE seats SET booked=true, version=version+1 WHERE id=? AND version=?
  Retry if version mismatch

Option 3: Distributed lock (cross-service)
  Redis: SET lock:resourceId requestId NX EX 30
  Process - Release lock
  Only when multiple services need to coordinate on same resource

10. Failure Handling Patterns

Circuit Breaker

States: CLOSED (normal) - OPEN (failing, reject all) - HALF-OPEN (test one request)

CLOSED: All requests pass through. Track failure count.
  If failures > threshold in window - switch to OPEN

OPEN: All requests immediately fail (do not hit downstream). Return fallback.
  After timeout - switch to HALF-OPEN

HALF-OPEN: Allow ONE request through.
  If success - back to CLOSED
  If fail - back to OPEN

When to use: calling external APIs (payment gateway, Maps API, carrier systems)
Implementation: Resilience4j (Java), Polly (.NET), custom with Redis counters

Retry with Exponential Backoff

Attempt 1: immediate
Attempt 2: wait 1s
Attempt 3: wait 2s
Attempt 4: wait 4s
Attempt 5: wait 8s (give up after this)

Formula: delay = base_delay * 2^(attempt-1) + random_jitter

ALWAYS add jitter (random 0-500ms). Without jitter:
  1000 clients all fail at same time - all retry at exactly same time - thundering herd

When to use: transient failures (network timeout, 503, connection refused)
When NOT to: 400 errors (bad request โ€” retrying will not help), 401/403 (auth โ€” retrying will not help)

Idempotency

Client generates unique request_id before sending.
Server:
  1. Check Redis: GET idempotency:request_id
  2. If exists - return cached response (duplicate detected)
  3. If not - process request - SET idempotency:request_id response EX 3600
  4. Return response

Key: request_id must be generated CLIENT-SIDE (not server-generated)
TTL: 1 hour is typical (covers retries within a session)
Scope: per-user + per-operation (not global)

Saga Pattern (Distributed Transactions)

Problem: Book ride requires: (1) charge payment (2) assign driver (3) create ride record
         These span 3 services. Cannot use a single DB transaction.

Saga: Execute steps in order. If any step fails, run compensating actions backwards.

  Step 1: Reserve payment - Success
  Step 2: Assign driver - Success
  Step 3: Create ride - FAILS
  Compensate: Release driver - Refund payment

Types:
  - Choreography: each service publishes events, next service reacts (simple, hard to debug)
  - Orchestration: central orchestrator calls each service in order (easier to debug, single point of failure)

When to use: multi-service writes that must be all-or-nothing
When NOT to: single-database operations (just use a DB transaction)

Dead Letter Queue (DLQ)

Normal flow: Kafka - Consumer - Process - Commit offset
Failure flow: Kafka - Consumer - Process FAILS 3x - Send to DLQ topic - Alert ops

DLQ gives you:
  - Messages are not lost (they are in the DLQ topic)
  - Main queue is not blocked (consumer moves on)
  - Ops can inspect, fix, and replay from DLQ later

When to use: any async consumer that cannot afford to lose messages

11. Communication Protocols

HTTP/1.1 vs HTTP/2 vs HTTP/3

Feature HTTP/1.1 HTTP/2 HTTP/3
Connections 1 request per connection (or keep-alive with head-of-line blocking) Multiplexed: many requests on 1 TCP connection Multiplexed over QUIC (UDP), no head-of-line blocking
Header compression None HPACK (significant savings for repeated headers) QPACK
Server push No Yes (server can send resources before client asks) Yes
Best for Legacy, simple APIs, microservices, real-time apps Mobile, unreliable networks

gRPC vs REST

Aspect REST (JSON/HTTP) gRPC (Protobuf/HTTP2)
Payload size Large (JSON text) Small (binary protobuf, 3-10x smaller)
Speed Slower (JSON parse) Faster (binary deserialize)
Streaming Not native (need WebSocket) Native bidirectional streaming
Browser support Universal Limited (needs grpc-web proxy)
Schema Optional (OpenAPI) Required (proto files) โ€” strong typing
Best for Public APIs, browser clients Internal microservices, high-throughput service-to-service

MQTT (for IoT/GPS devices)

Why MQTT for truck GPS:
  - Designed for low-bandwidth, unreliable networks (trucks in rural areas)
  - Tiny packet overhead (2 bytes minimum header vs HTTP 200 bytes)
  - Persistent session: if device disconnects and reconnects, queued messages are delivered
  - QoS levels: 0 (at most once), 1 (at least once), 2 (exactly once)

Typical setup: GPS device - MQTT broker (Mosquitto/EMQX) - Bridge to Kafka - Consumer

12. Scaling By Use Case

Real-Time Tracking (Uber, Freight, Delivery)

Device - API Gateway - Kafka - Location Consumer - Redis Geo
                                                 - Redis Pub/Sub - WebSocket - Viewer

Key decisions:
- Kafka for durability and replay (device pings are fire-and-forget)
- Redis Geo for proximity queries (GEOSEARCH)
- Pub/Sub for push to watchers (not polling)
- Partition Kafka by deviceId (ordering per device)
- Store only latest position in Redis (bounded memory)
- Historical data - Cassandra/TimescaleDB (time-series, cold)

Analytics Dashboard (Metrics, Monitoring)

Events - Kafka - Stream Processor (Flink) - Pre-aggregated rollups - Redis (hot) + Cassandra (cold)
                                                                          |
                                                                   Dashboard API reads from here

Key decisions:
- Never query raw events at read time (too slow)
- Pre-aggregate in time windows: 1-min, 1-hour, 1-day buckets
- Redis for last-hour data (TTL 2 hours)
- Cassandra/ClickHouse for historical (last 30 days)
- Kafka windowing or Flink for stream aggregation

Notification System (Multi-channel, Priority)

Event - Kafka (topic: notifications) - Priority Router
  - High priority: dedicated fast consumers - immediate push (FCM/APN)
  - Low priority: batch consumers - aggregate - email digest

Key decisions:
- Kafka for durability (never lose a notification)
- Separate topics or priority headers for routing
- User preference DB (cached in Redis) โ€” check before sending
- Idempotency key per notification (prevent duplicates on retry)
- Dead-letter queue for failed deliveries

Marketplace and Matching (Uber, Freight, Dating)

Supply (drivers/trucks) - Location Store (Redis Geo)
Demand (riders/loads) - Matching Service - Query Redis Geo - Score - Assign

Key decisions:
- Redis Geo for find nearest N in O(log M + N)
- Scoring formula: distance x availability x rating x demand
- Atomic assignment: UPDATE ... WHERE status=AVAILABLE (prevent double-booking)
- If high contention: retry with next candidate from pre-sorted list
- Fan-out new demand to eligible supply via push notification

Bidding and Auction (Real-Time, Competitive)

Bid - Kafka (durable) - Bid Processor - Redis Sorted Set (leaderboard)
                                       - Redis Pub/Sub - WebSocket - All watchers

Key decisions:
- Kafka between bid submission and processing (durability + replay)
- Redis Sorted Set for O(log N) ranking, O(1) best bid
- Pub/Sub for live updates to all watchers
- Auction close: single consumer per partition + DB transaction (no double-award)
- Anti-sniping: extend auction if last-minute bid arrives

Chat and Messaging

Sender - API - Kafka (ordered per conversationId) - Message Service - DB (Cassandra)
                                                                    - Pub/Sub - WebSocket - Receiver

Key decisions:
- Kafka partitioned by conversationId (ordering within conversation)
- Cassandra for message storage (write-heavy, partition by conversationId + time)
- WebSocket for online delivery, push notification for offline
- Delivery receipts: sent - delivered - read (state machine per message)
- Idempotency: message_id generated client-side, server deduplicates

Payment and Financial

Request - API - Idempotency check (Redis) - Payment Service - DB (Postgres with TX)
                                                            - Kafka (event for downstream)
                                                            - External payment gateway

Key decisions:
- Idempotency key: SET request_id response NX EX 3600 (prevent double-charge)
- Postgres for ledger (double-entry bookkeeping, ACID)
- Saga pattern for multi-step: charge - transfer - settle (compensating actions on failure)
- Kafka for event trail (audit, reconciliation)
- Never call external gateway without idempotency key

Rate Limiter

Request - API Gateway - Rate Limit Check (Redis) - Allow/Reject - Backend Service

Redis key: rate:userId:minute_bucket
Algorithm options:
  - Token Bucket: INCR + EXPIRE. If count > limit - reject (simplest)
  - Sliding Window Counter: HINCRBY on minute sub-buckets, sum last N
  - Sliding Window Log: ZADD timestamp, ZREMRANGEBYSCORE to trim, ZCARD for count
Algorithm Pros Cons Best for
Fixed Window Simple, O(1) Burst at window edges Low-precision limiting
Sliding Window Counter Smooth, no edge bursts Slightly more complex API rate limiting
Token Bucket Allows controlled bursts Needs refill logic Network traffic shaping
Leaky Bucket Constant output rate No bursts allowed Queue-based processing

URL Shortener

Write: long_url - hash/counter - short_code - Store in DB + Cache
Read: short_code - Cache hit? Return. Miss? DB lookup - Cache - 301 Redirect

Key decisions:
- Base62 encoding of auto-increment ID (no collisions, predictable)
- OR MD5/SHA hash truncated to 7 chars (collision possible, need check)
- Read-heavy (1000:1 read:write) - aggressive caching (Redis, CDN)
- Analytics: log each redirect to Kafka - analytics pipeline

Feed Generation (Twitter / Instagram)

Fan-out on Write (for regular users):
  User posts - Write to followers pre-built feed in Redis/Cassandra
  Read: just fetch pre-built feed. O(1) per page.

Fan-out on Read (for celebrities):
  User opens feed - Merge posts from all followed users in real-time
  Read: expensive (query N users). But write is O(1).

Hybrid (production):
  - Regular users (less than 10K followers): fan-out on write
  - Celebrities (more than 10K followers): fan-out on read, merge at read time
  - Feed cache per user in Redis, TTL 5 min

Search Autocomplete

User types - Debounce (300ms) - API - Trie/ES prefix query - Top 10 results

Options:
  - Trie in Redis: ZRANGEBYLEX for prefix matching (fast, limited)
  - Elasticsearch: prefix + completion suggester (flexible, heavier)
  - Pre-computed: Top 1000 queries cached in CDN/Redis (cheapest for hot queries)

Key decisions:
- Debounce on client (do not send every keystroke)
- Cache popular prefixes aggressively (80% of searches hit top 1000)
- Personalization: blend global popular + user recent searches

13. Capacity Estimation

Quick formulas

Metric Quick formula
QPS from DAU DAU x actions_per_user / 86400
Peak QPS Average x 3 (rule of thumb)
Storage/year daily_records x record_size x 365
Bandwidth QPS x avg_response_size
Servers needed Peak_QPS / single_server_capacity (typically 1K-10K req/sec)

Typical capacity per instance

System Typical capacity per instance
Postgres 5K-20K reads/sec, 1K-5K writes/sec
Redis 100K-200K ops/sec
Kafka broker 100K-500K messages/sec
WebSocket server 50K-100K concurrent connections
Elasticsearch 5K-20K searches/sec
Single API server 1K-10K req/sec (depends on complexity)

Worked Example 1: Uber-scale Location Tracking

Given: 2M active drivers, GPS ping every 4 seconds
Write QPS: 2,000,000 / 4 = 500,000 writes/sec
Payload: ~100 bytes per ping
Bandwidth: 500K x 100B = 50 MB/sec ingest
Daily storage: 500K x 100B x 86,400 = 4.3 TB/day (if storing all)
Redis memory: 2M entries x 100 bytes = 200 MB (only latest position โ€” fits in one instance)

Worked Example 2: Chat System (WhatsApp-scale)

Given: 100M DAU, average 40 messages/day per user
Write QPS: 100M x 40 / 86,400 = ~46K messages/sec
Message size: ~1KB (text + metadata)
Daily storage: 100M x 40 x 1KB = 4 TB/day
Peak QPS: 46K x 3 = ~140K messages/sec
Cassandra cluster: 140K writes/sec / 20K per node = 7 nodes minimum

Worked Example 3: URL Shortener

Given: 100M new URLs/month, 10:1 read:write ratio
Write QPS: 100M / (30 x 86,400) = ~40 writes/sec (trivial)
Read QPS: 40 x 10 = 400 reads/sec (still trivial for one DB)
After 5 years: 100M x 12 x 5 = 6 billion URLs
Storage: 6B x 500 bytes = 3 TB (fits on a single machine with SSDs)
Cache: 80/20 rule โ€” cache top 20% = 1.2B x 500B = 600 GB Redis cluster

Quick reference numbers

Thing Number to remember
1 day in seconds 86,400 (~100K)
1 month in seconds 2.6 million
1 year in seconds 31.5 million
1 KB 1,000 bytes
1 MB 1,000,000 bytes
1 GB 1 billion bytes
1 TB 1 trillion bytes
Single server QPS (API) 1K-10K req/sec
Redis ops/sec 100K-200K
Kafka throughput 1 million msg/sec (per cluster)
Postgres writes/sec 5K-20K (single primary)
Network bandwidth (1 Gbps) 125 MB/sec
SSD sequential read 500 MB/sec
SSD random read 200K IOPS
HDD sequential read 100 MB/sec
Memory access ~100 ns
SSD random access ~100 us (1000x memory)
Network round trip (same DC) ~0.5 ms
Network round trip (cross-region) ~50-150 ms

14. Trade-off One-Liners

Memorize these for when the interviewer asks โ€œwhy this over that?โ€

Trade-off One-line answer
SQL vs NoSQL โ€œSQL when I need transactions and joins. NoSQL when I need horizontal write scale and flexible schema.โ€
Redis vs Memcached โ€œRedis โ€” it has data structures (sorted sets, geo, pub/sub). Memcached is only key-value.โ€
Kafka vs RabbitMQ โ€œKafka for high-throughput event streaming with replay. RabbitMQ for simple task queues with routing.โ€
Postgres vs DynamoDB โ€œPostgres for complex queries and ACID. DynamoDB for single-key lookups at infinite scale.โ€
WebSocket vs SSE โ€œWebSocket for bidirectional (chat, bidding). SSE for server-push-only (notifications, tracking).โ€
Consistent hashing vs Mod N โ€œConsistent hashing โ€” adding/removing a node only moves K/N keys instead of reshuffling everything.โ€
Optimistic vs Pessimistic lock โ€œOptimistic when conflicts are rare (read-heavy). Pessimistic when conflicts are frequent (booking).โ€
Push vs Pull โ€œPush when latency matters and data changes often. Pull when data changes rarely or clients are many.โ€
Monolith vs Microservices โ€œMonolith to start (faster development). Microservices when teams scale and services have different scaling needs.โ€
Sync vs Async โ€œSync when user needs immediate confirmation (payment). Async when result can be delivered later (email, analytics).โ€
Strong vs Eventual consistency โ€œStrong for money and bookings. Eventual for feeds, search, and counters.โ€
Cache-aside vs Write-through โ€œCache-aside for read-heavy with tolerable staleness. Write-through when reads must always be fresh.โ€
Range vs Hash sharding โ€œRange for time-series and scans. Hash for uniform distribution and point lookups.โ€
CDN vs Origin โ€œCDN for static assets and cacheable responses (images, JS, API responses with TTL). Origin for dynamic/personalized.โ€

15. Interview Anti-Patterns

Do not do this Do this instead
โ€œLets use Kafkaโ€ without saying why โ€œI will use Kafka here because we need durability and the consumer might lag โ€” we cannot afford to lose these eventsโ€
Jump to microservices immediately Start with 2-3 services, split only when you explain WHY a new service is needed
Say โ€œwe can scale horizontallyโ€ without details Say HOW: โ€œpartition by userId, 10 shards, consistent hashingโ€
Forget the hot path Identify WHICH path has the highest load and design for that first
Over-design the cold path 10K bookings/day = 0.1 TPS โ€” a single Postgres handles this trivially. Do not shard it.
No numbers Always give rough estimates: โ€œat 50K trucks x 6 pings/min = 300K writes/sec, we need Kafka + Redisโ€
โ€œI do not knowโ€ โ€œMy instinct is X because Y. Does that make sense?โ€ โ€” ALWAYS guess with reasoning

16. Interview Phrases

Opening (after hearing the question)

โ€œBefore I jump in, let me clarify a few things to make sure I scope this rightโ€ฆโ€

โ€œLet me start with the functional requirements โ€” I want to make sure we align on what the system must do versus nice-to-haves.โ€

Scoping

โ€œThe shipper side is a straightforward CRUD form. The interesting engineering challenge is on the carrier side โ€” matching, real-time updates, and concurrency. Iโ€™ll focus my design there and keep the shipper side simple. Does that sound right?โ€

Drawing the first diagram

โ€œLet me start with the simplest possible architecture and then evolve it as requirements demand more.โ€

Introducing a new component

โ€œAt this point we have a problem: 300K writes per second hitting a single store. Thatโ€™s why Iโ€™m introducing Kafka here โ€” it buffers the writes and lets the consumer process at its own pace.โ€

Trade-off discussion

โ€œThereโ€™s a trade-off here. Iโ€™m choosing eventual consistency for the tracking display โ€” a shipper might see location thatโ€™s 10 seconds stale. The alternative is strong consistency which would require synchronous writes and reduce our throughput by 10x. Since this is a display counter and not a financial transaction, eventual consistency is acceptable.โ€

When unsure

โ€œI havenโ€™t worked with this exact scenario before, but my instinct is to use X because of Y. If that turns out to be wrong, we could fall back to Z. Does this direction make sense?โ€

Deep dive justification

โ€œLet me go deeper on the allocation logic because thatโ€™s where the concurrency risk is highest. The question is: what happens when two bookings want the same truck simultaneously?โ€

Wrapping up

โ€œTo summarize โ€” the system has three main data paths: the booking path which is sync and strongly consistent, the location ingestion path which is async through Kafka, and the tracking path which pushes via WebSocket. The hardest problem is the 300K writes/sec on location โ€” Kafka plus Redis handles that. Given more time, Iโ€™d add ML-based ETA prediction and multi-stop route optimization.โ€

When interviewer asks โ€œwhat would you improve?โ€

โ€œThree things: first, Iโ€™d add a circuit breaker on the Maps API call so ETA computation degrades gracefully. Second, Iโ€™d add a reconciliation job that compares Redis state with Postgres to catch any drift. Third, Iโ€™d implement CDC from Postgres to keep the search index fresh without dual-writing.โ€


17. Common Follow-up Questions

They ask They are testing Your answer pattern
โ€œWhat if this service goes down?โ€ Fault tolerance โ€œThe data is in Kafka, so on recovery we replay from last offset. No loss.โ€
โ€œWhat if two users do this simultaneously?โ€ Concurrency โ€œI use WHERE status=AVAILABLE as an atomic CAS. Only one succeeds.โ€
โ€œHow would you scale this to 10x?โ€ Horizontal scaling knowledge โ€œPartition by X, add more consumers/shards. The bottleneck moves to Y, which Iโ€™d solve with Z.โ€
โ€œWhy not just use X?โ€ Trade-off reasoning โ€œX works but has limitation A. My choice handles B which is critical for our NFR.โ€
โ€œWhat is the consistency model here?โ€ CAP understanding โ€œFor reads: eventual (acceptable 5s staleness). For writes: strong (DB transaction).โ€
โ€œHow do you handle data that does not fit in memory?โ€ Storage tiering โ€œHot data in Redis (last hour). Warm in Postgres (last month). Cold in S3/Cassandra (archive).โ€
โ€œWhat happens on network partition?โ€ CAP theorem โ€œWe favor availability โ€” the system continues to accept writes. We reconcile inconsistencies later via async job.โ€
โ€œHow would you test this?โ€ Engineering maturity โ€œLoad test the hot path (location ingestion) with 500K synthetic pings/sec. Chaos test: kill a Redis node during peak and verify recovery from Kafka.โ€

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access