Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll

Quick Revision β€” All 35 System Designs

Scan this before your interview. Each card gives you the core idea, the tech that matters, and the one insight that separates β€œgood” from β€œgreat” answers.


Real-Time & Communication

Chat System

1-on-1 and group messaging with delivery guarantees and presence.

Concepts: WebSocket, Message ordering, Presence, Fan-out on write, Offline sync Stack: WebSocket Gateway (connection mgmt), Kafka (message routing), Cassandra (chat history), Redis (presence + unread counts) Data Model: chat_messages: (channel_id, message_id timeuuid) β†’ {sender, body, timestamp} β€” Cassandra partition per channel, time-sorted within partition for O(1) append and range reads. Money Insight: Fan-out on write for small groups (<500 members), fan-out on read for large channels. The hybrid model is what WhatsApp/Discord actually use β€” pure fan-out-on-write explodes storage for 100K-member groups.


Notification System

Multi-channel delivery engine β€” push, SMS, email β€” with priority, dedup, and rate limiting.

Concepts: Priority queue, Template rendering, Rate limiting, Delivery tracking, Channel routing Stack: Kafka (ingest), Redis (dedup + rate limit), Postgres (templates + preferences), APNs/FCM/Twilio (delivery) Data Model: notifications: (user_id, notification_id) β†’ {channel, template_id, payload, status, retry_count} β€” status FSM tracks sent/delivered/failed per channel. Money Insight: The hard problem isn’t sending β€” it’s NOT sending. Dedup (idempotency key in Redis with TTL), per-user rate limits (sliding window), and quiet hours (timezone-aware suppression) prevent the system from becoming spam infrastructure.


Uber / Ride Sharing

Real-time driver matching and ride lifecycle management.

Concepts: Geospatial indexing, WebSocket, Fan-out, Pub/Sub, Surge pricing Stack: Redis Geo (location), Kafka (events), Postgres (rides), WebSocket (tracking) Data Model: GEOADD drivers:{city} lng lat driverId β€” 2M location pings/sec indexed by geohash for O(log N) radius search. Money Insight: You can’t query β€œnearest drivers” from Postgres at 500K QPS. Redis Geo with geohash sharding solves the hot write + proximity read problem simultaneously.


Leaderboard

Real-time ranking of millions of players with score updates and rank queries.

Concepts: Sorted sets, Sharding, Pagination, Time-windowed partitioning Stack: Redis Sorted Sets (ranking), Kafka (score events), Postgres (historical snapshots) Data Model: ZADD leaderboard:{game}:{window} score playerId β€” ZREVRANK gives rank in O(log N), ZREVRANGE gives top-K in O(K+log N). Money Insight: A single Redis sorted set handles ~25M members. Beyond that, shard by score range (not user hash!) so top-K queries hit only 1 shard. Time-windowed keys (daily/weekly) avoid expensive ZREMRANGEBYSCORE cleanups.


Freight Logistics

Truck allocation, route optimization, and real-time shipment tracking with live ETA.

Concepts: Redis Geo + geohash, Atomic CAS allocation, Kafka ingestion buffer, Traffic-aware ETA recomputation, Movement debounce Stack: Redis Geo (live truck positions), Kafka (GPS pings partitioned by truckId), Postgres (shipments + assignments), Cassandra/TimescaleDB (historical paths), Maps API (route + traffic ETA) Data Model: trucks:available:{region} -> Geo Set, member truckId, score geohash β€” keyed by region so 300K pings/sec spread across shards. truck:meta:{truckId} carries a 60s TTL, so a truck that stops pinging expires itself offline rather than needing a reaper. Money Insight: Resist the distributed lock. At ~10K bookings/day (0.1 TPS) a plain compare-and-swap wins: UPDATE trucks SET status='ASSIGNED' WHERE truck_id=? AND status='AVAILABLE', check rowsAffected, and on 0 fall through a pre-sorted top-10 candidate list. The naive check-then-update is what actually breaks β€” two allocator instances both see the truck as free and both assign it.


Content & Feed

Twitter Feed

Fan-out timeline generation with mixed media, likes, retweets at read-heavy scale.

Concepts: Fan-out on write vs read, Timeline cache, Celebrity problem, Ranking Stack: Redis (timeline cache), Kafka (fan-out workers), Cassandra (tweets), ML ranking service Data Model: timeline:{userId} β†’ [tweetId1, tweetId2, ...] in Redis list β€” pre-computed for 99% of users. Celebrity tweets merged at read time. Money Insight: Pure fan-out on write breaks for users with 50M followers (Lady Gaga problem). Hybrid: fan-out on write for users <10K followers, merge at read for celebrities. This is Twitter’s actual architecture.


Instagram

Photo/video sharing with feed, stories, explore, and social graph.

Concepts: CDN, Object storage, Feed ranking, Social graph, Content moderation Stack: S3 (media), CloudFront/CDN (delivery), Postgres (users/relationships), Cassandra (feed), ML (ranking + moderation) Data Model: posts: (user_id, post_id) β†’ {media_url, caption, location} + feed: (user_id, timestamp) β†’ [post_ids] β€” feed materialized via async fan-out workers. Money Insight: The real bottleneck is the Explore page, not the feed. Feed is pull-from-cache, but Explore needs collaborative filtering over billions of interactions in near-real-time. Pre-compute candidate pools per interest cluster, then rank at request time.


Netflix

Video streaming platform with adaptive bitrate, content delivery, and personalization.

Concepts: Adaptive bitrate (ABR), CDN, Transcoding pipeline, Recommendation engine Stack: S3 (master copies), Open Connect CDN (edge delivery), Kafka (viewing events), Spark (recommendations), FFmpeg (transcoding) Data Model: content: content_id β†’ {title, metadata} + encodings: (content_id, profile) β†’ {cdn_url, bitrate, resolution} β€” each title stored in 100+ encoding profiles. Money Insight: Netflix doesn’t stream from S3. They push content to 17,000+ Open Connect Appliances (OCAs) inside ISPs. The real design problem is the content placement algorithm β€” predicting what to pre-position where based on regional viewing patterns.


News Aggregator

Personalized news feed with source crawling, dedup, and ranking.

Concepts: Web crawling, NLP dedup, Personalization, Pub/Sub, Content ranking Stack: Kafka (article ingest), Elasticsearch (search + dedup), Redis (user feed cache), ML (personalization), Postgres (sources + metadata) Data Model: articles: article_id β†’ {source, title, body_hash, entities[], topic_vector} β€” body_hash (SimHash) for near-duplicate detection across sources. Money Insight: The same news event gets reported by 500 sources. SimHash (locality-sensitive hashing) clusters near-duplicates in O(1) per article, then you pick the best source per cluster by authority score. Without this, your feed is 90% duplicate noise.


QA Forum

Stack Overflow-style Q&A with voting, reputation, search, and anti-gaming.

Concepts: Full-text search, Reputation system, Vote aggregation, Anti-abuse Stack: Elasticsearch (search), Postgres (questions/answers/votes), Redis (hot question cache + reputation cache), Kafka (vote events) Data Model: questions: question_id β†’ {title, body, tags[], vote_count, author_id} + votes: (entity_id, user_id) β†’ {direction, timestamp} β€” composite key prevents double-voting at DB level. Money Insight: Vote count updates are the hidden hot path. Don’t UPDATE a counter per vote β€” use Kafka to batch vote events, aggregate per 1-second window, then do a single atomic increment. This turns 10K individual writes into 1 batched write per question.


Storage & Infrastructure

Dropbox

Cloud file sync with conflict resolution, chunked upload, and cross-device consistency.

Concepts: Chunked upload, Content-addressable storage, Sync protocol, Conflict resolution Stack: S3 (chunks), Postgres (metadata + file tree), Redis (sync state), Block server (chunking + dedup) Data Model: files: (user_id, path) β†’ {file_id, chunks: [sha256_1, sha256_2, ...], version} β€” content-addressable chunks enable cross-user dedup (50%+ storage savings). Money Insight: Files are split into 4MB chunks hashed by SHA-256. If two users upload the same PDF, it’s stored once. Delta sync means editing one byte in a 1GB file only uploads the changed chunk. This is why Dropbox’s storage cost is a fraction of naive object storage.


Key-Value Store

Distributed KV store with tunable consistency (like DynamoDB/Cassandra).

Concepts: Consistent hashing, Replication, Quorum reads/writes, Merkle trees, Vector clocks Stack: Custom storage engine (LSM-tree or B-tree), Gossip protocol (membership), WAL (durability) Data Model: hash(key) β†’ vnode β†’ {value, vector_clock, tombstone} β€” consistent hash ring maps keys to N replicas. W+R>N gives strong consistency; W+R≀N gives eventual. Money Insight: The quorum formula W+R>N is table stakes. The real insight is handling temporary failures: hinted handoff (write to a neighbor, replay later) prevents data loss during node outages without blocking the write path.


Message Queue

Distributed message broker with ordering, at-least-once delivery, and consumer groups.

Concepts: Partitioning, Consumer groups, Offset tracking, Replication, Backpressure Stack: Append-only log (segments on disk), ZooKeeper/Raft (leader election), OS page cache (zero-copy reads) Data Model: topic:{partition} β†’ [offset_0: msg, offset_1: msg, ...] β€” append-only segment files. Consumers track offsets independently. Retention by time or size. Money Insight: Kafka’s magic isn’t the queue β€” it’s sequential disk I/O + OS page cache + zero-copy sendfile(). A single broker does 600MB/s because it never deserializes messages. The broker is just a dumb commit log; all intelligence lives in producers and consumers.


URL Shortener

Generate short URLs, redirect at scale, track analytics.

Concepts: Base62 encoding, Distributed ID generation, 301 vs 302 redirect, Read-heavy caching Stack: Redis (URL cache), Postgres (URL mappings), Snowflake/TSID (ID generation), CDN (redirect at edge) Data Model: short_urls: base62(id) β†’ {long_url, user_id, created_at, expires_at} β€” 7-char base62 gives 3.5 trillion unique URLs. Money Insight: The redirect is the product β€” every millisecond of latency is lost ad revenue. Put the hot 20% of URLs in a Redis cache at the edge, serve 302 (not 301) so you retain analytics control, and the entire read path is cache-hit β†’ redirect in <5ms.


Pastebin

Text snippet sharing with expiration and syntax highlighting.

Concepts: Object storage, TTL-based expiration, Content-addressable storage, Read-heavy caching Stack: S3 (paste content), Postgres (metadata), Redis (hot paste cache), CDN (static delivery) Data Model: pastes: paste_id β†’ {s3_key, language, expires_at, visibility} β€” content stored in S3, metadata in Postgres. Lazy expiration via TTL index. Money Insight: Don’t store paste content in your database. At 10M pastes/day averaging 5KB each, that’s 50GB/day in your OLTP store. S3 costs $0.023/GB/month vs $0.10+/GB for RDS. Separate metadata (fast queries) from content (cheap storage).


Unique ID Generator

Distributed, sortable, unique ID generation at millions/sec with no coordination.

Concepts: Snowflake IDs, Clock skew, Epoch-based timestamps, Worker registration Stack: In-process ID generation (no network hop), ZooKeeper (worker ID assignment), NTP (clock sync) Data Model: [41-bit timestamp | 10-bit worker_id | 12-bit sequence] β€” 64-bit ID, sortable by time, 4096 IDs/ms per worker, no coordination needed after startup. Money Insight: The breakthrough is ZERO network calls for ID generation. Each worker pre-registers its 10-bit ID, then generates locally using timestamp + sequence. Compared to a centralized ID service (single point of failure, network latency), this gives sub-microsecond generation at any scale.


Web Crawler

Distributed crawler for billions of pages with politeness, dedup, and priority scheduling.

Concepts: URL frontier, Politeness (robots.txt), Content fingerprinting, Priority scheduling, Trap detection Stack: Kafka (URL frontier), Redis (seen-URL bloom filter), S3 (raw HTML storage), DNS resolver cache, Headless Chrome (JS rendering) Data Model: url_frontier: priority_queue per domain β†’ {url, depth, last_crawl, retry_count} + seen_urls: BloomFilter(url_hash) β€” bloom filter gives O(1) dedup with <1% false positive. Money Insight: The bottleneck isn’t bandwidth β€” it’s politeness. You must rate-limit per domain (1 req/sec). A single-threaded crawler is idle 99% of the time. The URL frontier must be partitioned by domain so each worker owns a set of domains and can pipeline requests without violating per-host limits.


Commerce & Payments

BookMyShow

Ticket booking with seat selection, hold-and-pay, and flash sale handling.

Concepts: Distributed locking, Seat hold with TTL, Idempotent payments, Queue-based load leveling Stack: Redis (seat locks with TTL), Postgres (bookings + inventory), Kafka (payment events), Rate limiter (flash sale protection) Data Model: seats: (show_id, seat_id) β†’ {status: available|held|booked, held_by, held_until} β€” Redis SETNX with TTL creates a 10-min hold. Expired holds auto-release. Money Insight: The double-booking problem isn’t solved by β€œjust use transactions.” At 100K concurrent users for a popular show, DB-level locks create a thundering herd. Use Redis SETNX (atomic set-if-not-exists) with a 10-min TTL as a lightweight distributed lock β€” it fails fast for losers and auto-releases on timeout.


Cart System

E-commerce cart with inventory reservation, price consistency, and session merge.

Concepts: Soft reservation, Price snapshot, Cart merge (guestβ†’logged-in), Eventual consistency Stack: Redis (active cart), Postgres (inventory + orders), Kafka (inventory events), DynamoDB (session store) Data Model: cart:{user_id} β†’ {items: [{sku, qty, price_snapshot, reserved_until}]} in Redis β€” TTL-based soft reservation prevents overselling without hard locks on inventory. Money Insight: Never trust the client’s price. Store a price_snapshot at add-to-cart time, re-validate at checkout against the catalog service. The cart is a β€œpromise” not a β€œcontract” β€” prices can change, items can go OOS. Show the diff at checkout, don’t silently charge the wrong amount.


Payment System

Payment processing with idempotency, ledger, reconciliation, and multi-PSP routing.

Concepts: Idempotency, Double-entry ledger, Saga pattern, PSP routing, Reconciliation Stack: Postgres (ledger), Kafka (payment events), Redis (idempotency keys), Temporal (payment orchestration), Stripe/Adyen (PSP) Data Model: ledger_entries: (tx_id, entry_id) β†’ {debit_account, credit_account, amount, currency} β€” every transaction creates exactly 2 entries (double-entry). Sum of all entries = 0 invariant. Money Insight: Idempotency isn’t optional β€” it’s the entire design. Network failures mean you WILL send duplicate charge requests to PSPs. An idempotency key (client-generated UUID) stored in Redis with the PSP response means retries return the cached result, not a double charge. This is Stripe’s actual pattern.


Digital Wallet

Stored-value wallet with top-up, P2P transfers, and balance consistency.

Concepts: Double-entry bookkeeping, Optimistic locking, Saga for transfers, Compliance (KYC/AML) Stack: Postgres (ledger + balances), Kafka (transaction events), Redis (balance cache for reads), Temporal (transfer orchestration) Data Model: balances: (wallet_id, currency) β†’ {available, pending, version} + transactions: tx_id β†’ {from_wallet, to_wallet, amount, status, idempotency_key} β€” version column enables optimistic concurrency control. Money Insight: A P2P transfer is NOT UPDATE balance SET amount = amount - 100 WHERE wallet_id = sender. That’s a lost update waiting to happen. Use optimistic locking: read version, compute new balance, UPDATE ... WHERE version = expected_version. If version changed, retry. This gives serializable consistency without pessimistic locks.


Stock Broker

Trading platform with real-time market data, order matching, and portfolio management.

Concepts: Order book, Price-time priority matching, Event sourcing, Market data streaming Stack: In-memory matching engine (Java/C++), Kafka (order events), TimescaleDB (OHLCV candles), WebSocket (market data streaming), Redis (portfolio cache) Data Model: order_book: (symbol, side) β†’ SortedMap<price, Queue<Order>> β€” price-time priority. Bids sorted descending, asks ascending. Match when best_bid β‰₯ best_ask. Money Insight: The matching engine MUST be single-threaded per symbol. Multi-threaded matching introduces race conditions that cause incorrect fills. One thread per symbol, processing orders sequentially, gives deterministic matching at 100K orders/sec β€” this is how real exchanges (LMAX, Nasdaq) work.


Zomato

Food delivery platform with restaurant discovery, ordering, and real-time delivery tracking.

Concepts: Geospatial search, ETA estimation, Order state machine, Delivery assignment, Menu management Stack: Elasticsearch (restaurant search), Redis Geo (delivery partner location), Postgres (orders + menus), Kafka (order events), WebSocket (live tracking) Data Model: restaurants: {restaurant_id, name, location, cuisine[], rating, delivery_radius} in Elasticsearch + orders: order_id β†’ {status_fsm, items[], restaurant_id, delivery_partner_id, eta} in Postgres. Money Insight: ETA isn’t just distance/speed. It’s prep_time (per restaurant ML model) + pickup_wait (driver location + traffic) + delivery_time (route-based, not crow-flies). Getting ETA wrong by even 5 minutes kills retention. The system needs separate ML models for each phase, retrained hourly on real delivery data.


Freight Bidding Platform

Real-time auction marketplace where carriers compete on freight loads posted by shippers.

Concepts: Redis Sorted Set leaderboard, Kafka write-ahead buffer, Exactly-one-winner semantics, WebSocket fan-out per auction, Anti-sniping Stack: Redis Sorted Sets (ranked bids per load), Kafka (bid events partitioned by loadId), Postgres (loads, auctions, bookings), WebSocket + Redis Pub/Sub (live bid fan-out) Data Model: bids:{loadId} -> Sorted Set, member {carrierId}:{bidId}, score = composite (0.7*price + 0.2*distance_penalty + 0.1*(1/rating)) β€” best bid is an O(1) ZRANGE 0 0, a carrier’s own rank an O(log N) ZRANK. bookings.load_id is UNIQUE, which is where one-winner is actually guaranteed. Money Insight: No distributed lock is needed to stop a double-award. UPDATE loads SET status='AWARDED' WHERE load_id=? AND status='OPEN' is sufficient on its own, because the second transaction’s WHERE clause matches no row; partitioning Kafka by loadId then removes the contention entirely by giving each auction a single consumer. Worth naming the mechanism precisely in an interview: this is an open composite-score auction, not second-price.


Scheduling & Monitoring

Job Scheduler

Distributed cron with exactly-once execution, retries, and DAG dependencies.

Concepts: Distributed locking, Exactly-once semantics, DAG execution, Dead-letter queue Stack: Postgres (job definitions + state), Redis (distributed locks), Kafka (job events), Worker pool (execution) Data Model: jobs: job_id β†’ {cron_expr, handler, status, locked_by, locked_until, retry_count, max_retries} β€” locked_by + locked_until prevents double-execution across workers. Money Insight: β€œExactly-once” doesn’t exist in distributed systems β€” you get β€œat-least-once + idempotent handlers.” The scheduler uses leader election (Redis SETNX with TTL) to claim a job, but if the worker dies mid-execution, the lock expires and another worker retries. Your handlers MUST be idempotent.


Delayed Trigger Service

Fire callbacks/webhooks at a precise future time (e.g., β€œremind in 30 min,” β€œexpire hold in 10 min”).

Concepts: Timer wheel, Delayed queue, At-least-once delivery, Clock skew handling Stack: Redis Sorted Sets (near-term timers), Kafka (durable delayed queue), Postgres (trigger definitions), Worker pool (dispatcher) Data Model: ZADD triggers:{shard} fire_timestamp triggerId β€” sorted set scored by fire time. Workers poll with ZRANGEBYSCORE triggers:{shard} 0 {now} every second. Money Insight: Don’t use setTimeout or cron for millions of timers. Redis sorted set scored by fire_timestamp gives you O(log N) insert and O(1) poll for due items. For timers >24h out, spill to Kafka with a delay topic β€” keeps Redis memory bounded while handling both seconds-scale and days-scale delays.


Metrics & Monitoring

Time-series ingestion, aggregation, alerting, and dashboarding at scale.

Concepts: Time-series DB, Pre-aggregation, Downsampling, Push vs Pull, Alert evaluation Stack: Prometheus/VictoriaMetrics (TSDB), Kafka (metrics ingest), Grafana (visualization), AlertManager (alerting), S3 (long-term cold storage) Data Model: metric: {name, labels{}} β†’ [(timestamp, value), ...] β€” labels give cardinality. Stored as compressed time-series chunks (Gorilla encoding: 1.37 bytes/point). Money Insight: Cardinality explosion is the silent killer. A metric with labels {host, endpoint, status_code, user_id} creates NΓ—MΓ—KΓ—U unique series. Cap cardinality by NEVER using unbounded values (user_id, request_id) as labels. Pre-aggregate at ingestion, downsample to 1min/5min/1hr windows for cold storage.


Rate Limiter

Distributed rate limiting for APIs β€” token bucket, sliding window, with quota management.

Concepts: Token bucket, Sliding window log, Fixed window counter, Distributed coordination Stack: Redis (counters + Lua scripts), Envoy/nginx (edge enforcement), Postgres (quota config) Data Model: Sliding window counter: MULTI; INCR ratelimit:{userId}:{window}; EXPIRE ratelimit:{userId}:{window} 60; EXEC β€” atomic increment + TTL in one round trip. Money Insight: Fixed window has the boundary problem (2x burst at window edges). Sliding window log is exact but memory-expensive (stores every request timestamp). The sweet spot is sliding window counter: interpolate between current and previous fixed windows. count = prev_window_count * overlap_percentage + current_window_count. Accurate within 0.003% error, O(1) memory.


Rules Engine

Business rules authored by product teams, evaluated against live events in sub-50ms with no redeploy.

Concepts: Rete algorithm, Pre-parsed AST, ANTLR parsing, Hot reload via atomic swap, Two-tier cache Stack: Postgres (rules + immutable versions, AST as JSONB), local process memory (compiled Rete network, L1), Redis (serialized rule set, L2), Kafka (rule.updated propagation + action dispatch), ANTLR (parse conditions at write time) Data Model: rule_versions: (rule_id, version_number) -> {condition_ast JSONB, actions JSONB} β€” append-only, so a rollback is just repointing rules.current_version, and storing the AST pre-parsed means the hot path never runs a parser. Audit rows in rule_executions are PARTITION BY RANGE (evaluated_at) so old days drop cheaply. Money Insight: Per-rule optimization plateaus around 30ms because rules share sub-conditions β€” 200 rules all checking user.country IN high_risk_list evaluate it 200 times. Rete gives each unique condition exactly one alpha node, so asserting the event once propagates the shared result everywhere: at 10K rules with 60% overlap, effective evaluations drop from 50K to ~8K and P99 lands at 12-20ms, for ~50MB per namespace.


Collaboration & AI

Google Docs

Real-time collaborative document editing with conflict resolution.

Concepts: OT (Operational Transform) / CRDT, WebSocket, Cursor presence, Version history Stack: WebSocket (real-time sync), Redis Pub/Sub (presence), Postgres (document snapshots), S3 (revision history), OT server (conflict resolution) Data Model: operations: (doc_id, version) β†’ {type: insert|delete, position, content, author} β€” OT transforms concurrent ops against each other to maintain consistency. Money Insight: OT requires a central server to determine operation ordering (total order). CRDTs (like Yjs/Automerge) are eventually consistent without a server but produce larger payloads. Google chose OT because it gives smaller ops and the server was already there. For an interview, know BOTH and state the tradeoff: OT = simpler ops but needs coordination; CRDT = no coordination but larger state.


ChatGPT / LLM Serving

Serving large language models with streaming responses, context management, and multi-turn conversations.

Concepts: Token streaming (SSE), KV-cache, Model sharding (tensor/pipeline parallelism), Context window management, GPU scheduling Stack: vLLM/TensorRT-LLM (inference), Redis (session + KV-cache), Kafka (request queue), S3 (model weights), GPU cluster (A100/H100) Data Model: sessions: session_id β†’ {messages: [{role, content}], model, temperature} + kv_cache: (session_id, layer) β†’ {key_tensor, value_tensor} β€” KV-cache avoids recomputing attention for prior tokens. Money Insight: GPU memory is the bottleneck, not compute. A 70B parameter model needs 140GB just for weights (FP16). The KV-cache for a 4K context window adds ~2GB per concurrent request. Continuous batching (vLLM’s PagedAttention) lets you serve 10x more concurrent requests by treating KV-cache like virtual memory β€” paging blocks in/out instead of reserving max_seq_len upfront.


Search Autocomplete

Type-ahead suggestions with ranking, personalization, and <100ms latency.

Concepts: Trie, Prefix matching, Precomputation, Ranking by frequency/recency Stack: Redis (precomputed suggestion lists), Kafka (query log ingestion), Spark (offline suggestion recomputation), Elasticsearch (fallback full-text) Data Model: autocomplete:{prefix} β†’ [suggestion_1, suggestion_2, ..., suggestion_10] in Redis β€” precomputed top-10 per prefix. 26^3 = 17K three-char prefixes cover most queries. Money Insight: Don’t build a trie and traverse it at query time β€” that’s O(prefix_length + K) per request and impossible to shard. Instead, precompute the top-10 suggestions for every observed prefix offline (MapReduce over query logs), store in Redis as a flat list. Query becomes a single GET in O(1). This is how Google’s autocomplete actually works.


Ad System

Real-time ad auction, targeting, and delivery with budget pacing.

Concepts: RTB (Real-Time Bidding), Auction (second-price), Budget pacing, CTR prediction, Frequency capping Stack: Aerospike/Redis (ad index + budget counters), Kafka (impression events), Spark/Flink (CTR model training), Elasticsearch (targeting index) Data Model: ad_index: {targeting_criteria} β†’ [ad_candidates] + budgets: (campaign_id, hour) β†’ {spent, limit} β€” budget counters sharded by time window to avoid hot key on popular campaigns. Money Insight: The entire auction must complete in <100ms (including network). You can’t score all 10M ads. Use a funnel: targeting filter (1Mβ†’1K candidates in 5ms) β†’ lightweight CTR model (1Kβ†’100 in 20ms) β†’ full ranking model (100β†’10 in 30ms) β†’ auction (10β†’1 in 5ms). Each stage reduces candidates by 10x.


Nearby Service

Find nearby points of interest (restaurants, stores, friends) within a radius.

Concepts: Geohash, QuadTree, Spatial indexing, Proximity search, Geofencing Stack: Redis Geo / PostGIS (spatial queries), Elasticsearch (geo + attribute filtering), Postgres (POI metadata), CDN (static POI tiles) Data Model: GEOADD pois:{category} lng lat poiId or PostGIS: CREATE INDEX ON pois USING GIST(location geography) β€” geohash prefix matching gives O(1) cell lookup for β€œwhat’s in this area.” Money Insight: Don’t do SELECT * WHERE ST_DWithin(location, point, 5km) ORDER BY distance on 100M rows. Geohash the search area into cells, query only those cells (neighbors included to avoid edge effects), then do exact distance filtering in-memory on the small result set. This turns a full table scan into a bounded index lookup.


Image Processing Service

Async image transformation β€” resize, compress, convert, watermark β€” over an autoscaling worker pool.

Concepts: Presigned-URL direct upload, Queue-depth autoscaling, Idempotency keys, Failure classification + DLQ, Streaming tile pipeline Stack: S3 (presigned uploads + output storage), SQS (job queue with visibility timeout + DLQ), Postgres (job state), Redis (idempotency hashes, 24h TTL), libvips (processing), K8s HPA on queue depth Data Model: jobs: job_id -> {tenant_id, status, operations JSONB, idempotency_key, attempts} with UNIQUE(tenant_id, idempotency_key) plus a partial index ON jobs(status) WHERE status IN ('PENDING','PROCESSING') β€” the index covers only in-flight work instead of bloating with every completed row. Money Insight: Scale on queue depth, not CPU. CPU lags a traffic spike by 2-3 minutes, by which point the latency SLA is already blown, whereas queue_depth / messages_per_worker reacts in 15-30s. And canonicalize before hashing the idempotency key: {"width":800,"height":600} and {"height":600,"width":800} must hash identically or dedup silently does nothing.

Last updated: Quick Revision v1.1 β€” 39 designs, one page, 15-minute review.

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access