Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 38 min read

Designing a Ticket Booking Platform (BookMyShow / Ticketmaster)

Difficulty: Intermediate Topics: Distributed Locking, Seat Reservation, Payment Saga, TTL Holds, Inventory Management Asked at: Ticketmaster, BookMyShow, Amazon, Flipkart, PhonePe Prerequisites:Caching, Message Queues, and Saga Pattern


1. Understanding the Problem

A ticket booking platform lets users browse movies and events, view available seats on a seat map, temporarily hold selected seats while completing payment, and receive confirmed tickets. The hard part? When 100K users rush to book seats for a popular movie premiere at the same time, no two users should ever successfully book the same seat, held seats must auto-release if payment isn’t completed within 10 minutes, and the system must handle payment failures gracefully without leaving seats in a limbo state.


2. Naive First Cut

flowchart LR
    User["User Browser"]:::client
    API["API Server"]:::service
    DB["Postgres DB"]:::data
    PG["Payment Gateway"]:::external

    User --> API
    API --> DB
    API --> PG

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
    classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0
Color Meaning
🟠 Purple-Orange Client apps
πŸ”΅ Blue Edge / Gateway
🟒 Green Backend services
🟑 Yellow Data stores
🟣 Purple Async (Kafka)
πŸ”΄ Pink External services

How this breaks:

The rest of the doc evolves this into a production-grade booking system with distributed locks, TTL-based holds, and payment saga patterns.


3. Prior Art We’re Drawing From


4. Functional Requirements

Core (Top 3)

  1. Browse and select seats - users view available shows, see a real-time seat map, and select specific seats for booking
  2. Hold seats and complete payment - selected seats are temporarily held (10 min TTL) while user completes payment; seats auto-release on expiry
  3. Confirm booking - after successful payment, seats are permanently marked as booked and user receives a ticket/confirmation

Below the Line


5. Non-Functional Requirements

Core

NFR Target
No double-booking Two users must never successfully book the same seat - strong consistency on seat state
Hold expiry Held seats auto-release exactly at TTL expiry (10 min); no manual intervention needed
Peak concurrency Handle 100K+ concurrent booking attempts for a single hot show (movie premiere, concert)
Booking latency Seat hold acquired in < 500ms; end-to-end booking (select β†’ pay β†’ confirm) < 30 seconds

Below the Line

6. Scale Estimation (Back-of-Envelope)


7. Core Entities


8. API / System Interface

GET /api/v1/shows/{showId}/seats
  Response: { seats: [{ seatId, row, number, category, status, price }] }
  Auth: JWT Bearer token
  Note: Returns current availability. HELD seats show as unavailable.

POST /api/v1/bookings/hold
  Body: { showId, seatIds: ["A1", "A2", "A3"] }
  Response: { bookingId, status: "HELD", expiresAt, totalAmount }
  Auth: JWT Bearer token
  Note: Idempotency via clientRequestId header. Hold TTL = 10 minutes.

POST /api/v1/bookings/{bookingId}/pay
  Body: { paymentMethod: "upi", idempotencyKey: "uuid-v4" }
  Response: { bookingId, status: "CONFIRMED", paymentId, tickets[] }
  Auth: JWT Bearer token
  Note: Idempotency key prevents double-charge on retry.

DELETE /api/v1/bookings/{bookingId}
  Response: { status: "CANCELLED", seatsReleased: ["A1", "A2", "A3"] }
  Auth: JWT Bearer token

GET /api/v1/movies?city=bangalore&date=2026-01-15
  Response: { movies: [{ id, title, shows: [{ showId, time, cinema, availability }] }] }
  Auth: Optional (public endpoint, rate limited)

9. High-Level Design

FR1: Browse Shows and View Seat Map

The first interaction: user opens the app, picks a city, selects a movie, chooses a show time, and sees a seat map with real-time availability (green = available, red = booked, yellow = held by someone else). This is a read-heavy path - thousands of users viewing the same show’s seat map simultaneously.

Build it out of one database. A hall’s seat layout is static, and each seat for a given show is in one of three states, so the whole feature is a join and a render. No cache and no separate lock store yet β€” both answer non-functional requirements and both are argued for in the deep dives.

New components we need:

  1. API Gateway - entry point for all client requests. Auth, rate limiting, routing.
  2. Catalog Service - movie listings, show schedules, cinema information.
  3. Seat Service - returns the seat map for one show: the static layout joined against each seat’s current state.
  4. Postgres - cinemas, shows, and a seats row per seat per show carrying status of AVAILABLE, HELD or BOOKED.
flowchart LR
    User["User Browser"]:::client
    GW["API Gateway"]:::edge
    CAT["Catalog Service"]:::service
    SS["Seat Service"]:::service
    DB[("Postgres<br/>shows and seats")]:::data

    User -->|"1. Browse city and movie"| GW
    GW -->|"2. Ask for shows"| CAT
    CAT -->|"3. Read shows"| DB
    GW -->|"4. Ask for a seat map"| SS
    SS -->|"5. Read seats for this show"| DB

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Step-by-step flow:

  1. User picks a city and a movie; the Catalog Service returns the shows from Postgres
  2. User picks a show time
  3. The Seat Service reads the hall’s seat layout, which never changes for a given cinema
  4. It reads the seats rows for this show β€” a few hundred rows, one indexed query
  5. It returns each seat with its status, and the client colours the map green, yellow or red

Why one row per seat per show rather than a list of booked seats on the show? Because the next requirement needs to change one seat’s state without touching its neighbours. A row per seat gives each one its own identity to lock, its own status, and its own holder, which is exactly what FR2 is about to need.

What we have deliberately left broken. This renders a correct seat map, and it is correct only at the instant it is read:


FR2: Hold Seats with TTL

When a user selects seats and clicks β€œProceed to Payment,” we need to temporarily reserve those seats so no one else can book them during the 10-minute payment window. If payment isn’t completed, seats auto-release.

The seats have to stop being available to everyone else while one user goes off to pay, and they have to come back if that payment never happens. Two things: a state change, and a deadline.

New components we need:

  1. Booking Service - owns the booking lifecycle. Checks the seats are free, marks them held, and records the booking.

The deadline needs somewhere to live. The simplest place is the process that created the hold: schedule a timer for ten minutes from now that sets the seats back to AVAILABLE.

flowchart LR
    User["User Browser"]:::client
    GW["API Gateway"]:::edge
    BS["Booking Service"]:::service
    DB[("Postgres<br/>seats and bookings")]:::data
    T["In-process timer<br/>ten minutes"]:::async

    User -->|"1. Proceed to payment"| GW
    GW -->|"2. Hold these seats"| BS
    BS -->|"3. Read the seat rows"| DB
    BS -->|"4. Mark them held"| DB
    BS -->|"5. Schedule the release"| T
    T -->|"6. Set them available again"| DB

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
    classDef async fill:#3a2a4c,stroke:#c084fc,color:#e2e8f0

Step-by-step flow:

  1. User picks A1, A2, A3 and taps Proceed to Payment
  2. The Booking Service reads those three seats rows and checks all three are AVAILABLE
  3. If any is not, it returns an error naming the seats that are gone
  4. Otherwise it writes a bookings row with status HELD and expires_at = now + 10 minutes
  5. It updates the three seat rows to HELD with that booking’s id
  6. It schedules an in-process timer for ten minutes’ time
  7. When the timer fires, if the booking is still HELD, it flips the booking to EXPIRED and the seats back to AVAILABLE

Why hold the seats at all rather than take payment first? Because the alternative is charging someone for a seat we then discover is taken, and a refund is a much worse experience than a seat that was unavailable in the first place. The hold is what lets us promise the user their selection is safe for as long as the countdown says it is.

What we have deliberately left broken. This is the honest version, and both of its halves are unsound:


FR3: Complete Payment and Confirm Booking

After seats are held, the user has 10 minutes to complete payment. Payment can fail (insufficient funds, gateway timeout, OTP expired). We need to handle all failure modes without losing the booking or double-charging.

New components we need:

  1. Payment Service - takes the amount and the booking, calls the gateway, and reports what happened.
  2. Payment Gateway (external) - Razorpay, Stripe or similar. Moves the actual money. We do not control it, cannot see inside it, and cannot include it in our transaction.
flowchart LR
    User["User Browser"]:::client
    BS["Booking Service"]:::service
    PS["Payment Service"]:::service
    PGW["Payment Gateway<br/>external"]:::external
    DB[("Postgres<br/>seats and bookings")]:::data

    User -->|"1. Pay for the booking"| BS
    BS -->|"2. Check the hold is live"| DB
    BS -->|"3. Charge this amount"| PS
    PS -->|"4. Take the money"| PGW
    PGW -->|"5. Approved"| PS
    BS -->|"6. Mark booked and confirmed"| DB

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Step-by-step flow:

  1. User taps Pay while the countdown from FR2 is still running
  2. The Booking Service confirms the booking is still HELD and expires_at has not passed
  3. It calls the Payment Service with the amount and the booking id
  4. The Payment Service calls the gateway and waits for it to come back
  5. On approval, control returns to the Booking Service
  6. It sets the booking to CONFIRMED and the three seats to BOOKED, and the ticket is issued
  7. On decline it leaves the hold alone, so the user can try another card inside their remaining time

Why check the hold again in step 2 rather than trusting the countdown on screen? The timer in the user’s browser is a decoration. The authority is expires_at, and the ten minutes may have elapsed while the payment screen sat open, so the only safe moment to test it is on the server immediately before spending the user’s money.

What we have deliberately left broken. The happy path is right and there is a hole in the middle of it:

All three are Deep Dive 3.


10. Technology Choices

Tier Purpose Stores Access Pattern Primary Alternatives
Booking DB Booking lifecycle state Bookings, seat assignments, payment refs Read/write by bookingId, showId Postgres CockroachDB, TiDB
Seat Lock Store Temporary seat holds with TTL seatId β†’ userId + expiryTime Atomic CAS, TTL-based expiry Redis Cluster DynamoDB (conditional writes)
Event Catalog Movies, shows, cinemas, schedules Catalog metadata Read-heavy, filtered by city/date Postgres (read replicas) MongoDB
Event Bus Async events (booking confirmed, seat released) Booking lifecycle events Pub/sub per show Kafka or Redpanda Kinesis, RabbitMQ
Cache Show listings, seat availability snapshots Aggregated seat counts, catalog data High-QPS reads, short TTL Redis Cluster Memcached
Queue Booking requests during peak (virtual waiting room) User positions, pending requests FIFO with priorities SQS or Kafka RabbitMQ, Redis Streams
Payment Gateway External payment processing Transactions Request-response with webhooks Razorpay or Stripe PayU, PayTM

Why Redis for seat locks, not Postgres row-level locks? During a hot event, 100K users hit the system in 10 seconds. Each seat lock attempt is a SET seatId NX EX 600 (atomic check-and-set with 10-min TTL). Redis handles 300K+ ops/sec per shard in-memory. Postgres row-level locks would create massive lock contention, connection pool exhaustion, and deadlocks. Redis gives us sub-millisecond lock acquisition with built-in TTL for auto-release.

Why Kafka for booking events? A single booking triggers 5+ downstream actions: send confirmation email, update seat map cache, record analytics, trigger invoice generation, update show occupancy counter. Kafka’s consumer groups let each downstream service process independently without blocking the booking path.


11. Data Modeling

Postgres (Booking DB β€” booking lifecycle):

CREATE TABLE bookings (
    booking_id UUID PRIMARY KEY,
    user_id UUID NOT NULL,
    show_id UUID NOT NULL,
    status VARCHAR(15) NOT NULL,  -- INITIATED, HELD, CONFIRMED, CANCELLED, EXPIRED
    total_amount DECIMAL(10,2),
    payment_ref VARCHAR(128),
    idempotency_key UUID UNIQUE,
    created_at TIMESTAMP NOT NULL,
    confirmed_at TIMESTAMP
);
CREATE INDEX idx_bookings_user ON bookings(user_id, created_at DESC);
CREATE INDEX idx_bookings_show ON bookings(show_id, status);

Redis (Seat Lock Store β€” temporary holds with TTL):

Key: "seat:hold:{showId}:{seatId}" β†’ Hash { userId, bookingId, expiresAt }
SET via: SET seat:hold:{showId}:{seatId} NX EX 600  (atomic, 10-min TTL)

Postgres (Seat Availability β€” source of truth after payment):

CREATE TABLE show_seats (
    show_id UUID NOT NULL,
    seat_id VARCHAR(10) NOT NULL,  -- e.g., "A1", "B12"
    category VARCHAR(10) NOT NULL, -- SILVER, GOLD, PLATINUM
    price DECIMAL(8,2) NOT NULL,
    status VARCHAR(10) NOT NULL DEFAULT 'AVAILABLE',  -- AVAILABLE, BOOKED
    booked_by UUID,
    PRIMARY KEY (show_id, seat_id)
);

Access Patterns:

Query Data Source How
Show available seats Redis (cache) + Postgres Cache seat map per show with 5s TTL; fallback to DB
Hold seats (atomic) Redis SET seat:hold:{showId}:{seatId} NX EX 600 per seat β€” all-or-nothing in Lua script
Confirm booking Postgres In txn: UPDATE show_seats SET status='BOOKED' + UPDATE bookings SET status='CONFIRMED'
Auto-release expired holds Redis TTL Keys auto-expire after 600s; no sweeper needed
Check if seat is held Redis EXISTS seat:hold:{showId}:{seatId}

How Seat Locking Prevents Double-Booking During a Hot Event (100K concurrent users):

  1. User selects seats A1, A2, A3 β†’ API sends POST /bookings/hold
  2. Server executes a Redis Lua script (atomic, no race conditions):
    for each seatId in [A1, A2, A3]:
      result = SET seat:hold:{showId}:{seatId} {userId} NX EX 600
      if result == nil: ROLLBACK all previous SETs, return CONFLICT
    
  3. All-or-nothing: if ANY seat is already held, none are held (no partial locks)
  4. On success: create booking in Postgres with status=HELD, return booking_id + payment URL
  5. User has 10 minutes to pay. On payment success β†’ update Postgres show_seats to BOOKED
  6. If user abandons β†’ Redis keys auto-expire after 10 min, seats become available again
  7. No background sweeper needed for Redis path β€” TTL handles cleanup automatically

12. Deep Dives

1) A thousand people tap the same seat in the same second. How does exactly one get it?

Problem: 1000 users click β€œHold Seat A5” within the same second for a hot show. Exactly one must succeed; 999 must fail cleanly.

Bad: What FR2 built β€” read the seat rows, check they say AVAILABLE, then write HELD. The check and the write are two separate statements, and nothing stops other requests running between them. Every concurrent request reads the same AVAILABLE, every one concludes it may proceed, and every one writes. The last write decides who the row says owns the seat, while every one of those users has already been told their hold succeeded. The result is not a near miss: on a hot release with a thousand taps in one second, the expected outcome is that we sell one seat many times and find out at the door. Adding a re-read after the write does not fix it either, because two requests can both re-read after both have written.

Good: Postgres row-level lock. SELECT ... FOR UPDATE on the seat row, then update status. Works for moderate concurrency, but under 1000 concurrent transactions, you get lock contention, connection pool starvation, and 5+ second response times. Postgres serializes all competing transactions - they queue up behind each other.

Great: Redis SET NX EX (atomic conditional set with TTL).
πŸ’‘ SET NX = β€œSet only if Not eXists” - an atomic compare-and-swap in one round trip. Combined with EX (expire), it gives us a lock that auto-releases.

flowchart LR
    LM["Lock Manager"]:::service
    R1["Redis Shard 1"]:::data
    R2["Redis Shard 2"]:::data
    R3["Redis Shard 3"]:::data
    BS["Booking Service"]:::service

    BS -->|"1. Acquire seats"| LM
    LM -->|"2. Show:1:seat:A*"| R1
    LM -->|"3. Show:1:seat:B*"| R2
    LM -->|"4. Show:2:seat:*"| R3

    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Mechanism:

  1. Lock key format: show:{showId}:seat:{seatId} β†’ value: {bookingId}:{userId}:{timestamp}
  2. Lock acquisition (Lua script for multi-seat atomicity):
    -- Atomic multi-seat lock
    for each seatKey in requested_seats:
        result = SET seatKey bookingId NX EX 600
        if result == nil: -- someone else holds it
            -- rollback: delete all keys we just set
            for each acquired_key: DEL acquired_key
            return FAILED + which seat was taken
    return SUCCESS
    
  3. NX guarantees only one caller succeeds for each seat - Redis is single-threaded per shard
  4. EX 600 (10 minutes) ensures auto-release if payment isn’t completed
  5. Sharding: keys are sharded by showId across Redis cluster nodes. One hot show maps to one shard - that shard handles all lock contention for that show

Why not Redlock?

Redlock (consensus across N Redis nodes) adds 3-5ms latency and complexity. For seat booking, a single Redis shard with persistence (AOF every second) is sufficient. The worst case of a Redis crash is: users who had holds lose them (10-minute window restarts). This is acceptable - they can re-select and re-hold. We don’t need the durability guarantees of Redlock.

Backstop: Postgres as final gate

Even though Redis handles the fast path, the Booking Service writes the confirmed booking to Postgres with a unique constraint: UNIQUE(showId, seatId) WHERE status = 'CONFIRMED'. If Redis somehow fails (split-brain, data loss), Postgres prevents actual double-booking at the persistence layer.


2) A deploy restarts the server mid-hold. Who releases those seats?

Problem: User holds seats, goes to make tea, never pays. Those seats must become available again exactly at TTL expiry. But Redis TTL deletion is lazy (checked on access or via periodic sampling) - there’s no guarantee of exact-millisecond release.

Bad: What FR2 built β€” an in-process timer on whichever instance served the hold. The timer’s lifetime is the process’s lifetime, so a restart, a deploy, an autoscaler scale-in or a crash takes it with them, and no other instance has any idea the hold existed. Those seats sit HELD permanently: not paid for, not purchasable, gone from inventory with no record of why. A routine deploy therefore quietly destroys seats on every show with a hold open at that moment, which on a busy evening is most of them. The damage also accumulates silently, because nothing in the system is looking for holds that should have expired.

Good: Rely purely on Redis TTL. When another user queries seat status, the key is gone (expired), so they can lock it. Works for the lock itself, but: the booking record in Postgres still says β€œHELD” - inconsistency.

Great: Redis TTL for lock release + background reconciler for state consistency + Kafka event for downstream notifications.

In simple terms: When you select a seat, we lock it for 10 minutes using Redis with an auto-expiry timer. If you don’t pay in time, the lock disappears automatically and the seat becomes available again. A background job double-checks that Redis and the database agree.

Mechanism:

  1. Redis TTL (primary): The lock key expires after exactly 600 seconds. After expiry, any new SET NX on that key will succeed - seat is available for others.
  2. Reconciler (consistency): Runs every 30 seconds. Queries Postgres: SELECT * FROM bookings WHERE status = 'HELD' AND expires_at < NOW(). For each:
    • Updates status to EXPIRED
    • Publishes booking.expired event to Kafka
    • Kafka consumers: notify user (β€œYour hold expired”), update analytics
  3. Redis Keyspace Notifications (optional enhancement): Subscribe to __keyevent@0__:expired events. When a lock key expires, immediately trigger the Postgres update instead of waiting for the 30-second reconciler sweep.

Edge case - race between payment and expiry:

User pays at minute 9:58 (2 seconds before expiry). Payment gateway takes 5 seconds to process. By the time we get the success response, the Redis TTL has expired and someone else might have locked the seat.

Solution: Before calling the payment gateway, extend the Redis lock by 2 minutes (safety buffer): EXPIRE show:{showId}:seat:{seatId} 720. This gives us time to process the payment response. If payment fails, we explicitly DEL the lock.


3) The money moved and then our server died. How does the customer get their ticket?

Problem: Booking involves two systems: our seat lock (Redis + Postgres) and an external payment gateway. These can’t be in a single database transaction. What if payment succeeds but our server crashes before confirming the booking? What if we confirm the booking but payment actually failed (gateway network error)?

Bad: What FR3 built β€” call the gateway inline, then write CONFIRMED when it returns. The two steps are a charge at another company and a write in our database, and there is no way to make them one operation. If the process dies in between, the customer has paid and has no ticket, and nothing is left that knows to look. A timed-out gateway call is the same wound from the other side: we cannot tell whether the money moved, so retrying may charge twice and not retrying may strand a payment.

The instinctive fix is a distributed transaction across both β€” two-phase commit. It is not available: no public payment gateway exposes a 2PC participant interface, so there is nothing to enlist. Even between systems we control, 2PC holds locks for the duration of the round trip and leaves participants blocked indefinitely if the coordinator dies mid-commit, which is a poor trade for a ten-minute human checkout. The gap has to be closed by making each step individually recoverable rather than by pretending the two are one.

Good: Optimistic approach - assume payment will succeed, confirm booking first, then process payment. If payment fails, roll back the booking. Problem: user already has a β€œconfirmed” ticket for a brief moment (bad UX and potential fraud vector).

Great: Saga pattern with explicit compensating transactions. (Borrowing from Stripe’s idempotency key pattern and Razorpay’s orchestration.)

In simple terms: Booking involves multiple steps (lock seat β†’ charge card β†’ confirm). If step 2 fails (card declined), we automatically β€œundo” step 1 (release the seat). Each step has a pre-defined rollback action.

flowchart LR
    BS["Booking Service Orchestrator"]:::service
    PS["Payment Service"]:::service
    RL["Redis Locks"]:::data
    DB["Postgres"]:::data
    PGW["Payment Gateway"]:::external
    KF["Kafka"]:::async

    BS -->|"1. Verify hold"| RL
    BS -->|"2. Initiate payment"| PS
    PS -->|"3. Charge"| PGW
    PGW -->|"4. Success/Fail"| PS
    PS -->|"5. Result"| BS
    BS -->|"6. Confirm or rollback"| DB
    BS -->|"7. Event"| KF

    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
    classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0
    classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0

Saga steps:

Step Action Compensating Transaction
1 Verify hold is still valid (Redis lock exists) - (read-only check)
2 Extend lock TTL by 2 min (safety buffer) Restore original TTL
3 Call payment gateway with idempotencyKey Refund payment
4 On success: update booking to CONFIRMED -
5 On failure: release lock explicitly -

Idempotency key lifecycle:

  1. Client generates a UUID v4 as idempotencyKey before clicking Pay
  2. Server stores {idempotencyKey: bookingId, status: PROCESSING} in Redis with 24hr TTL
  3. If the same request arrives again (retry), check Redis: if status = PROCESSING β†’ return β€œstill processing”; if status = SUCCESS β†’ return cached response; if status = FAILED β†’ allow retry with same key
  4. Gateway uses the same key: same charge is never processed twice

Handling β€œzombie payments” (success webhook arrives after hold expired):


4) 500K people arrive in ten seconds for one premiere. Who gets served and who waits?

Problem: Avengers premiere - 500K users hit the booking page in 10 seconds. The Redis lock store for that show gets 500K SET NX attempts/sec. Even Redis will struggle, and the API servers will be overwhelmed.

Bad: What the design so far does β€” nothing. Every one of the 500K arrivals is admitted straight to the seat map and the hold endpoint. FR1’s reads alone are 50K queries/sec against one show’s few hundred rows, and FR2’s holds all contend on those same rows, so the database saturates and requests start timing out. Then it compounds: every timed-out client retries, so load rises after the failure, and the system stays down for minutes rather than seconds. What the users experience is worse than slowness β€” it is arbitrary. Whoever has the lowest latency and the most aggressive retry loop wins the seats, so the outcome is decided by network proximity and client behaviour rather than by who arrived first.

Good: Rate limit at the API gateway (500 requests/sec). Most users get rejected. Better for the system, but terrible UX - β€œtry again later” for 499K users.

Great: Virtual waiting room with fair queue. (Borrowing from Ticketmaster’s approach.)

In simple terms: During a hot event (Avengers premiere), instead of letting 500K people hammer the system simultaneously, we put them in a virtual queue. We let them through in batches of 1000, preventing the system from crashing while keeping it fair (first come, first served).

flowchart LR
    Users["500K Users"]:::client
    WR["Waiting Room Service"]:::edge
    Q["Priority Queue SQS"]:::async
    BS["Booking Service"]:::service
    RL["Redis Locks"]:::data

    Users -->|"1. Enter waiting room"| WR
    WR -->|"2. Assign position"| Q
    Q -->|"3. Dequeue 100/sec"| BS
    BS -->|"4. Acquire seat lock"| RL

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Mechanism:

  1. Trigger: When a show’s booking page hits > 10K concurrent viewers (tracked via WebSocket connections or session counter), activate the waiting room for that show.
  2. Assign position: Each user arriving at the booking page gets a random position in the queue (not first-come-first-served - avoids bot advantage). Position = hash(userId + salt + timestamp_bucket).
  3. Drip processing: Waiting Room Service dequeues users at a controlled rate (100-500 users/sec) based on the downstream system’s capacity.
  4. User experience: User sees β€œYou’re #4,521 in line. Estimated wait: 2 minutes.” Position updates in real-time via SSE.
  5. When it’s their turn: User gets a time-limited token (2 minutes validity) that allows them to access the seat selection page. Token is validated at the API Gateway.
  6. Overflow: If all seats are sold while users are in queue, remaining queue members are notified β€œSold Out” and queue is drained.

Why random position instead of arrival order?

Arrival-order queues reward bots and users with faster network connections. Random assignment is fairer and eliminates the incentive to DDoS the system at second zero.

Capacity math: A cinema hall has 300 seats. Even if all 300 are booked in one go (unlikely - most users book 2-4 seats), we only need ~100 successful booking attempts. Processing 500 users/sec means the entire queue is served in ~17 minutes for a 500K-user queue. With 80% dropping off or failing, actual booking completes in 3-5 minutes.


5) Someone takes A5 while a thousand people are looking at it. How do their screens find out?

Problem: 1000 users are viewing the same seat map. When User A holds seat A5, all other users should see it turn yellow (held) within 2-3 seconds. Otherwise, they’ll select the same seat and get frustrated when their hold fails.

Bad: What FR1 leaves the client to do β€” re-request the whole seat map on a timer. At 1000 concurrent viewers polling every 2 seconds that is 500 requests/sec for a single show, each one returning a few hundred seats where typically nothing has changed, and it multiplies by every show running concurrently. The waste is the smaller complaint. A 2-second interval means a seat can be taken and up to 2 seconds pass before anyone else’s screen shows it, which is exactly the window in which someone picks it, submits, and is told no. Shortening the interval to close that window multiplies a request rate that is already almost entirely redundant.

Good: Short polling with aggressive caching. Cache the seat map in Redis with 3-second TTL. 500 requests/sec all hit cache. Works, but users still see stale data for up to 3 seconds.

Great: Server-Sent Events (SSE) for seat status push + event-driven invalidation.

Mechanism:

  1. When a user opens the seat map page, browser opens an SSE connection: GET /sse/v1/shows/{showId}/seats (long-lived HTTP connection with text/event-stream)
  2. Gateway registers this connection to a show-specific channel
  3. When any seat’s status changes (held, released, booked), the Booking Service publishes to Kafka topic show.{showId}.seat-updates
  4. A Seat Update Consumer reads from Kafka and pushes to all SSE connections for that show via Redis Pub/Sub β†’ SSE Gateway
  5. Client receives: data: {"seatId": "A5", "status": "held", "heldBy": "someone"} and updates the seat map UI immediately

Why SSE instead of WebSocket?

Seat map updates are serverβ†’client only (users don’t send data back over this connection). SSE is simpler, works over HTTP/2 multiplexing, auto-reconnects, and requires no upgrade handshake. WebSocket would be overkill for a one-directional push.

Scaling: With SSE over HTTP/2, a single gateway can hold 100K+ connections (far cheaper than WebSocket). For hot shows with 50K concurrent viewers, 1 gateway instance suffices for the seat map push path.


13. Design Self-Audit

Question Answer
Dedicated search index? Not needed for core booking flow. Movie/event discovery can use Postgres full-text search or Elasticsearch for advanced filtering (genre, language, nearby cinemas). Low priority - not on the critical booking path.
Stale reads after writes? Seat map has 2-3 second lag via SSE push (acceptable). After YOUR hold succeeds, your own UI updates immediately (optimistic). Other users may attempt the same seat and fail - clean error handling.
Single points of failure? Redis lock store uses Redis Cluster (3+ shards, each with a replica). Booking Service is stateless, horizontally scaled. Postgres uses primary + synchronous standby for booking confirmations.
Dead-letter / reconciliation? Reconciler every 30s cleans expired holds. Failed payment webhooks go to DLQ with exponential retry (1min, 5min, 30min). Orphaned bookings (INITIATED for > 15 min) are auto-cancelled.
Data freshness across caches? Catalog cache TTL = 60s (acceptable for movie listings). Seat map is real-time via SSE. Aggregated availability (β€œ45 seats left”) is eventually consistent (5s lag).
Cost at scale? Redis Cluster (6 nodes for locks): ~$2000/month. Kafka (3 brokers): ~$1500/month. API + Booking Service (20 instances): ~$4000/month. Postgres RDS (primary + standby): ~$2000/month. Total: ~$10K/month for 1M bookings/day platform.

14. Core Flows

Flow 1: Seat Hold with Concurrent Competition

sequenceDiagram
    participant U1 as User 1
    participant U2 as User 2
    participant GW as API Gateway
    participant BS as Booking Service
    participant LM as Lock Manager
    participant Redis as Redis Lock Store

    Note over U1,U2: Both want seat A5 for same show
    U1->>GW: POST /bookings/hold (seats: [A5, A6])
    U2->>GW: POST /bookings/hold (seats: [A5, A7])
    GW->>BS: User1 hold request
    GW->>BS: User2 hold request
    BS->>LM: Lock [A5, A6] for User1
    BS->>LM: Lock [A5, A7] for User2
    LM->>Redis: SET show:1:seat:A5 booking1 NX EX 600
    Redis-->>LM: OK (User1 wins)
    LM->>Redis: SET show:1:seat:A6 booking1 NX EX 600
    Redis-->>LM: OK
    LM-->>BS: All locks acquired for User1
    LM->>Redis: SET show:1:seat:A5 booking2 NX EX 600
    Redis-->>LM: nil (FAILED - already locked)
    LM-->>BS: Lock failed for User2 on seat A5
    BS-->>U1: 200 OK - Seats held. Pay within 10 min.
    BS-->>U2: 409 Conflict - Seat A5 unavailable

Non-obvious failure path: What if the Booking Service crashes after acquiring locks in Redis but before writing the booking to Postgres? The Redis locks have a 10-minute TTL - they’ll auto-expire. The reconciler won’t find a matching booking in Postgres (it was never written), so no cleanup needed. The seats become available again after TTL expires. Worst case: 10 minutes of phantom unavailability for those seats.

Flow 2: Payment Completion with Failure Handling

sequenceDiagram
    participant User
    participant BS as Booking Service
    participant PS as Payment Service
    participant PGW as Payment Gateway
    participant Redis as Redis Locks
    participant DB as Postgres

    User->>BS: POST /bookings/{id}/pay
    BS->>Redis: Check lock exists for bookingId
    Redis-->>BS: Lock valid (5 min remaining)
    BS->>PS: Process payment (amount + idempotencyKey)
    PS->>PGW: Create payment intent
    PGW-->>PS: Payment processing...
    alt Payment succeeds
        PGW-->>PS: Success + transactionId
        PS-->>BS: Payment confirmed
        BS->>DB: Update booking: CONFIRMED
        BS->>Redis: PERSIST lock (remove TTL)
        BS-->>User: 200 OK - Booking confirmed + ticket
    else Payment fails (insufficient funds)
        PGW-->>PS: Failed: insufficient funds
        PS-->>BS: Payment failed
        BS-->>User: 402 - Payment failed. Retry with another method.
        Note over Redis: Hold still valid. User can retry.
    else Hold expired during payment
        PGW-->>PS: Success + transactionId
        PS->>BS: Payment confirmed
        BS->>Redis: Check lock exists
        Redis-->>BS: Lock expired (nil)
        BS->>PS: Trigger refund
        PS->>PGW: Refund transactionId
        BS-->>User: 410 Gone - Hold expired. Refund initiated.
    end

Non-obvious failure path: Payment gateway sends a success webhook but our server crashes before processing it. The gateway will retry the webhook (typically 3-5 times over 24 hours). Payment Service uses the idempotencyKey to detect duplicate webhooks and skip re-processing. If all webhook retries fail, a reconciler job polls the gateway every 5 minutes for recent payments and matches them against pending bookings.

Booking Lifecycle State Machine

stateDiagram-v2
    [*] --> INITIATED : User selects seats
    INITIATED --> HELD : All locks acquired
    INITIATED --> FAILED : Lock conflict (seats taken)
    HELD --> CONFIRMED : Payment success
    HELD --> EXPIRED : TTL expired (10 min)
    HELD --> CANCELLED : User cancels
    CONFIRMED --> REFUNDED : Admin refund or cancellation policy
    EXPIRED --> [*]
    FAILED --> [*]
    CONFIRMED --> [*]
    CANCELLED --> [*]
    REFUNDED --> [*]

Each transition emits a Kafka event consumed by: Notification Service (user updates), Seat Map Cache (invalidation), Analytics (conversion tracking), and the Reconciler (consistency checks).


15. Final Architecture

flowchart TB
    UA(["User Browser and App"]):::client
    GW["API Gateway<br>auth and rate limiting"]:::edge
    WR["Waiting Room<br>admits a trickle"]:::edge
    SSE["SSE Gateway<br>seat map updates"]:::edge
    CAT["Catalog Service<br>shows and seat maps"]:::service
    SS["Seat Service<br>holds and releases seats"]:::service
    BS["Booking Service<br>owns booking lifecycle"]:::service
    PS["Payment Service"]:::service
    NS["Notification Service"]:::service
    KF[["Kafka<br>booking events"]]:::async
    PG[("Postgres<br>bookings and seats")]:::data
    RD[("Redis<br>seat locks and cache")]:::data
    PGW[/"Payment Gateway"/]:::external

    UA -->|"browse and book"| GW
    GW -->|"peak traffic"| WR
    WR --> CAT
    GW --> CAT
    CAT --> RD
    GW --> SS
    SS -->|"lock the seat"| RD
    SS --> BS
    BS --> PG
    BS --> PS
    PS --> PGW
    BS -->|"booking confirmed"| KF
    KF --> NS
    KF --> SSE
    SSE -->|"seats just went"| UA
    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
    classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0

How it works end-to-end:

  1. User sends request β€” hits the Load Balancer, routed through API Gateway
  2. Waiting Room queues during surge β€” hot event on-sales throttle via the Booking Queue
  3. Catalog Service serves seat map β€” reads from Redis Cache (or Postgres on miss)
  4. Seat Service checks availability β€” queries Redis Lock Store for real-time seat state
  5. Booking Service acquires lock β€” Lock Manager does a Redis SET NX with TTL on selected seats
  6. Payment is processed β€” Payment Service calls the external Payment Gateway with idempotency key
  7. Booking confirmed and persisted β€” writes to Postgres, emits event to Kafka
  8. Downstream consumers react β€” Confirmation Service generates ticket, Notification Service sends email/push, SSE Gateway updates live seat maps
  9. Reconciler handles orphans β€” releases expired locks and syncs Redis with Postgres

Want a deep dive on multi-cinema franchise inventory aggregation, dynamic pricing (surge for hot shows), or fraud detection (bot-booking prevention)? Drop a comment below πŸ‘‡


Key Technologies

Term What it is
Redis SET NX (distributed lock) Atomic β€œset if not exists” command used to acquire a seat lock - only one caller succeeds, preventing double-booking.
TTL-based hold Lock keys expire automatically after 10 minutes, releasing seats if payment isn’t completed.
CQRS Command Query Responsibility Segregation - separating the write path (seat locks, bookings) from read path (seat map browsing) for independent scaling.
Kafka Event bus carrying booking lifecycle events to downstream services (notifications, analytics, seat map invalidation).
WebSocket Persistent connection for pushing real-time seat availability updates to users viewing the same show.
Postgres ACID-compliant relational DB serving as the booking source of truth with unique constraints as a double-booking backstop.
Idempotency Key Client-generated UUID ensuring payment retries never double-charge - gateway returns the same result for repeated keys.

What’s Expected at Each Level

This section helps you calibrate your depth. You don’t need to cover everything - just know what’s expected for your level.

Mid-level

Design a basic seat selection and booking system with a database. Recognize the concurrency problem - two users selecting the same seat simultaneously. Propose a locking mechanism with prompting. You should articulate why naive check-then-act creates race conditions and sketch a happy-path flow from seat selection through payment.

Senior

Propose Redis SET NX with TTL for seat holds. Explain atomic multi-seat locking via a Lua script (all-or-nothing semantics). Discuss the payment saga pattern and idempotency without prompting. Recognize the need for a hold expiry reconciler to keep Postgres consistent with Redis TTL state. Articulate why Postgres row-level locks fail under 100K concurrent users.

Staff+

Address the thundering herd problem on hot events by proposing a virtual waiting room or queue-based admission control. Discuss Postgres as a backstop with unique constraints even when Redis is the fast path. Proactively mention how CDN invalidation works for seat maps, zombie payment handling (success webhook after hold expiry), and the cost trade-off of keeping expired holds in Redis vs background cleanup. Quantify the peak load numbers and explain capacity planning.


🎯 Key Takeaways



Understand the building blocks used in this design:

Discussion

Newest first
You

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access