Designing a Ticket Booking Platform (BookMyShow / Ticketmaster)
Difficulty: Intermediate Topics: Distributed Locking, Seat Reservation, Payment Saga, TTL Holds, Inventory Management Asked at: Ticketmaster, BookMyShow, Amazon, Flipkart, PhonePe Prerequisites:Caching, Message Queues, and Saga Pattern
1. Understanding the Problem
A ticket booking platform lets users browse movies and events, view available seats on a seat map, temporarily hold selected seats while completing payment, and receive confirmed tickets. The hard part? When 100K users rush to book seats for a popular movie premiere at the same time, no two users should ever successfully book the same seat, held seats must auto-release if payment isnβt completed within 10 minutes, and the system must handle payment failures gracefully without leaving seats in a limbo state.
2. Naive First Cut
flowchart LR
User["User Browser"]:::client
API["API Server"]:::service
DB["Postgres DB"]:::data
PG["Payment Gateway"]:::external
User --> API
API --> DB
API --> PG
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0
| Color | Meaning |
|---|---|
| π Purple-Orange | Client apps |
| π΅ Blue | Edge / Gateway |
| π’ Green | Backend services |
| π‘ Yellow | Data stores |
| π£ Purple | Async (Kafka) |
| π΄ Pink | External services |
How this breaks:
- Two users select the same seat simultaneously β both see it as βavailableβ β both attempt to book β double-booking
- If payment fails after marking seat as βbooked,β the seat is stuck - no one can book it (orphaned reservation)
- Single API server canβt handle 100K concurrent users rushing for a hot event (Avengers premiere, IPL final)
- No seat hold mechanism - user selects seats, goes to payment, comes back 5 minutes later and seats were taken by someone else
- No way to show real-time seat availability updates to other users viewing the same show
- Single Postgres DB becomes a write bottleneck when 50K users try to lock seats simultaneously
The rest of the doc evolves this into a production-grade booking system with distributed locks, TTL-based holds, and payment saga patterns.
3. Prior Art Weβre Drawing From
- Ticketmaster Virtual Queue - Uses a virtual waiting room during high-demand on-sales. Users are assigned random positions in a queue, preventing thundering herd on the booking system. Processes users in controlled batches. (Ticketmaster Tech Blog)
- Stripe Idempotency Keys - Guarantees exactly-once payment processing using client-generated idempotency keys. If a payment request is retried, the same result is returned without double-charging. (Stripe Engineering)
- BookMyShow Engineering - Handles 20M+ users for major movie releases using Redis-based seat locking with TTL, event-driven architecture with Kafka, and eventual consistency for non-critical reads. (BookMyShow Engineering Blog)
- Razorpay Payment Orchestration - Multi-gateway payment routing with automatic failover. Implements saga pattern for coordinating booking + payment as a distributed transaction. (Razorpay Engineering)
- Amazon DynamoDB Transactions - Demonstrates conditional writes with version checks for exactly-once operations in distributed systems. Used internally for inventory management at Amazon retail. (AWS Blog)
4. Functional Requirements
Core (Top 3)
- Browse and select seats - users view available shows, see a real-time seat map, and select specific seats for booking
- Hold seats and complete payment - selected seats are temporarily held (10 min TTL) while user completes payment; seats auto-release on expiry
- Confirm booking - after successful payment, seats are permanently marked as booked and user receives a ticket/confirmation
Below the Line
- Browse movies/events by city, genre, language
- Cancellation and refunds
- Promotional codes and discounts
- Notifications (booking confirmation, show reminders)
- Reviews and ratings
- Waitlist for sold-out shows
5. Non-Functional Requirements
Core
| NFR | Target |
|---|---|
| No double-booking | Two users must never successfully book the same seat - strong consistency on seat state |
| Hold expiry | Held seats auto-release exactly at TTL expiry (10 min); no manual intervention needed |
| Peak concurrency | Handle 100K+ concurrent booking attempts for a single hot show (movie premiere, concert) |
| Booking latency | Seat hold acquired in < 500ms; end-to-end booking (select β pay β confirm) < 30 seconds |
Below the Line
- 99.99% availability for browsing (eventual consistency acceptable)
- Seat map renders in < 2 seconds with real-time availability
- Payment processing SLA: < 10 seconds
- Support 50K+ shows across 1000+ cinemas
6. Scale Estimation (Back-of-Envelope)
- Users: 10M DAU, 100K+ concurrent during hot event launch (Avengers premiere, IPL final)
- Write QPS: 50K booking attempts/min peak (~833/sec sustained, bursty during on-sale windows)
- Read QPS: 300K seat-status reads/sec during on-sale (50K users refreshing seat maps every 2-3s)
- Storage: ~500GB booking data/year (5M bookings/day Γ booking + payment metadata)
- Bandwidth: ~1 Gbps at peak (seat map pushes via SSE + API responses)
7. Core Entities
- Movie/Event - id, title, genre, language, duration, poster, rating
- Show - id, movieId, cinemaHallId, startTime, endTime, pricing tiers
- Seat - id, hallId, row, number, category (Silver/Gold/Platinum), status
- ShowSeat - showId + seatId composite, state (AVAILABLE/HELD/BOOKED), heldBy, heldUntil, bookedBy
- Booking - id, userId, showId, seatIds[], status (INITIATED/HELD/CONFIRMED/CANCELLED/EXPIRED), paymentRef, totalAmount
- Payment - id, bookingId, amount, status (PENDING/SUCCESS/FAILED/REFUNDED), gateway, idempotencyKey
8. API / System Interface
GET /api/v1/shows/{showId}/seats
Response: { seats: [{ seatId, row, number, category, status, price }] }
Auth: JWT Bearer token
Note: Returns current availability. HELD seats show as unavailable.
POST /api/v1/bookings/hold
Body: { showId, seatIds: ["A1", "A2", "A3"] }
Response: { bookingId, status: "HELD", expiresAt, totalAmount }
Auth: JWT Bearer token
Note: Idempotency via clientRequestId header. Hold TTL = 10 minutes.
POST /api/v1/bookings/{bookingId}/pay
Body: { paymentMethod: "upi", idempotencyKey: "uuid-v4" }
Response: { bookingId, status: "CONFIRMED", paymentId, tickets[] }
Auth: JWT Bearer token
Note: Idempotency key prevents double-charge on retry.
DELETE /api/v1/bookings/{bookingId}
Response: { status: "CANCELLED", seatsReleased: ["A1", "A2", "A3"] }
Auth: JWT Bearer token
GET /api/v1/movies?city=bangalore&date=2026-01-15
Response: { movies: [{ id, title, shows: [{ showId, time, cinema, availability }] }] }
Auth: Optional (public endpoint, rate limited)
9. High-Level Design
FR1: Browse Shows and View Seat Map
The first interaction: user opens the app, picks a city, selects a movie, chooses a show time, and sees a seat map with real-time availability (green = available, red = booked, yellow = held by someone else). This is a read-heavy path - thousands of users viewing the same showβs seat map simultaneously.
Build it out of one database. A hallβs seat layout is static, and each seat for a given show is in one of three states, so the whole feature is a join and a render. No cache and no separate lock store yet β both answer non-functional requirements and both are argued for in the deep dives.
New components we need:
- API Gateway - entry point for all client requests. Auth, rate limiting, routing.
- Catalog Service - movie listings, show schedules, cinema information.
- Seat Service - returns the seat map for one show: the static layout joined against each seatβs current state.
- Postgres -
cinemas,shows, and aseatsrow per seat per show carryingstatusofAVAILABLE,HELDorBOOKED.
flowchart LR
User["User Browser"]:::client
GW["API Gateway"]:::edge
CAT["Catalog Service"]:::service
SS["Seat Service"]:::service
DB[("Postgres<br/>shows and seats")]:::data
User -->|"1. Browse city and movie"| GW
GW -->|"2. Ask for shows"| CAT
CAT -->|"3. Read shows"| DB
GW -->|"4. Ask for a seat map"| SS
SS -->|"5. Read seats for this show"| DB
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Step-by-step flow:
- User picks a city and a movie; the Catalog Service returns the shows from Postgres
- User picks a show time
- The Seat Service reads the hallβs seat layout, which never changes for a given cinema
- It reads the
seatsrows for this show β a few hundred rows, one indexed query - It returns each seat with its status, and the client colours the map green, yellow or red
Why one row per seat per show rather than a list of booked seats on the show? Because the next requirement needs to change one seatβs state without touching its neighbours. A row per seat gives each one its own identity to lock, its own status, and its own holder, which is exactly what FR2 is about to need.
What we have deliberately left broken. This renders a correct seat map, and it is correct only at the instant it is read:
- The map is stale the moment it reaches the screen. A thousand people looking at the same popular show all hold a snapshot from whenever their request landed. Someone else takes A5 a second later and nobodyβs view changes, so several users pick a seat they cannot have and only discover it when their booking fails. That is Deep Dive 5.
- Popular shows concentrate all reads on a few hundred rows. 50K people opening the same premiereβs seat map is 50K queries/sec against one showβs rows, and the requirement is a page that loads instantly. That load, and what to do about the surge in general, is Deep Dive 4.
FR2: Hold Seats with TTL
When a user selects seats and clicks βProceed to Payment,β we need to temporarily reserve those seats so no one else can book them during the 10-minute payment window. If payment isnβt completed, seats auto-release.
The seats have to stop being available to everyone else while one user goes off to pay, and they have to come back if that payment never happens. Two things: a state change, and a deadline.
New components we need:
- Booking Service - owns the booking lifecycle. Checks the seats are free, marks them held, and records the booking.
The deadline needs somewhere to live. The simplest place is the process that created the hold:
schedule a timer for ten minutes from now that sets the seats back to AVAILABLE.
flowchart LR
User["User Browser"]:::client
GW["API Gateway"]:::edge
BS["Booking Service"]:::service
DB[("Postgres<br/>seats and bookings")]:::data
T["In-process timer<br/>ten minutes"]:::async
User -->|"1. Proceed to payment"| GW
GW -->|"2. Hold these seats"| BS
BS -->|"3. Read the seat rows"| DB
BS -->|"4. Mark them held"| DB
BS -->|"5. Schedule the release"| T
T -->|"6. Set them available again"| DB
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
classDef async fill:#3a2a4c,stroke:#c084fc,color:#e2e8f0
Step-by-step flow:
- User picks A1, A2, A3 and taps Proceed to Payment
- The Booking Service reads those three
seatsrows and checks all three areAVAILABLE - If any is not, it returns an error naming the seats that are gone
- Otherwise it writes a
bookingsrow with statusHELDandexpires_at = now + 10 minutes - It updates the three seat rows to
HELDwith that bookingβs id - It schedules an in-process timer for ten minutesβ time
- When the timer fires, if the booking is still
HELD, it flips the booking toEXPIREDand the seats back toAVAILABLE
Why hold the seats at all rather than take payment first? Because the alternative is charging someone for a seat we then discover is taken, and a refund is a much worse experience than a seat that was unavailable in the first place. The hold is what lets us promise the user their selection is safe for as long as the countdown says it is.
What we have deliberately left broken. This is the honest version, and both of its halves are unsound:
- Step 2 and step 5 are separate. Between reading βA5 is availableβ and writing βA5 is heldβ, any number of other requests can read the same row and reach the same conclusion. Every one of them then writes, and the last write wins, so several users each get a confirmed hold on the same seat and the show sells more tickets than it has seats. On a hot release, a thousand people press the button in the same second and this is not an unlikely race but the expected outcome. That is Deep Dive 1.
- The deadline lives in a process. A timer in memory is lost when that process restarts, is deployed, or crashes β and no other instance knows the hold existed. Those seats then stay
HELDforever: never paid for, never released, invisible to everyone. A routine deploy silently destroys inventory. That is Deep Dive 2.
FR3: Complete Payment and Confirm Booking
After seats are held, the user has 10 minutes to complete payment. Payment can fail (insufficient funds, gateway timeout, OTP expired). We need to handle all failure modes without losing the booking or double-charging.
New components we need:
- Payment Service - takes the amount and the booking, calls the gateway, and reports what happened.
- Payment Gateway (external) - Razorpay, Stripe or similar. Moves the actual money. We do not control it, cannot see inside it, and cannot include it in our transaction.
flowchart LR
User["User Browser"]:::client
BS["Booking Service"]:::service
PS["Payment Service"]:::service
PGW["Payment Gateway<br/>external"]:::external
DB[("Postgres<br/>seats and bookings")]:::data
User -->|"1. Pay for the booking"| BS
BS -->|"2. Check the hold is live"| DB
BS -->|"3. Charge this amount"| PS
PS -->|"4. Take the money"| PGW
PGW -->|"5. Approved"| PS
BS -->|"6. Mark booked and confirmed"| DB
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Step-by-step flow:
- User taps Pay while the countdown from FR2 is still running
- The Booking Service confirms the booking is still
HELDandexpires_athas not passed - It calls the Payment Service with the amount and the booking id
- The Payment Service calls the gateway and waits for it to come back
- On approval, control returns to the Booking Service
- It sets the booking to
CONFIRMEDand the three seats toBOOKED, and the ticket is issued - On decline it leaves the hold alone, so the user can try another card inside their remaining time
Why check the hold again in step 2 rather than trusting the countdown on screen? The timer
in the userβs browser is a decoration. The authority is expires_at, and the ten minutes may
have elapsed while the payment screen sat open, so the only safe moment to test it is on the
server immediately before spending the userβs money.
What we have deliberately left broken. The happy path is right and there is a hole in the middle of it:
- Step 5 and step 6 are not atomic, and they cannot be. The gateway is a different company across a network; it cannot join a Postgres transaction. So there is a window between the money moving and our database recording it, and if the process dies in that window the customer has been charged and owns no ticket. Nothing in this design will ever notice, because the only thing that knew a payment was in flight was the request that died.
- A retry charges twice. If the gateway call times out at step 4 we genuinely do not know whether the money moved. Retrying risks a second charge; not retrying risks a paid customer with no seat. As written we have no way to ask βdid this already happen?β.
- An expired hold plus a successful payment has no resolution. If the hold lapsed while the user was on the gatewayβs page and someone else took the seats, step 6 cannot succeed β and the money is already gone with no path to send it back.
All three are Deep Dive 3.
10. Technology Choices
| Tier | Purpose | Stores | Access Pattern | Primary | Alternatives |
|---|---|---|---|---|---|
| Booking DB | Booking lifecycle state | Bookings, seat assignments, payment refs | Read/write by bookingId, showId | Postgres | CockroachDB, TiDB |
| Seat Lock Store | Temporary seat holds with TTL | seatId β userId + expiryTime | Atomic CAS, TTL-based expiry | Redis Cluster | DynamoDB (conditional writes) |
| Event Catalog | Movies, shows, cinemas, schedules | Catalog metadata | Read-heavy, filtered by city/date | Postgres (read replicas) | MongoDB |
| Event Bus | Async events (booking confirmed, seat released) | Booking lifecycle events | Pub/sub per show | Kafka or Redpanda | Kinesis, RabbitMQ |
| Cache | Show listings, seat availability snapshots | Aggregated seat counts, catalog data | High-QPS reads, short TTL | Redis Cluster | Memcached |
| Queue | Booking requests during peak (virtual waiting room) | User positions, pending requests | FIFO with priorities | SQS or Kafka | RabbitMQ, Redis Streams |
| Payment Gateway | External payment processing | Transactions | Request-response with webhooks | Razorpay or Stripe | PayU, PayTM |
Why Redis for seat locks, not Postgres row-level locks?
During a hot event, 100K users hit the system in 10 seconds. Each seat lock attempt is a SET seatId NX EX 600 (atomic check-and-set with 10-min TTL). Redis handles 300K+ ops/sec per shard in-memory. Postgres row-level locks would create massive lock contention, connection pool exhaustion, and deadlocks. Redis gives us sub-millisecond lock acquisition with built-in TTL for auto-release.
Why Kafka for booking events? A single booking triggers 5+ downstream actions: send confirmation email, update seat map cache, record analytics, trigger invoice generation, update show occupancy counter. Kafkaβs consumer groups let each downstream service process independently without blocking the booking path.
11. Data Modeling
Postgres (Booking DB β booking lifecycle):
CREATE TABLE bookings (
booking_id UUID PRIMARY KEY,
user_id UUID NOT NULL,
show_id UUID NOT NULL,
status VARCHAR(15) NOT NULL, -- INITIATED, HELD, CONFIRMED, CANCELLED, EXPIRED
total_amount DECIMAL(10,2),
payment_ref VARCHAR(128),
idempotency_key UUID UNIQUE,
created_at TIMESTAMP NOT NULL,
confirmed_at TIMESTAMP
);
CREATE INDEX idx_bookings_user ON bookings(user_id, created_at DESC);
CREATE INDEX idx_bookings_show ON bookings(show_id, status);
Redis (Seat Lock Store β temporary holds with TTL):
Key: "seat:hold:{showId}:{seatId}" β Hash { userId, bookingId, expiresAt }
SET via: SET seat:hold:{showId}:{seatId} NX EX 600 (atomic, 10-min TTL)
Postgres (Seat Availability β source of truth after payment):
CREATE TABLE show_seats (
show_id UUID NOT NULL,
seat_id VARCHAR(10) NOT NULL, -- e.g., "A1", "B12"
category VARCHAR(10) NOT NULL, -- SILVER, GOLD, PLATINUM
price DECIMAL(8,2) NOT NULL,
status VARCHAR(10) NOT NULL DEFAULT 'AVAILABLE', -- AVAILABLE, BOOKED
booked_by UUID,
PRIMARY KEY (show_id, seat_id)
);
Access Patterns:
| Query | Data Source | How |
|---|---|---|
| Show available seats | Redis (cache) + Postgres | Cache seat map per show with 5s TTL; fallback to DB |
| Hold seats (atomic) | Redis | SET seat:hold:{showId}:{seatId} NX EX 600 per seat β all-or-nothing in Lua script |
| Confirm booking | Postgres | In txn: UPDATE show_seats SET status='BOOKED' + UPDATE bookings SET status='CONFIRMED' |
| Auto-release expired holds | Redis TTL | Keys auto-expire after 600s; no sweeper needed |
| Check if seat is held | Redis | EXISTS seat:hold:{showId}:{seatId} |
How Seat Locking Prevents Double-Booking During a Hot Event (100K concurrent users):
- User selects seats A1, A2, A3 β API sends
POST /bookings/hold - Server executes a Redis Lua script (atomic, no race conditions):
for each seatId in [A1, A2, A3]: result = SET seat:hold:{showId}:{seatId} {userId} NX EX 600 if result == nil: ROLLBACK all previous SETs, return CONFLICT - All-or-nothing: if ANY seat is already held, none are held (no partial locks)
- On success: create booking in Postgres with status=HELD, return booking_id + payment URL
- User has 10 minutes to pay. On payment success β update Postgres
show_seatsto BOOKED - If user abandons β Redis keys auto-expire after 10 min, seats become available again
- No background sweeper needed for Redis path β TTL handles cleanup automatically
12. Deep Dives
1) A thousand people tap the same seat in the same second. How does exactly one get it?
Problem: 1000 users click βHold Seat A5β within the same second for a hot show. Exactly one must succeed; 999 must fail cleanly.
Bad: What FR2 built β read the seat rows, check they say AVAILABLE, then write HELD.
The check and the write are two separate statements, and nothing stops other requests running
between them. Every concurrent request reads the same AVAILABLE, every one concludes it may
proceed, and every one writes. The last write decides who the row says owns the seat, while
every one of those users has already been told their hold succeeded. The result is not a near
miss: on a hot release with a thousand taps in one second, the expected outcome is that we sell
one seat many times and find out at the door. Adding a re-read after the write does not fix it
either, because two requests can both re-read after both have written.
Good: Postgres row-level lock. SELECT ... FOR UPDATE on the seat row, then update status. Works for moderate concurrency, but under 1000 concurrent transactions, you get lock contention, connection pool starvation, and 5+ second response times. Postgres serializes all competing transactions - they queue up behind each other.
Great: Redis SET NX EX (atomic conditional set with TTL).
π‘ SET NX = βSet only if Not eXistsβ - an atomic compare-and-swap in one round trip. Combined with EX (expire), it gives us a lock that auto-releases.
flowchart LR
LM["Lock Manager"]:::service
R1["Redis Shard 1"]:::data
R2["Redis Shard 2"]:::data
R3["Redis Shard 3"]:::data
BS["Booking Service"]:::service
BS -->|"1. Acquire seats"| LM
LM -->|"2. Show:1:seat:A*"| R1
LM -->|"3. Show:1:seat:B*"| R2
LM -->|"4. Show:2:seat:*"| R3
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Mechanism:
- Lock key format:
show:{showId}:seat:{seatId}β value:{bookingId}:{userId}:{timestamp} - Lock acquisition (Lua script for multi-seat atomicity):
-- Atomic multi-seat lock for each seatKey in requested_seats: result = SET seatKey bookingId NX EX 600 if result == nil: -- someone else holds it -- rollback: delete all keys we just set for each acquired_key: DEL acquired_key return FAILED + which seat was taken return SUCCESS - NX guarantees only one caller succeeds for each seat - Redis is single-threaded per shard
- EX 600 (10 minutes) ensures auto-release if payment isnβt completed
- Sharding: keys are sharded by showId across Redis cluster nodes. One hot show maps to one shard - that shard handles all lock contention for that show
Why not Redlock?
Redlock (consensus across N Redis nodes) adds 3-5ms latency and complexity. For seat booking, a single Redis shard with persistence (AOF every second) is sufficient. The worst case of a Redis crash is: users who had holds lose them (10-minute window restarts). This is acceptable - they can re-select and re-hold. We donβt need the durability guarantees of Redlock.
Backstop: Postgres as final gate
Even though Redis handles the fast path, the Booking Service writes the confirmed booking to Postgres with a unique constraint: UNIQUE(showId, seatId) WHERE status = 'CONFIRMED'. If Redis somehow fails (split-brain, data loss), Postgres prevents actual double-booking at the persistence layer.
2) A deploy restarts the server mid-hold. Who releases those seats?
Problem: User holds seats, goes to make tea, never pays. Those seats must become available again exactly at TTL expiry. But Redis TTL deletion is lazy (checked on access or via periodic sampling) - thereβs no guarantee of exact-millisecond release.
Bad: What FR2 built β an in-process timer on whichever instance served the hold. The
timerβs lifetime is the processβs lifetime, so a restart, a deploy, an autoscaler scale-in or a
crash takes it with them, and no other instance has any idea the hold existed. Those seats sit
HELD permanently: not paid for, not purchasable, gone from inventory with no record of why. A
routine deploy therefore quietly destroys seats on every show with a hold open at that moment,
which on a busy evening is most of them. The damage also accumulates silently, because nothing
in the system is looking for holds that should have expired.
Good: Rely purely on Redis TTL. When another user queries seat status, the key is gone (expired), so they can lock it. Works for the lock itself, but: the booking record in Postgres still says βHELDβ - inconsistency.
Great: Redis TTL for lock release + background reconciler for state consistency + Kafka event for downstream notifications.
In simple terms: When you select a seat, we lock it for 10 minutes using Redis with an auto-expiry timer. If you donβt pay in time, the lock disappears automatically and the seat becomes available again. A background job double-checks that Redis and the database agree.
Mechanism:
- Redis TTL (primary): The lock key expires after exactly 600 seconds. After expiry, any new SET NX on that key will succeed - seat is available for others.
- Reconciler (consistency): Runs every 30 seconds. Queries Postgres:
SELECT * FROM bookings WHERE status = 'HELD' AND expires_at < NOW(). For each:- Updates status to
EXPIRED - Publishes
booking.expiredevent to Kafka - Kafka consumers: notify user (βYour hold expiredβ), update analytics
- Updates status to
- Redis Keyspace Notifications (optional enhancement): Subscribe to
__keyevent@0__:expiredevents. When a lock key expires, immediately trigger the Postgres update instead of waiting for the 30-second reconciler sweep.
Edge case - race between payment and expiry:
User pays at minute 9:58 (2 seconds before expiry). Payment gateway takes 5 seconds to process. By the time we get the success response, the Redis TTL has expired and someone else might have locked the seat.
Solution: Before calling the payment gateway, extend the Redis lock by 2 minutes (safety buffer): EXPIRE show:{showId}:seat:{seatId} 720. This gives us time to process the payment response. If payment fails, we explicitly DEL the lock.
3) The money moved and then our server died. How does the customer get their ticket?
Problem: Booking involves two systems: our seat lock (Redis + Postgres) and an external payment gateway. These canβt be in a single database transaction. What if payment succeeds but our server crashes before confirming the booking? What if we confirm the booking but payment actually failed (gateway network error)?
Bad: What FR3 built β call the gateway inline, then write CONFIRMED when it returns. The
two steps are a charge at another company and a write in our database, and there is no way to
make them one operation. If the process dies in between, the customer has paid and has no
ticket, and nothing is left that knows to look. A timed-out gateway call is the same wound from
the other side: we cannot tell whether the money moved, so retrying may charge twice and not
retrying may strand a payment.
The instinctive fix is a distributed transaction across both β two-phase commit. It is not available: no public payment gateway exposes a 2PC participant interface, so there is nothing to enlist. Even between systems we control, 2PC holds locks for the duration of the round trip and leaves participants blocked indefinitely if the coordinator dies mid-commit, which is a poor trade for a ten-minute human checkout. The gap has to be closed by making each step individually recoverable rather than by pretending the two are one.
Good: Optimistic approach - assume payment will succeed, confirm booking first, then process payment. If payment fails, roll back the booking. Problem: user already has a βconfirmedβ ticket for a brief moment (bad UX and potential fraud vector).
Great: Saga pattern with explicit compensating transactions. (Borrowing from Stripeβs idempotency key pattern and Razorpayβs orchestration.)
In simple terms: Booking involves multiple steps (lock seat β charge card β confirm). If step 2 fails (card declined), we automatically βundoβ step 1 (release the seat). Each step has a pre-defined rollback action.
flowchart LR
BS["Booking Service Orchestrator"]:::service
PS["Payment Service"]:::service
RL["Redis Locks"]:::data
DB["Postgres"]:::data
PGW["Payment Gateway"]:::external
KF["Kafka"]:::async
BS -->|"1. Verify hold"| RL
BS -->|"2. Initiate payment"| PS
PS -->|"3. Charge"| PGW
PGW -->|"4. Success/Fail"| PS
PS -->|"5. Result"| BS
BS -->|"6. Confirm or rollback"| DB
BS -->|"7. Event"| KF
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0
classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
Saga steps:
| Step | Action | Compensating Transaction |
|---|---|---|
| 1 | Verify hold is still valid (Redis lock exists) | - (read-only check) |
| 2 | Extend lock TTL by 2 min (safety buffer) | Restore original TTL |
| 3 | Call payment gateway with idempotencyKey | Refund payment |
| 4 | On success: update booking to CONFIRMED | - |
| 5 | On failure: release lock explicitly | - |
Idempotency key lifecycle:
- Client generates a UUID v4 as idempotencyKey before clicking Pay
- Server stores
{idempotencyKey: bookingId, status: PROCESSING}in Redis with 24hr TTL - If the same request arrives again (retry), check Redis: if status = PROCESSING β return βstill processingβ; if status = SUCCESS β return cached response; if status = FAILED β allow retry with same key
- Gateway uses the same key: same charge is never processed twice
Handling βzombie paymentsβ (success webhook arrives after hold expired):
- Payment Service receives success webhook for a booking thatβs already EXPIRED
- Immediately triggers refund via gateway API
- Records this as an auto-refunded transaction
- Alerts ops dashboard (unusual but not a bug)
4) 500K people arrive in ten seconds for one premiere. Who gets served and who waits?
Problem: Avengers premiere - 500K users hit the booking page in 10 seconds. The Redis lock store for that show gets 500K SET NX attempts/sec. Even Redis will struggle, and the API servers will be overwhelmed.
Bad: What the design so far does β nothing. Every one of the 500K arrivals is admitted straight to the seat map and the hold endpoint. FR1βs reads alone are 50K queries/sec against one showβs few hundred rows, and FR2βs holds all contend on those same rows, so the database saturates and requests start timing out. Then it compounds: every timed-out client retries, so load rises after the failure, and the system stays down for minutes rather than seconds. What the users experience is worse than slowness β it is arbitrary. Whoever has the lowest latency and the most aggressive retry loop wins the seats, so the outcome is decided by network proximity and client behaviour rather than by who arrived first.
Good: Rate limit at the API gateway (500 requests/sec). Most users get rejected. Better for the system, but terrible UX - βtry again laterβ for 499K users.
Great: Virtual waiting room with fair queue. (Borrowing from Ticketmasterβs approach.)
In simple terms: During a hot event (Avengers premiere), instead of letting 500K people hammer the system simultaneously, we put them in a virtual queue. We let them through in batches of 1000, preventing the system from crashing while keeping it fair (first come, first served).
flowchart LR
Users["500K Users"]:::client
WR["Waiting Room Service"]:::edge
Q["Priority Queue SQS"]:::async
BS["Booking Service"]:::service
RL["Redis Locks"]:::data
Users -->|"1. Enter waiting room"| WR
WR -->|"2. Assign position"| Q
Q -->|"3. Dequeue 100/sec"| BS
BS -->|"4. Acquire seat lock"| RL
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Mechanism:
- Trigger: When a showβs booking page hits > 10K concurrent viewers (tracked via WebSocket connections or session counter), activate the waiting room for that show.
- Assign position: Each user arriving at the booking page gets a random position in the queue (not first-come-first-served - avoids bot advantage). Position = hash(userId + salt + timestamp_bucket).
- Drip processing: Waiting Room Service dequeues users at a controlled rate (100-500 users/sec) based on the downstream systemβs capacity.
- User experience: User sees βYouβre #4,521 in line. Estimated wait: 2 minutes.β Position updates in real-time via SSE.
- When itβs their turn: User gets a time-limited token (2 minutes validity) that allows them to access the seat selection page. Token is validated at the API Gateway.
- Overflow: If all seats are sold while users are in queue, remaining queue members are notified βSold Outβ and queue is drained.
Why random position instead of arrival order?
Arrival-order queues reward bots and users with faster network connections. Random assignment is fairer and eliminates the incentive to DDoS the system at second zero.
Capacity math: A cinema hall has 300 seats. Even if all 300 are booked in one go (unlikely - most users book 2-4 seats), we only need ~100 successful booking attempts. Processing 500 users/sec means the entire queue is served in ~17 minutes for a 500K-user queue. With 80% dropping off or failing, actual booking completes in 3-5 minutes.
5) Someone takes A5 while a thousand people are looking at it. How do their screens find out?
Problem: 1000 users are viewing the same seat map. When User A holds seat A5, all other users should see it turn yellow (held) within 2-3 seconds. Otherwise, theyβll select the same seat and get frustrated when their hold fails.
Bad: What FR1 leaves the client to do β re-request the whole seat map on a timer. At 1000 concurrent viewers polling every 2 seconds that is 500 requests/sec for a single show, each one returning a few hundred seats where typically nothing has changed, and it multiplies by every show running concurrently. The waste is the smaller complaint. A 2-second interval means a seat can be taken and up to 2 seconds pass before anyone elseβs screen shows it, which is exactly the window in which someone picks it, submits, and is told no. Shortening the interval to close that window multiplies a request rate that is already almost entirely redundant.
Good: Short polling with aggressive caching. Cache the seat map in Redis with 3-second TTL. 500 requests/sec all hit cache. Works, but users still see stale data for up to 3 seconds.
Great: Server-Sent Events (SSE) for seat status push + event-driven invalidation.
Mechanism:
- When a user opens the seat map page, browser opens an SSE connection:
GET /sse/v1/shows/{showId}/seats(long-lived HTTP connection withtext/event-stream) - Gateway registers this connection to a show-specific channel
- When any seatβs status changes (held, released, booked), the Booking Service publishes to Kafka topic
show.{showId}.seat-updates - A Seat Update Consumer reads from Kafka and pushes to all SSE connections for that show via Redis Pub/Sub β SSE Gateway
- Client receives:
data: {"seatId": "A5", "status": "held", "heldBy": "someone"}and updates the seat map UI immediately
Why SSE instead of WebSocket?
Seat map updates are serverβclient only (users donβt send data back over this connection). SSE is simpler, works over HTTP/2 multiplexing, auto-reconnects, and requires no upgrade handshake. WebSocket would be overkill for a one-directional push.
Scaling: With SSE over HTTP/2, a single gateway can hold 100K+ connections (far cheaper than WebSocket). For hot shows with 50K concurrent viewers, 1 gateway instance suffices for the seat map push path.
13. Design Self-Audit
| Question | Answer |
|---|---|
| Dedicated search index? | Not needed for core booking flow. Movie/event discovery can use Postgres full-text search or Elasticsearch for advanced filtering (genre, language, nearby cinemas). Low priority - not on the critical booking path. |
| Stale reads after writes? | Seat map has 2-3 second lag via SSE push (acceptable). After YOUR hold succeeds, your own UI updates immediately (optimistic). Other users may attempt the same seat and fail - clean error handling. |
| Single points of failure? | Redis lock store uses Redis Cluster (3+ shards, each with a replica). Booking Service is stateless, horizontally scaled. Postgres uses primary + synchronous standby for booking confirmations. |
| Dead-letter / reconciliation? | Reconciler every 30s cleans expired holds. Failed payment webhooks go to DLQ with exponential retry (1min, 5min, 30min). Orphaned bookings (INITIATED for > 15 min) are auto-cancelled. |
| Data freshness across caches? | Catalog cache TTL = 60s (acceptable for movie listings). Seat map is real-time via SSE. Aggregated availability (β45 seats leftβ) is eventually consistent (5s lag). |
| Cost at scale? | Redis Cluster (6 nodes for locks): ~$2000/month. Kafka (3 brokers): ~$1500/month. API + Booking Service (20 instances): ~$4000/month. Postgres RDS (primary + standby): ~$2000/month. Total: ~$10K/month for 1M bookings/day platform. |
14. Core Flows
Flow 1: Seat Hold with Concurrent Competition
sequenceDiagram
participant U1 as User 1
participant U2 as User 2
participant GW as API Gateway
participant BS as Booking Service
participant LM as Lock Manager
participant Redis as Redis Lock Store
Note over U1,U2: Both want seat A5 for same show
U1->>GW: POST /bookings/hold (seats: [A5, A6])
U2->>GW: POST /bookings/hold (seats: [A5, A7])
GW->>BS: User1 hold request
GW->>BS: User2 hold request
BS->>LM: Lock [A5, A6] for User1
BS->>LM: Lock [A5, A7] for User2
LM->>Redis: SET show:1:seat:A5 booking1 NX EX 600
Redis-->>LM: OK (User1 wins)
LM->>Redis: SET show:1:seat:A6 booking1 NX EX 600
Redis-->>LM: OK
LM-->>BS: All locks acquired for User1
LM->>Redis: SET show:1:seat:A5 booking2 NX EX 600
Redis-->>LM: nil (FAILED - already locked)
LM-->>BS: Lock failed for User2 on seat A5
BS-->>U1: 200 OK - Seats held. Pay within 10 min.
BS-->>U2: 409 Conflict - Seat A5 unavailable
Non-obvious failure path: What if the Booking Service crashes after acquiring locks in Redis but before writing the booking to Postgres? The Redis locks have a 10-minute TTL - theyβll auto-expire. The reconciler wonβt find a matching booking in Postgres (it was never written), so no cleanup needed. The seats become available again after TTL expires. Worst case: 10 minutes of phantom unavailability for those seats.
Flow 2: Payment Completion with Failure Handling
sequenceDiagram
participant User
participant BS as Booking Service
participant PS as Payment Service
participant PGW as Payment Gateway
participant Redis as Redis Locks
participant DB as Postgres
User->>BS: POST /bookings/{id}/pay
BS->>Redis: Check lock exists for bookingId
Redis-->>BS: Lock valid (5 min remaining)
BS->>PS: Process payment (amount + idempotencyKey)
PS->>PGW: Create payment intent
PGW-->>PS: Payment processing...
alt Payment succeeds
PGW-->>PS: Success + transactionId
PS-->>BS: Payment confirmed
BS->>DB: Update booking: CONFIRMED
BS->>Redis: PERSIST lock (remove TTL)
BS-->>User: 200 OK - Booking confirmed + ticket
else Payment fails (insufficient funds)
PGW-->>PS: Failed: insufficient funds
PS-->>BS: Payment failed
BS-->>User: 402 - Payment failed. Retry with another method.
Note over Redis: Hold still valid. User can retry.
else Hold expired during payment
PGW-->>PS: Success + transactionId
PS->>BS: Payment confirmed
BS->>Redis: Check lock exists
Redis-->>BS: Lock expired (nil)
BS->>PS: Trigger refund
PS->>PGW: Refund transactionId
BS-->>User: 410 Gone - Hold expired. Refund initiated.
end
Non-obvious failure path: Payment gateway sends a success webhook but our server crashes before processing it. The gateway will retry the webhook (typically 3-5 times over 24 hours). Payment Service uses the idempotencyKey to detect duplicate webhooks and skip re-processing. If all webhook retries fail, a reconciler job polls the gateway every 5 minutes for recent payments and matches them against pending bookings.
Booking Lifecycle State Machine
stateDiagram-v2
[*] --> INITIATED : User selects seats
INITIATED --> HELD : All locks acquired
INITIATED --> FAILED : Lock conflict (seats taken)
HELD --> CONFIRMED : Payment success
HELD --> EXPIRED : TTL expired (10 min)
HELD --> CANCELLED : User cancels
CONFIRMED --> REFUNDED : Admin refund or cancellation policy
EXPIRED --> [*]
FAILED --> [*]
CONFIRMED --> [*]
CANCELLED --> [*]
REFUNDED --> [*]
Each transition emits a Kafka event consumed by: Notification Service (user updates), Seat Map Cache (invalidation), Analytics (conversion tracking), and the Reconciler (consistency checks).
15. Final Architecture
flowchart TB
UA(["User Browser and App"]):::client
GW["API Gateway<br>auth and rate limiting"]:::edge
WR["Waiting Room<br>admits a trickle"]:::edge
SSE["SSE Gateway<br>seat map updates"]:::edge
CAT["Catalog Service<br>shows and seat maps"]:::service
SS["Seat Service<br>holds and releases seats"]:::service
BS["Booking Service<br>owns booking lifecycle"]:::service
PS["Payment Service"]:::service
NS["Notification Service"]:::service
KF[["Kafka<br>booking events"]]:::async
PG[("Postgres<br>bookings and seats")]:::data
RD[("Redis<br>seat locks and cache")]:::data
PGW[/"Payment Gateway"/]:::external
UA -->|"browse and book"| GW
GW -->|"peak traffic"| WR
WR --> CAT
GW --> CAT
CAT --> RD
GW --> SS
SS -->|"lock the seat"| RD
SS --> BS
BS --> PG
BS --> PS
PS --> PGW
BS -->|"booking confirmed"| KF
KF --> NS
KF --> SSE
SSE -->|"seats just went"| UA
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0
How it works end-to-end:
- User sends request β hits the Load Balancer, routed through API Gateway
- Waiting Room queues during surge β hot event on-sales throttle via the Booking Queue
- Catalog Service serves seat map β reads from Redis Cache (or Postgres on miss)
- Seat Service checks availability β queries Redis Lock Store for real-time seat state
- Booking Service acquires lock β Lock Manager does a Redis SET NX with TTL on selected seats
- Payment is processed β Payment Service calls the external Payment Gateway with idempotency key
- Booking confirmed and persisted β writes to Postgres, emits event to Kafka
- Downstream consumers react β Confirmation Service generates ticket, Notification Service sends email/push, SSE Gateway updates live seat maps
- Reconciler handles orphans β releases expired locks and syncs Redis with Postgres
Want a deep dive on multi-cinema franchise inventory aggregation, dynamic pricing (surge for hot shows), or fraud detection (bot-booking prevention)? Drop a comment below π
Key Technologies
| Term | What it is |
|---|---|
| Redis SET NX (distributed lock) | Atomic βset if not existsβ command used to acquire a seat lock - only one caller succeeds, preventing double-booking. |
| TTL-based hold | Lock keys expire automatically after 10 minutes, releasing seats if payment isnβt completed. |
| CQRS | Command Query Responsibility Segregation - separating the write path (seat locks, bookings) from read path (seat map browsing) for independent scaling. |
| Kafka | Event bus carrying booking lifecycle events to downstream services (notifications, analytics, seat map invalidation). |
| WebSocket | Persistent connection for pushing real-time seat availability updates to users viewing the same show. |
| Postgres | ACID-compliant relational DB serving as the booking source of truth with unique constraints as a double-booking backstop. |
| Idempotency Key | Client-generated UUID ensuring payment retries never double-charge - gateway returns the same result for repeated keys. |
Whatβs Expected at Each Level
This section helps you calibrate your depth. You donβt need to cover everything - just know whatβs expected for your level.
Mid-level
Design a basic seat selection and booking system with a database. Recognize the concurrency problem - two users selecting the same seat simultaneously. Propose a locking mechanism with prompting. You should articulate why naive check-then-act creates race conditions and sketch a happy-path flow from seat selection through payment.
Senior
Propose Redis SET NX with TTL for seat holds. Explain atomic multi-seat locking via a Lua script (all-or-nothing semantics). Discuss the payment saga pattern and idempotency without prompting. Recognize the need for a hold expiry reconciler to keep Postgres consistent with Redis TTL state. Articulate why Postgres row-level locks fail under 100K concurrent users.
Staff+
Address the thundering herd problem on hot events by proposing a virtual waiting room or queue-based admission control. Discuss Postgres as a backstop with unique constraints even when Redis is the fast path. Proactively mention how CDN invalidation works for seat maps, zombie payment handling (success webhook after hold expiry), and the cost trade-off of keeping expired holds in Redis vs background cleanup. Quantify the peak load numbers and explain capacity planning.
π― Key Takeaways
- Redis SET NX EX gives atomic seat locking with auto-release via TTL
- All-or-nothing Lua script prevents partial seat holds
- Saga pattern with idempotency keys handles payment failures safely
- Virtual waiting room protects the system during hot event on-sales
Related Designs
- Stock Broker (Robinhood) - exactly-once processing, order matching
- Digital Wallet (PhonePe) - payment orchestration, saga pattern, idempotency
- Job Scheduler - TTL expiry management, delayed triggers
Related Concepts
Understand the building blocks used in this design:
- Distributed Locking β β holds seats during checkout so two users canβt book the same seat
- Saga Pattern β β coordinates seat hold, payment, and confirmation with compensating rollbacks
- Idempotency β β retried payment callbacks never double-charge a booking
- Database Replication β β keeps seat inventory durable and available across nodes
Discussion
Newest first