Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
โฑ๏ธ 42 min read

Designing an E-Commerce Platform

Difficulty: Advanced Topics: Inventory Consistency, Search, Order Orchestration, Idempotency Asked at: Amazon, Flipkart, Walmart, Meesho Prerequisites:Idempotency and Transactions and Isolation


1. Understanding the Problem

An e-commerce platform lets people find a product among a hundred million of them, put it in a cart, pay, and receive it. The browsing half is a read-scaling problem and it is largely solved by caching. The buying half is where the design earns its difficulty rating, because it has to be correct about two things that cannot be approximated: how many units exist, and how much money changed hands.

The one sentence to say out loud early: you can sell a unit you do not have exactly zero times. Everything structural in this design follows from that. Overselling one PlayStation during a flash sale is not a rounding error you reconcile later, it is a cancelled order, a refund, a support ticket and a customer who buys from someone else next time.

This page is the end-to-end platform. Two neighbouring designs go deeper on their own slice and are not repeated here: Cart Management covers cart persistence, guest-cart merging and cross-device sync, and Payment System covers the payment rails, PSP integration and settlement. Read them alongside this.

Real examples: Amazon, Flipkart, Walmart, Meesho, Shopify, Zalando.


2. Naive First Cut

flowchart LR
    USER["Shopper"]:::client
    APP["Monolith<br/>catalog cart orders"]:::service
    DB[("Postgres<br/>products stock orders")]:::data

    USER -->|"1. Browse and search"| APP
    APP -->|"2. SELECT and UPDATE"| DB
    USER -->|"3. Place order"| APP

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Color Meaning
๐ŸŸฃ Purple Clients
๐ŸŸข Green Services
๐ŸŸก Yellow Data stores
๐Ÿ”ต Blue Edge / CDN

One application, one database. Products in a table, a stock integer on each row, orders in another table. Search is WHERE title LIKE '%query%'. Checkout reads the stock, checks it in application code, and writes the decremented value.

Why this breaks:

The rest of this doc evolves that into a platform that cannot oversell, searches properly, and survives a partial failure mid-checkout.


3. Prior Art Weโ€™re Drawing From


4. Functional Requirements

Core (Top 3)

  1. Browse and search the catalog - find a product among 100M by text query, filter by brand, price and category
  2. Place an order without overselling - a unit is sold at most once, ever, under any amount of concurrency
  3. Track order state - the buyer sees where their order is from placement through to delivery

Below the Line


5. Non-Functional Requirements

Core

NFR Target
No oversell Hard invariant. Units sold never exceeds units held. Not eventually - never.
Search latency P95 under 200 ms at 50K queries/sec
Order durability A placed order is never lost, even if a downstream service is down
Catalog read scale 1000:1 read-to-write, served mostly from cache and CDN

Below the Line


6. Scale Estimation (Back-of-Envelope)

The shape: an enormous cacheable read path, a modest but strictly correct write path, and one adversarial spike pattern that breaks the modest write path if you do not plan for it.


7. Core Entities


8. API / System Interface

GET /v1/search?q=running+shoes&brand=nike&maxPrice=5000&page=1
  Response: { total, facets: {...}, results: [{ productId, title, price, imageUrl, rating }] }

GET /v1/products/{productId}
  Response: { productId, title, description, attributes, skus: [{ skuId, size, price, inStock }] }

POST /v1/cart/items
  Body: { skuId, quantity }
  Response: { cartId, items: [...], subtotal }

POST /v1/orders
  Headers: Idempotency-Key: <client-generated uuid>
  Body: { cartId, addressId, paymentMethodId }
  Response: { orderId, state: "PAID", total, estimatedDelivery }

GET /v1/orders/{orderId}
  Response: { orderId, state, items: [...], total, timeline: [{ state, at }] }

Security note: The buyerโ€™s identity comes from the session, never from the request body - an endpoint that accepts a userId lets anyone read anyoneโ€™s orders. More subtly, price must never be accepted from the client. The server reads the current price at order time and captures it on the order row. A checkout that trusts a client-supplied total is a free-shopping bug, and it is a real one that ships regularly.


9. High-Level Design

Three requirements, three passes. The third one is the whole design.

FR1: Browse the Catalog

100M products with wildly different shapes, read a thousand times for every write.

Why not one relational table. A products table needs columns for every attribute any category might have. Books need isbn and page_count, shoes need size and width, televisions need panel_type and refresh_rate. You end up with either hundreds of nullable columns or an entity-attribute-value side table that turns every product page render into a pivot across dozens of rows. Neither is pleasant, and neither is necessary: a product is naturally a document, and nothing about reading it needs a join.

New components:

  1. Catalog Service - owns product documents. Validates seller writes, serves product reads.
  2. Catalog Store (document DB) - one document per product, with category-specific attributes nested inside it. Read by primary key, which is the only access pattern the product page needs.
    ๐Ÿ’ก A document store here is not about โ€œunstructured dataโ€ - the data is perfectly structured. It is about the structure differing per category, and about every read being a single-key lookup that a join would only slow down.
  3. CDN plus cache - product pages and images are near-static. At 1000:1 read-to-write, almost every read should be answered before it reaches the Catalog Service.
flowchart LR
    USER["Shopper"]:::client
    CDN["CDN<br/>product pages and images"]:::edge
    CAT["Catalog Service"]:::service
    CDB[("Catalog Store<br/>product documents")]:::data
    SELLER["Seller"]:::client

    USER -->|"1. View product"| CDN
    CDN -->|"2. Forward on miss"| CAT
    CAT -->|"3. Read product by id"| CDB
    SELLER -->|"4. Update product"| CAT

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Step-by-step flow:

  1. Shopper opens a product page. The request hits a CDN edge.
  2. On a hit, the page is served in under 30 ms and nothing downstream is touched.
  3. On a miss, the Catalog Service reads the document by primary key and returns it with cache headers.
  4. A seller updating a price writes through the Catalog Service, which invalidates the cached page for that product.

Note what is deliberately not on the product page: a live stock count. Rendering โ€œ3 left in stockโ€ requires a read of the inventory tier on every page view, which couples your cheapest path to your most contended one. Show a boolean inStock from a cached value instead and let the truth be established at checkout, where it has to be anyway.


FR2: Search the Catalog

50K queries/sec with relevance, facets and typo tolerance. The catalog store cannot do this and neither can Postgres.

Why a database cannot serve search. Three separate reasons, and all three are fatal. A leading-wildcard LIKE cannot use a B-tree index, so every query scans. There is no notion of relevance - matching โ€œnike running shoesโ€ should rank a Nike running shoe above a shoe-cleaning kit that happens to mention Nike, and LIKE has no opinion. And a shopper typing โ€œnkieโ€ expects results, which requires fuzzy matching the database has no mechanism for.

New components:

  1. Search Service - translates a shopperโ€™s query and filters into an index query, applies business rules (hide out-of-stock sellers, boost sponsored results).
  2. Search Index (Elasticsearch) - an inverted index over the catalog with an analysis pipeline for tokenising, lowercasing and stemming, plus faceting and relevance scoring. See search indexing for how the inverted index and BM25 ranking actually work.
  3. CDC Pipeline - streams catalog changes into the index asynchronously.
flowchart LR
    USER["Shopper"]:::client
    SS["Search Service"]:::service
    ES[("Search Index<br/>inverted index")]:::data
    CAT["Catalog Service"]:::service
    CDB[("Catalog Store<br/>product documents")]:::data
    CDC["CDC Pipeline"]:::async

    USER -->|"1. Search query"| SS
    SS -->|"2. Query index"| ES
    CAT -->|"3. Write product"| CDB
    CDB -->|"4. Stream changes"| CDC
    CDC -->|"5. Index document"| ES

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Step-by-step flow:

  1. Shopper searches โ€œrunning shoesโ€ with a brand filter.
  2. Search Service builds an index query with the filter as a facet constraint and returns ranked results plus facet counts.
  3. Separately, a seller updates a productโ€™s price through the Catalog Service.
  4. The write lands in the Catalog Store, and the CDC pipeline reads it from the storeโ€™s own change log.
  5. The pipeline transforms the change into an index document and writes it to Elasticsearch.

This makes search eventually consistent, and that is a product decision, not an accident. The lag from a catalog write to a searchable change is typically a few seconds - the CDC poll plus the index refresh interval. A seller who changes a price and immediately searches for their own product may see the old price. That is acceptable. What is not acceptable is the checkout path trusting the index: the price charged and the stock checked must come from the authoritative stores, never from Elasticsearch.


FR3: Place an Order Without Overselling

This is the design. Everything above was setup.

An order placement has to do several things that span services: confirm the units exist and claim them, charge the buyer, and record an immutable order. Any of those can fail, and the failure cannot leave a unit claimed but unpaid, or a payment taken with no order.

New components:

  1. Inventory Service - the only writer to inventory. Owns the available and reserved counts per SKU and the reservation records.
  2. Order Service - owns orders and their state machine. Entry point for checkout, holds the idempotency table.
  3. Workflow Engine (Temporal) - runs the order as a durable workflow so a crash mid-checkout resumes rather than stranding the order.
  4. Inventory Store (Postgres) - strongly consistent, supports conditional updates. Holds inventory, reservations, orders and idempotency keys in one transactional boundary.
  5. Reservation Sweeper - a background job releasing reservations whose TTL expired.
flowchart LR
    USER["Shopper"]:::client
    OS["Order Service"]:::service
    WF["Workflow Engine"]:::async
    INV["Inventory Service"]:::service
    PAY["Payment Service"]:::service
    PG[("Postgres<br/>inventory orders keys")]:::data
    SWEEP["Reservation Sweeper"]:::async

    USER -->|"1. Place order with key"| OS
    OS -->|"2. Start order workflow"| WF
    WF -->|"3. Reserve units"| INV
    INV -->|"4. Conditional update"| PG
    WF -->|"5. Charge buyer"| PAY
    WF -->|"6. Commit order"| OS
    SWEEP -->|"7. Release expired holds"| PG

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Step-by-step flow:

  1. Shopper submits checkout with a client-generated Idempotency-Key. Order Service checks the key - if it has seen it, it returns the original response and stops.
  2. Order Service captures the current price from the catalog, writes an order in state CREATED, and starts a durable workflow.
  3. The workflow asks Inventory Service to reserve the units. Inventory performs a conditional update and creates a reservation with a 15-minute TTL.
  4. If the reservation fails, the workflow marks the order CANCELLED and returns an out-of-stock error.
  5. The workflow calls Payment Service. On success the order moves to PAID and inventory converts the reservation into a permanent decrement.
  6. On payment failure or timeout, the workflow runs its compensation: release the reservation, mark the order CANCELLED.
  7. Independently, the sweeper releases any reservation whose TTL passed without resolution - the backstop for a workflow that died in a way the engine could not recover.

Why reserve at all, rather than just decrementing at payment time? Because payment takes seconds and can involve a redirect to a bank. Decrementing only after payment means two buyers can both pass the stock check, both go to their bank, and both come back expecting a unit that one of them no longer has. Reserving converts a long, uncertain window into a short, bounded claim.

The order state machine. Every state an order can occupy, and every legal move between them. Writing it out explicitly is what stops an order ending up somewhere nobody designed:

stateDiagram-v2
    [*] --> CREATED
    CREATED --> RESERVED : units held
    CREATED --> CANCELLED : out of stock
    RESERVED --> PAID : payment captured
    RESERVED --> CANCELLED : payment failed or hold expired
    PAID --> PACKED : warehouse picked
    PACKED --> SHIPPED : carrier accepted
    SHIPPED --> DELIVERED : delivery confirmed
    PAID --> REFUNDED : cancelled after payment
    DELIVERED --> REFUNDED : return accepted
    CANCELLED --> [*]
    DELIVERED --> [*]
    REFUNDED --> [*]

Two transitions carry most of the design weight. RESERVED --> CANCELLED is the one that fires when a hold expires without payment, and it is what returns stock to the market. PAID --> REFUNDED is the compensation path, and it has to exist because PAID is the point at which the buyerโ€™s money has actually moved.


10. Technology Choices

Tier What it stores Access pattern Primary pick Alternatives
Catalog 100M product documents Single-key read, heavy cache MongoDB DynamoDB, Couchbase, Postgres with JSONB
Search Inverted index over the catalog Fuzzy text, facets, ranking Elasticsearch OpenSearch, Vespa, Solr, Algolia
Inventory and orders Stock counts, reservations, orders, idempotency keys Conditional updates, read-your-writes Postgres MySQL, CockroachDB, Spanner
Cart Per-user cart contents Read-modify-write per user, short-lived Redis with persistence DynamoDB, Memcached
Event bus Order and catalog change events Append-only, replayable, many consumers Kafka Kinesis, Pub-Sub, RabbitMQ
Workflow Order workflow state and timers Durable, resumable, timer-driven Temporal Cadence, Step Functions, Netflix Conductor

Why Postgres holds inventory and orders, and not Cassandra or DynamoDB. This is the choice the whole page turns on, and it is the opposite of the reflex answer at this scale.

The oversell invariant needs a conditional update with a reliable affected-row count - โ€œdecrement only if the result stays non-negative, and tell me whether you did it.โ€ Postgres gives that in one statement. DynamoDB can do it with a conditional expression on a single item, which works but ties you to one item per SKU and therefore to one partitionโ€™s throughput ceiling. Cassandraโ€™s last-write-wins model has no safe read-modify-write at all; lightweight transactions exist but are a Paxos round per operation and explicitly discouraged for hot paths.

The order path also needs read-your-writes, because a buyer who places an order and lands on the confirmation page must see it. And it needs multi-row atomicity: the order row, its item rows and the idempotency key have to commit together or the idempotency guarantee is a lie.

The volume argument that usually pushes people away from relational stores does not apply. 5K orders/sec of small rows is unremarkable for a tuned Postgres primary, and orders partition cleanly by time. Cassandra earns its place on this platform for things like the order event history and per-seller analytics, where volume is high and the access pattern is a partitioned append. It does not belong on the stock count.

Why a workflow engine instead of a chain of queue consumers. You can build an order saga with Kafka and a state column, and plenty of teams do. What you then write by hand is: a timer for every step that might hang, retry with backoff per step, a compensation path per failure mode, and recovery for a process that died between two steps. A workflow engine provides all four as primitives, and the code reads as the business process rather than as a message-handling graph. The cost is another stateful system to operate, which is the honest trade.


11. Data Modeling

Inventory and orders live in Postgres, and the schema is where the correctness guarantees actually live.

-- Stock per SKU. One row per SKU, and the CHECK is the invariant.
CREATE TABLE inventory (
    sku_id        UUID PRIMARY KEY,
    available     INT  NOT NULL,
    reserved      INT  NOT NULL DEFAULT 0,
    version       BIGINT NOT NULL DEFAULT 0,
    updated_at    TIMESTAMPTZ NOT NULL DEFAULT now(),
    CONSTRAINT available_non_negative CHECK (available >= 0),
    CONSTRAINT reserved_non_negative  CHECK (reserved  >= 0)
);

-- A short-lived claim on units. The sweeper reads expires_at.
CREATE TABLE reservations (
    reservation_id UUID PRIMARY KEY,
    sku_id         UUID NOT NULL REFERENCES inventory(sku_id),
    order_id       UUID NOT NULL,
    quantity       INT  NOT NULL,
    state          TEXT NOT NULL,   -- HELD, CONFIRMED, RELEASED
    expires_at     TIMESTAMPTZ NOT NULL
);
CREATE INDEX reservations_expiry ON reservations (expires_at) WHERE state = 'HELD';

-- Orders. Price is captured here, never read live at fulfilment time.
CREATE TABLE orders (
    order_id     UUID PRIMARY KEY,
    user_id      UUID NOT NULL,
    state        TEXT NOT NULL,   -- CREATED, RESERVED, PAID, PACKED, SHIPPED, DELIVERED, CANCELLED, REFUNDED
    total_minor  BIGINT NOT NULL, -- integer minor units, never a float
    currency     CHAR(3) NOT NULL,
    created_at   TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX orders_by_user ON orders (user_id, created_at DESC);

CREATE TABLE order_items (
    order_id        UUID NOT NULL REFERENCES orders(order_id),
    sku_id          UUID NOT NULL,
    quantity        INT  NOT NULL,
    unit_price_minor BIGINT NOT NULL,  -- the price at order time
    PRIMARY KEY (order_id, sku_id)
);

-- Checkout idempotency. The UNIQUE constraint is the whole mechanism.
CREATE TABLE idempotency_keys (
    key          TEXT PRIMARY KEY,
    user_id      UUID NOT NULL,
    order_id     UUID,
    response     JSONB,
    created_at   TIMESTAMPTZ NOT NULL DEFAULT now()
);

Three deliberate choices worth defending:

Money is BIGINT in minor units, never a float. A price of 1299.50 is stored as 129950. Floating point cannot represent most decimal fractions exactly, and summing a cart in floats produces totals that are off by a paisa in ways customers notice and auditors escalate.

unit_price_minor is on the order item, not read from the catalog at fulfilment. The price the buyer agreed to is a property of the order. A seller changing their price an hour later must not retroactively change what someone was charged.

Orders are indexed by (user_id, created_at DESC) because the dominant query is โ€œmy orders, newest firstโ€. Seller-side and ops-side queries have a different shape and are served from a read replica or the warehouse rather than by adding indexes to the write path.

Partitioning. Orders partition by created_at monthly - the table grows forever, queries are overwhelmingly recent, and old partitions can be detached to cold storage. Inventory does not partition by anything useful; it is small (one row per SKU) and the problem is contention on individual rows, not table size. That distinction matters and is covered in Deep Dive 1.


12. Deep Dives

Deep Dive 1: Preventing Oversell

Problem: A SKU has one unit. Two buyers check out simultaneously. Exactly one must succeed, and this must hold at 50K attempts/sec against the same row.

Bad: Read the stock, check it in application code, write the new value.

Buyer A                          Buyer B
-------                          -------
SELECT available  -> 1
                                 SELECT available  -> 1
if (1 >= 1) ok                   if (1 >= 1) ok
UPDATE SET available = 0
                                 UPDATE SET available = 0
COMMIT                           COMMIT

Result: two orders, one unit, available = 0

Both transactions read before either wrote, so neither saw the other. Wrapping this in a transaction does not fix it - under the default READ COMMITTED isolation, both reads are perfectly legal. The bug is the read-then-write gap, not the lack of a transaction.

Good: Make the check and the write one atomic statement.

UPDATE inventory
   SET available = available - 1,
       version   = version + 1
 WHERE sku_id = $1
   AND available >= 1;

Then inspect the affected row count. One row means you got the unit; zero rows means you did not. The database evaluates available >= 1 and applies the decrement under a row lock it holds for the duration, so there is no window for a second transaction to interleave. The single statement is what makes this safe, and the CHECK (available >= 0) constraint is the belt-and-braces guarantee that no code path anywhere can drive it negative.

This is correct. Its limit is throughput: every concurrent buyer for that SKU serialises on one row lock. A lock acquire, update and release is on the order of a millisecond once contention and WAL flushes are included, which caps you at roughly a thousand attempts per second for that one SKU. Fine for normal trade. Not fine for a flash sale.

Great: Reservations, plus shard the hot counter.

Reservations solve the duration problem. Payment takes seconds; holding a row lock for that long is untenable. So the conditional update moves units from available to reserved and writes a reservation row with a TTL, all in one short transaction. Payment success converts the reservation to a permanent decrement. Payment failure, or TTL expiry caught by the sweeper, returns the units.

Counter sharding solves the contention problem. Split a hot SKUโ€™s stock across N rows:

CREATE TABLE inventory_shards (
    sku_id    UUID NOT NULL,
    shard_no  SMALLINT NOT NULL,
    available INT NOT NULL CHECK (available >= 0),
    PRIMARY KEY (sku_id, shard_no)
);

1,000 units become 20 shards of 50. A buyer picks a shard at random and runs the conditional update against it. Twenty independent row locks means roughly twenty times the throughput. The costs are real and worth stating: the true available count is now a SUM across shards rather than a single read, and a buyer can be told โ€œout of stockโ€ while units remain on a shard they did not try. The standard mitigation is to retry two or three other shards before giving up, and to rebalance shards in the background as they drain unevenly.

In simple terms: Never read the stock and then write it back - ask the database to decrease it only if there is enough, in one instruction, and check whether it actually did. For checkout, donโ€™t take the unit away permanently until payment clears; put a 15-minute hold on it instead. And if one product is so popular that everyone is fighting over the same database row, split its stock into twenty rows so twenty people can be served at once.

Why a reservation TTL rather than holding until the payment resolves: a payment can hang indefinitely - a bank redirect the user abandons, a PSP timing out with no callback. Without a TTL those units are stranded until a human notices. Fifteen minutes is long enough for any legitimate payment flow and short enough that abandoned carts return stock to the market while the sale is still live.


Deep Dive 2: Orchestrating the Order Across Services

Problem: Placing an order touches inventory, payment and fulfilment. Each is a separate service with its own database. Any step can fail, time out, or succeed-but-not-reply, and the result must never be a unit claimed with no payment, or money taken with no order.

Bad: A distributed transaction across all three with two-phase commit.

It is the intuitive answer and it does not survive contact with this problem. 2PC requires every participant to hold locks through the prepare phase until the coordinator decides - so an inventory row stays locked while a payment provider thinks about it, which is exactly the multi-second lock you were trying to avoid. A coordinator crash after prepare leaves participants blocked, holding locks, unable to decide alone. And the payment provider is a third party over HTTP; it will not enlist in your transaction at all. 2PC is not an option here regardless of whether it is a good idea.

Good: A saga - a sequence of local transactions, each with a compensating action for the earlier steps.

Forward path                      Compensation if a later step fails
------------                      ----------------------------------
1. Reserve units                  Release the reservation
2. Charge the buyer               Refund the charge
3. Confirm the order              Cancel the order

Each step commits locally, so no lock is held across a service boundary. The catch is that compensation is not a rollback. A refund is a new transaction that happens after the charge was real, so there is a window where the buyerโ€™s statement shows a charge for an order that no longer exists. That window has to be short and the messaging has to be honest. See the saga pattern for the orchestration-versus-choreography trade.

Great: An orchestrated saga on a durable workflow engine.

Hand-rolled sagas fail in a specific, predictable way: the process running the saga dies between step two and step three. The charge happened, nothing recorded that it happened, and no compensation will ever run because the thing that knew to run it is gone. You discover these as a finance reconciliation ticket a week later.

A durable workflow engine persists the workflowโ€™s state and its position after every step. If the worker dies mid-saga, another worker resumes from the last completed step. Timers, retries with backoff, and compensation handlers are declared rather than coded, and they survive deploys and crashes.

The workflow drives the order state machine shown in FR3. The two transitions worth noticing there are RESERVED --> CANCELLED on hold expiry, which is the sweeperโ€™s job, and PAID --> REFUNDED, which is the compensation path that has to exist because PAID is the point of no return for the buyerโ€™s money. Every transition is explicit, so an order can never be in a state nobody designed.

In simple terms: You cannot wrap โ€œtake the stock, take the money, create the orderโ€ in one database transaction, because the payment company is not part of your database. So do each step separately, record where you are after each one, and if a later step fails, run the undo for the earlier ones. Keep that progress record in something that survives a server crash, otherwise a crash halfway through leaves a charge nobody will ever refund.


Deep Dive 3: Checkout Idempotency

Problem: The buyer taps Place Order, nothing visibly happens for four seconds, and they tap again. Or the response is lost to a dropped connection and the app retries automatically. The server receives two identical checkout requests and must create one order and one charge.

Bad: Nothing. Each request is treated as new.

Two orders, two reservations, two charges. This is one of the most common real-world e-commerce bugs because it is invisible in testing - you do not double-tap your own checkout button on a fast local network.

Good: A client-supplied idempotency key with a uniqueness constraint.

The client generates a UUID per checkout attempt (not per request - a retry reuses it) and sends it as a header. The server inserts it into idempotency_keys, whose primary key is the key itself. The second insert violates the constraint, the server catches that, and rejects the duplicate.

This stops double orders. Its weakness is what the duplicate request receives: a constraint-violation error is not the right answer. The buyerโ€™s retry gets a failure for an order that actually succeeded, so they try again, and now they are genuinely confused about whether they bought anything.

Great: Write the key inside the same transaction as the order, store the response, and replay it.

BEGIN;
  INSERT INTO idempotency_keys (key, user_id) VALUES ($1, $2);  -- fails if duplicate
  INSERT INTO orders (order_id, user_id, state, total_minor, currency)
       VALUES ($3, $2, 'CREATED', $4, $5);
  INSERT INTO order_items (...) VALUES (...);
  UPDATE idempotency_keys SET order_id = $3, response = $6 WHERE key = $1;
COMMIT;

The key and the order commit together, so there is no state where the key exists but the order does not. On a duplicate, the server reads the stored response and returns it with the original status code. The retry is indistinguishable from the first call, which is what idempotency actually means - not โ€œrejects duplicatesโ€ but โ€œthe second call has the same effect and the same answer as the firstโ€.

Two operational details that matter. Scope the key to the user, or one customerโ€™s key collides with anotherโ€™s and leaks an order. And retain keys for a bounded window - 24 hours is typical - because keeping them forever grows a table that is only useful for minutes. State the window in your API docs so clients know when a retry stops being safe.

In simple terms: Have the app make up a unique code for each checkout attempt and send it with every retry. Save that code in the database in the same transaction as the order. If the same code shows up again, donโ€™t make a second order - look up what you answered the first time and say exactly that again. The customerโ€™s retry then just works instead of erroring.

See idempotency for the general pattern and its variations.


Deep Dive 4: Keeping Search in Sync with the Catalog

Problem: Product data lives in the catalog store and must also be in Elasticsearch. These are two separate systems, and a write has to reach both.

Bad: Dual-write - the application writes to the catalog store, then writes to Elasticsearch.

Every partial failure leaves them divergent. The catalog write succeeds and the index write fails, so search shows a stale price indefinitely with nothing to detect it. Or the index write succeeds and the catalog write rolls back, so search advertises a product that does not exist. Retrying the index write helps with transient failures and does nothing for the case where the process dies between the two calls. Worse, two concurrent updates to the same product can reach the two systems in opposite orders, so the catalog ends up with version B while the index has version A, permanently, with no error anywhere.

Good: The outbox pattern. The application writes the product change and an event row to the catalog store in one transaction; a relay reads the outbox and publishes to the index.

Now there is exactly one atomic write, so the event cannot be lost and cannot exist without the data. The costs are application code in every write path and an outbox table you have to drain and prune. See the outbox pattern.

Great: Log-based change data capture.

Instead of the application announcing its own changes, read the databaseโ€™s own replication log. Debezium acts as a replication client, converts each committed change into an event with before and after images, and publishes to Kafka. An indexer consumes the topic and writes to Elasticsearch.

What this buys over the outbox: no application change at all, every change captured including ones made by a migration script or a DBA, and strict per-row ordering straight from the log. It also gives you a clean reindex path - rewind the Kafka consumer offset and replay, or run Debeziumโ€™s initial snapshot to rebuild the index from scratch without touching the application.

The consumer must be idempotent, because CDC delivers at least once. Indexing is naturally idempotent if you use the product id as the document id and include a version so an out-of-order replay cannot overwrite newer data with older. See change data capture for the snapshot-then-stream handover and the schema-evolution hazard.

In simple terms: Donโ€™t have your code write to the database and then to the search index - sooner or later one of those two writes fails and the two disagree forever with nobody noticing. Instead, let a tool watch the databaseโ€™s own change log and copy every change into the search index. Your code only ever writes to one place, and the copying is somebody elseโ€™s reliable job.


13. Design Self-Audit

Question Answer
Payment succeeds but the order write fails? The workflow owns this. Payment capture and order commit are separate steps with the order write retried on resume; if it cannot be completed, the compensation refunds. The charge is never left without either an order or a refund.
Can a buyer see their own order immediately? Yes. Orders are written to the Postgres primary and the confirmation read goes to the primary, not a replica. This is why the order store is not eventually consistent.
Flash sale on one SKU? Shard that SKUโ€™s stock across N rows (Deep Dive 1), queue or rate-limit checkout at the edge, and fail the 98% of attempts that cannot succeed early and cheaply rather than letting them reach Postgres.
Price changes between add-to-cart and checkout? The order captures unit_price_minor at order time. The cart shows an indicative price and the checkout page re-reads the authoritative one, so the buyer is shown any change before they confirm.
What is strongly consistent, what is eventual? Strong: inventory, reservations, orders, idempotency keys - all in one Postgres boundary. Eventual: the search index, the product page cache, analytics, and recommendations.
Reservation sweeper fails? Units stay held until it recovers, which costs sales but never correctness. The sweeper is idempotent and stateless, so it is safe to run several and to re-run after a gap.
Single points of failure? Postgres primary is the real one. Mitigate with synchronous replication to a standby and automated failover; accept that a failover window blocks order placement while browsing continues to serve from cache.

14. Core Flows

Flow: Checkout End to End

sequenceDiagram
    participant U as Shopper
    participant OS as Order Service
    participant WF as Workflow Engine
    participant INV as Inventory Service
    participant PAY as Payment Service

    U->>OS: POST orders with Idempotency-Key
    OS->>OS: Check key, capture price
    OS->>WF: Start order workflow
    WF->>INV: Reserve units with 15m TTL
    alt Units available
        INV-->>WF: Reservation held
        WF->>PAY: Charge buyer
        alt Payment captured
            PAY-->>WF: Success
            WF->>INV: Confirm reservation
            WF->>OS: Mark order PAID
            OS-->>U: 201 orderId state PAID
        else Payment failed
            PAY-->>WF: Declined
            WF->>INV: Release reservation
            WF->>OS: Mark order CANCELLED
            OS-->>U: 402 payment declined
        end
    else Out of stock
        INV-->>WF: Rejected
        WF->>OS: Mark order CANCELLED
        OS-->>U: 409 out of stock
    end

Walkthrough:

  1. Checkout arrives with an idempotency key. A repeat of a key already seen returns the stored response immediately and none of the following happens.
  2. Order Service reads the authoritative price, writes the order as CREATED, and starts the workflow. Order and key commit together.
  3. The workflow reserves units. Inventory runs the conditional update and writes a reservation with a TTL.
  4. On success the workflow charges the buyer, then confirms the reservation into a permanent decrement and moves the order to PAID.
  5. On payment failure the workflow releases the reservation and cancels the order, and the units are immediately back on sale.
  6. Out of stock short-circuits before any payment attempt.

Non-obvious failure: the workflow charges the buyer successfully and then the worker dies before confirming the reservation. The charge is real, the order is still RESERVED, and the reservation TTL is ticking. If nothing intervenes, the sweeper releases the units for an order that has been paid for - selling the same unit twice, which is the exact thing this design exists to prevent. Two defences, and you want both: the workflow engine resumes the workflow on another worker and completes the confirmation, and the sweeper refuses to release a reservation whose order is in state PAID, escalating it for manual resolution instead. The sweeper checking order state before releasing is the one that saves you when the workflow engine itself is the thing that failed.

Flow: Search Query

sequenceDiagram
    participant U as Shopper
    participant SS as Search Service
    participant ES as Search Index
    participant CAT as Catalog Service

    U->>SS: GET search q=running shoes brand=nike
    SS->>ES: Query with filters and facets
    ES-->>SS: Ranked ids plus facet counts
    alt Results found
        SS->>CAT: Hydrate top 20 by id
        CAT-->>SS: Product documents from cache
        SS-->>U: 200 results and facets
    else No results
        SS->>ES: Retry with fuzzy matching
        ES-->>SS: Did-you-mean suggestions
        SS-->>U: 200 suggestions
    end

Walkthrough:

  1. The query and filters go to Search Service, which builds an index query.
  2. Elasticsearch returns ranked product ids and facet counts, not full documents.
  3. Search Service hydrates only the page being shown - twenty documents - from the Catalog Service, which answers from cache.
  4. An empty result set triggers a fuzzy retry so a typo returns suggestions instead of a blank page.

Non-obvious failure: the index returns a product id that the catalog no longer has, because a deletion propagated to the catalog but not yet to the index. Hydration returns nothing for that id. Dropping it silently from the page is the right behaviour - it gives nineteen results instead of twenty rather than a broken tile or an error - and the mismatch should increment a metric, because a rising rate of it means CDC is lagging or broken.


15. Final Architecture

flowchart LR
    USER["Shopper"]:::client
    CDN["CDN<br/>pages and images"]:::edge
    SS["Search Service"]:::service
    CAT["Catalog Service"]:::service
    OS["Order Service"]:::service
    INV["Inventory Service"]:::service
    PAY["Payment Service"]:::service
    WF["Workflow Engine"]:::async
    CDC["CDC Pipeline"]:::async
    SWEEP["Reservation Sweeper"]:::async
    ES[("Search Index")]:::data
    CDB[("Catalog Store")]:::data
    PG[("Postgres<br/>inventory orders keys")]:::data
    REDIS[("Redis<br/>carts")]:::data

    USER -->|"Browse product"| CDN
    CDN -->|"Forward on miss"| CAT
    USER -->|"Search"| SS
    SS -->|"Query index"| ES
    SS -->|"Hydrate results"| CAT
    CAT -->|"Read and write documents"| CDB
    CDB -->|"Stream changes"| CDC
    CDC -->|"Index documents"| ES
    USER -->|"Update cart"| REDIS
    USER -->|"Place order"| OS
    OS -->|"Start workflow"| WF
    WF -->|"Reserve and confirm"| INV
    WF -->|"Charge buyer"| PAY
    INV -->|"Conditional update"| PG
    OS -->|"Write order and key"| PG
    SWEEP -->|"Release expired holds"| PG

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

How it works end-to-end (browse and search path):

  1. Shopper opens a product โ€” the CDN serves the page and images, and the Catalog Service is only touched on a miss
  2. Shopper searches โ€” Search Service queries the inverted index for ranked ids and facets, then hydrates only the visible page from the cached catalog
  3. Seller updates a product โ€” the write lands in the Catalog Store, and CDC streams the change into the search index within seconds

How it works end-to-end (buy path):

  1. Cart edits โ€” held in Redis, cheap to write and tolerable to lose
  2. Checkout โ€” Order Service dedupes on the idempotency key, captures the price, writes the order, and starts a durable workflow
  3. Units reserved โ€” Inventory Service runs a conditional update against Postgres and holds the units with a TTL
  4. Payment โ€” on capture, the reservation converts to a permanent decrement and the order becomes PAID; on failure the workflow releases the hold and cancels
  5. Expired holds โ€” the sweeper returns units from abandoned checkouts, refusing to touch anything whose order is already paid

Key Technologies

Term What it is
Conditional update A single UPDATE ... WHERE available >= n whose affected-row count tells you whether it applied. The primitive that prevents oversell.
Reservation A short-lived, expiring claim on stock, created at checkout and resolved by payment or by TTL.
Saga A sequence of local transactions with compensating actions, used instead of a distributed transaction across services.
Durable workflow A long-running process whose state and position persist, so a crash resumes rather than stranding the work. Temporal, Cadence, Step Functions.
Idempotency key A client-supplied token stored with the order, making a retried checkout return the original response rather than creating a second order.
CDC Change data capture. Reading the databaseโ€™s own replication log to publish every committed change downstream.
Inverted index The term-to-document mapping that makes ranked, faceted, typo-tolerant text search possible.

Whatโ€™s Expected at Each Level

Mid-level

Split the platform into services and justify the split. Identify that stock is held per SKU, not per product. Spot the read-then-write race on inventory and fix it with a single conditional update, checking the affected row count. Propose a search index separate from the catalog store and explain why LIKE cannot do the job. Recognise that checkout needs an idempotency key.

Senior

Explain why reservations exist - that payment takes seconds and you cannot hold a row lock across it - and design the TTL plus sweeper. Choose a saga over two-phase commit and name the compensating actions. Draw the order state machine and defend every transition. Explain why inventory and orders belong in a strongly consistent store while the catalog and search can be eventual, and why dual-writing to the index is wrong - reaching for outbox or CDC. Store money as integer minor units and capture price at order time.

Staff+

Address the flash sale directly: shard the hot SKU counter, accept the false-negative rate it introduces, and shed load at the edge rather than at the database. Explain what happens when a saga crashes between charging and confirming, and why the sweeper must check order state before releasing - the two independent defences for the one failure that causes a genuine double-sell. Discuss CDC operationally: snapshot-then-stream handover, idempotent consumers, replay for reindexing, and schema evolution breaking downstream consumers. Be explicit about the Postgres primary as the remaining single point of failure and what a failover costs.


๐ŸŽฏ Key Takeaways



Understand the building blocks used in this design:

Discussion

Newest first
You

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access