Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 7 min read

QPS vs TPS vs RPS - Complete Deep Dive

Prerequisites: Performance Metrics, Back-of-Envelope Estimation Used in: Uber, Payment System, Digital Wallet


The Short Answer

All three measure throughput β€” how much work per second. They differ in what counts as one unit of work and which layer of the stack you are counting at.

Unit Stands for One unit is Measured at
RPS Requests Per Second One inbound HTTP or RPC call Load balancer, API gateway, web tier
QPS Queries Per Second One operation against a data store Database, cache, search index
TPS Transactions Per Second One business or ACID unit of work Application or database transaction layer

The relationship you should expect in a real system:

TPS  ≀  RPS  ≀  QPS

One business transaction takes several requests. Each request issues several queries. So the number gets larger as you move down the stack.

πŸ’‘ The one exception is caching. A request served entirely from Redis issues zero database queries, so a well-cached system can have lower database QPS than RPS.


Why the Distinction Matters

Because you will size infrastructure with the wrong number and be off by an order of magnitude.

If an interviewer says β€œthe system handles 10K requests per second” and you provision a database for 10K QPS, you have almost certainly undersized it. Each of those requests probably reads a user record, checks a permission, loads a list, and writes an audit row. That is 10K RPS but 40K QPS.

This is the single most common back-of-envelope mistake in system design interviews: taking the number the interviewer gives you and using it at every tier without applying the fan-out.


The Fan-Out, Concretely

Take a checkout flow. The numbers below are illustrative, but the shape is what matters.

flowchart LR
    T["1 checkout<br/>1 TPS"]:::client
    R1["POST cart validate"]:::service
    R2["POST payment authorize"]:::service
    R3["POST order create"]:::service
    R4["GET confirmation"]:::service
    Q["12 datastore ops<br/>12 QPS"]:::data

    T --> R1
    T --> R2
    T --> R3
    T --> R4
    R1 --> Q
    R2 --> Q
    R3 --> Q
    R4 --> Q

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
Color Meaning
🟠 Orange Business transaction
🟒 Green Service requests
🟑 Yellow Data store operations

One checkout becomes 4 requests becomes 12 queries. So:

If the product team says β€œwe need to support 500 checkouts per second,” your database needs to handle roughly 6,000 QPS, not 500. Stating that multiplier out loud is what separates a candidate who has sized a real system from one who has not.


Where Each Term Actually Comes From

RPS is the operational number. It is what your load balancer, ingress, and web server emit by default. When you autoscale on traffic, you are almost always scaling on RPS. It is layer-agnostic about what the request does β€” a health check and a full-text search both count as one.

QPS originated at Google describing search queries and stuck as the unit for read-heavy data access. In modern usage it means one operation against a store: one SQL statement, one Redis GET, one Elasticsearch query. Many teams use QPS loosely as a synonym for RPS at the service tier. That is common enough that you should not correct an interviewer over it, but the useful discipline is to reserve QPS for the data tier so the fan-out stays visible.

TPS carries the strongest guarantee. A transaction is a unit of work that either fully completes or fully rolls back β€” one payment, one order, one BEGIN ... COMMIT block. Because a transaction holds locks and must be durable before it acknowledges, TPS is bounded by disk fsync latency and lock contention in a way RPS is not. This is why database benchmarks are quoted in TPS and why it is always the smallest of the three.

πŸ’‘ A system doing 50K RPS and 800 TPS is not contradictory. Most of that traffic is reads; only a small fraction mutates state transactionally.


Converting Between Them

You need two multipliers, and you should ask for or state both:

RPS = TPS Γ— requests_per_transaction
QPS = RPS Γ— queries_per_request

In an interview you will not be handed these. Estimate and say so:

β€œOne ride request is one transaction. It takes about 3 API calls β€” fare estimate, request, then status polling. Each call hits the driver index and the ride record, so call it 2 to 3 queries per request. At 500 rides per second that’s roughly 1,500 RPS and 4,000 QPS on the data tier. I’ll size the location store against the 4,000 number.”

That is the whole skill. Name the unit, name the multiplier, carry it to the tier you are provisioning.


Peak vs Average

Whichever unit you use, the number you provision against is peak, not average. Daily traffic is not flat β€” most consumer systems see a peak of roughly 2x to 5x their daily average, and event-driven ones far more.

Two habits worth carrying into an interview:

  1. State which one you mean. β€œ40K QPS average” and β€œ40K QPS peak” describe very different clusters.
  2. Derive peak from average explicitly. If you compute 4,000 QPS average, say you are provisioning for 12,000 to 20,000 and why. An interviewer who wanted a different peak factor will tell you.

Common Mistakes

Mistake Why it hurts
Using the interviewer’s RPS figure as your database QPS Undersizes the data tier by the fan-out factor, often 3-10x
Quoting TPS for a read-heavy system Reads are not transactions; the number will look implausibly low
Sizing against average instead of peak The system falls over at exactly the moment it matters
Treating a batch write as one unit A bulk insert of 1,000 rows is one request but very much not one row of load
Never stating which unit you mean The interviewer cannot tell whether your arithmetic is right

What to Say in an Interview

Keep it to two sentences when the topic comes up:

β€œI want to be precise about units. I’m treating one user action as a transaction, which fans out to roughly N requests and M queries β€” so the number I size the database against is the query figure, not the request figure.”

Then apply it consistently for the rest of the design. Interviewers are not testing whether you can recite three acronyms. They are testing whether your capacity numbers survive being followed down a tier.



← Back to Fundamentals Next: Performance Metrics β†’

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access