QPS vs TPS vs RPS - Complete Deep Dive
Prerequisites: Performance Metrics, Back-of-Envelope Estimation Used in: Uber, Payment System, Digital Wallet
The Short Answer
All three measure throughput β how much work per second. They differ in what counts as one unit of work and which layer of the stack you are counting at.
| Unit | Stands for | One unit is | Measured at |
|---|---|---|---|
| RPS | Requests Per Second | One inbound HTTP or RPC call | Load balancer, API gateway, web tier |
| QPS | Queries Per Second | One operation against a data store | Database, cache, search index |
| TPS | Transactions Per Second | One business or ACID unit of work | Application or database transaction layer |
The relationship you should expect in a real system:
TPS β€ RPS β€ QPS
One business transaction takes several requests. Each request issues several queries. So the number gets larger as you move down the stack.
π‘ The one exception is caching. A request served entirely from Redis issues zero database queries, so a well-cached system can have lower database QPS than RPS.
Why the Distinction Matters
Because you will size infrastructure with the wrong number and be off by an order of magnitude.
If an interviewer says βthe system handles 10K requests per secondβ and you provision a database for 10K QPS, you have almost certainly undersized it. Each of those requests probably reads a user record, checks a permission, loads a list, and writes an audit row. That is 10K RPS but 40K QPS.
This is the single most common back-of-envelope mistake in system design interviews: taking the number the interviewer gives you and using it at every tier without applying the fan-out.
The Fan-Out, Concretely
Take a checkout flow. The numbers below are illustrative, but the shape is what matters.
flowchart LR
T["1 checkout<br/>1 TPS"]:::client
R1["POST cart validate"]:::service
R2["POST payment authorize"]:::service
R3["POST order create"]:::service
R4["GET confirmation"]:::service
Q["12 datastore ops<br/>12 QPS"]:::data
T --> R1
T --> R2
T --> R3
T --> R4
R1 --> Q
R2 --> Q
R3 --> Q
R4 --> Q
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef data fill:#fbbf24,stroke:#92400e,color:#000
| Color | Meaning |
|---|---|
| π Orange | Business transaction |
| π’ Green | Service requests |
| π‘ Yellow | Data store operations |
One checkout becomes 4 requests becomes 12 queries. So:
- 1 TPS at the business layer
- 4 RPS at the API tier β a 4x multiplier
- 12 QPS at the data tier β a 12x multiplier
If the product team says βwe need to support 500 checkouts per second,β your database needs to handle roughly 6,000 QPS, not 500. Stating that multiplier out loud is what separates a candidate who has sized a real system from one who has not.
Where Each Term Actually Comes From
RPS is the operational number. It is what your load balancer, ingress, and web server emit by default. When you autoscale on traffic, you are almost always scaling on RPS. It is layer-agnostic about what the request does β a health check and a full-text search both count as one.
QPS originated at Google describing search queries and stuck as the unit for read-heavy data access. In modern usage it means one operation against a store: one SQL statement, one Redis GET, one Elasticsearch query. Many teams use QPS loosely as a synonym for RPS at the service tier. That is common enough that you should not correct an interviewer over it, but the useful discipline is to reserve QPS for the data tier so the fan-out stays visible.
TPS carries the strongest guarantee. A transaction is a unit of work that either fully completes or fully rolls back β one payment, one order, one BEGIN ... COMMIT block. Because a transaction holds locks and must be durable before it acknowledges, TPS is bounded by disk fsync latency and lock contention in a way RPS is not. This is why database benchmarks are quoted in TPS and why it is always the smallest of the three.
π‘ A system doing 50K RPS and 800 TPS is not contradictory. Most of that traffic is reads; only a small fraction mutates state transactionally.
Converting Between Them
You need two multipliers, and you should ask for or state both:
RPS = TPS Γ requests_per_transaction
QPS = RPS Γ queries_per_request
In an interview you will not be handed these. Estimate and say so:
βOne ride request is one transaction. It takes about 3 API calls β fare estimate, request, then status polling. Each call hits the driver index and the ride record, so call it 2 to 3 queries per request. At 500 rides per second thatβs roughly 1,500 RPS and 4,000 QPS on the data tier. Iβll size the location store against the 4,000 number.β
That is the whole skill. Name the unit, name the multiplier, carry it to the tier you are provisioning.
Peak vs Average
Whichever unit you use, the number you provision against is peak, not average. Daily traffic is not flat β most consumer systems see a peak of roughly 2x to 5x their daily average, and event-driven ones far more.
Two habits worth carrying into an interview:
- State which one you mean. β40K QPS averageβ and β40K QPS peakβ describe very different clusters.
- Derive peak from average explicitly. If you compute 4,000 QPS average, say you are provisioning for 12,000 to 20,000 and why. An interviewer who wanted a different peak factor will tell you.
Common Mistakes
| Mistake | Why it hurts |
|---|---|
| Using the interviewerβs RPS figure as your database QPS | Undersizes the data tier by the fan-out factor, often 3-10x |
| Quoting TPS for a read-heavy system | Reads are not transactions; the number will look implausibly low |
| Sizing against average instead of peak | The system falls over at exactly the moment it matters |
| Treating a batch write as one unit | A bulk insert of 1,000 rows is one request but very much not one row of load |
| Never stating which unit you mean | The interviewer cannot tell whether your arithmetic is right |
What to Say in an Interview
Keep it to two sentences when the topic comes up:
βI want to be precise about units. Iβm treating one user action as a transaction, which fans out to roughly N requests and M queries β so the number I size the database against is the query figure, not the request figure.β
Then apply it consistently for the rest of the design. Interviewers are not testing whether you can recite three acronyms. They are testing whether your capacity numbers survive being followed down a tier.
Related Concepts
- Performance Metrics β latency, percentiles, and Littleβs Law, which converts throughput into the concurrency you need
- Back-of-Envelope Estimation β the full estimation workflow these units feed into
- Database Sharding β what you do once a single node cannot absorb your QPS
- Caching β the main lever for keeping database QPS below RPS
- Terminology and Fine Distinctions β more X-vs-Y answers in the same format
| β Back to Fundamentals | Next: Performance Metrics β |