Performance Metrics & Vocabulary - Complete Deep Dive
Prerequisites: Back-of-Envelope Estimation, Terminology Used in: every design β this is the language you use to state and defend non-functional requirements
Why this matters more than it looks
Every system design interview opens with non-functional requirements. Candidates say βit should be fast and scalable.β Strong candidates say βp99 read latency under 200ms at 50K rps, 99.9% availability, and Iβll size the thread pools from Littleβs Law.β
Same design, completely different signal. This page is that vocabulary.
Latency vs Response time vs Service time
These three get used interchangeably and they are not the same thing.
βββββββββββββββββ response time (what the user feels) ββββββββββββββββββ
β β
client ββ network out β queue wait β service time β serialize β network back ββ client
β βββ latency (waiting, not working) βββ β
| Term | Definition | Where you see it |
|---|---|---|
| Service time | Time actually spent processing | Your appβs own timer, server-side histogram |
| Latency | Strictly: time spent waiting (queue + network). Loosely: a synonym for response time. | Load balancer metrics |
| Response time | Total client-observed duration | RUM, client-side timing, synthetic checks |
Say this: βResponse time is what the user feels; service time is what my server logs. The gap is queueing β and under load queueing dominates, which is why p99 explodes long before CPU reaches 100%.β
β οΈ The classic gotcha: your server-side p99 is 40ms but users complain. The 40ms excludes queue wait at the LB, TLS handshakes, DNS, and mobile radio wake-up. Always be clear about where you measured.
Latency vs Throughput
Not opposites, and improving one usually costs the other.
Latency = time per operation (ms/op) β lower is better
Throughput = operations per second (ops/s) β higher is better
Throughput needs a unit before it means anything. Requests, queries, and transactions are three different counts at three different tiers β see QPS vs TPS vs RPS for the fan-out multipliers between them.
| Technique | Throughput | Latency |
|---|---|---|
| Batching (wait to fill a batch) | ββ | β (you pay the wait) |
| Pipelining (overlap work) | β | ~ |
| Adding replicas | β | ~ until coordination costs bite |
| Compression | β (less bytes on wire) | β CPU, β transfer β net win over WAN, loss over LAN |
| Caching | β | ββ (the rare win-win) |
The highway analogy that actually works: latency is how long your car takes to travel the road. Throughput is how many cars pass per minute. Adding lanes raises throughput without making your trip faster. Raising the speed limit lowers latency. And at high enough traffic, latency degrades even though throughput is at maximum β thatβs congestion, and itβs exactly what happens to a saturated service.
Littleβs Law β the single most useful formula
L = Ξ» Γ W
L = average number of requests in the system (concurrency)
Ξ» = arrival rate (requests per second)
W = average time in the system (response time)
Worked example:
10,000 rps Γ 0.2 s average response time = 2,000 concurrent requests in flight
β Thread-per-request model needs ~2,000 threads (at ~1MB stack = 2GB just for stacks)
β Across 20 servers: 100 threads each
β Each with a 20-connection DB pool: 400 total DB connections β fine
β If a downstream slows to 1s: L = 10,000 Γ 1 = 10,000 in flight
β 500 threads/server needed; pools exhaust; requests queue; p99 collapses
β An unrelated endpoint on the same pool now times out too
That last line is the whole point. One slow dependency causes a fleet-wide outage through pool exhaustion, and Littleβs Law is how you predict it. Itβs also the argument for bulkheads (separate pools per dependency) and for aggressive timeouts.
Say this: βI size pools with Littleβs Law: at 10K rps and 200ms, thatβs 2,000 requests in flight. Iβd also set a timeout budget so a slow downstream canβt inflate W without bound β otherwise concurrency grows until the pool exhausts and takes down endpoints that donβt even use that dependency.β
Percentiles β and why averages lie
| Metric | Meaning | Verdict |
|---|---|---|
| Average (mean) | Sum Γ· count | Actively misleading. One 30s timeout hides among thousands of 5ms requests. |
| p50 / median | Half are faster | Describes the typical case, says nothing about pain |
| p90 | 1 in 10 is slower | Where degradation first becomes visible |
| p99 | 1 in 100 is slower | The standard SLO target. At 10K rps thatβs 100 users/second suffering. |
| p99.9 | 1 in 1,000 is slower | Often your biggest customers β they make the most requests, so they hit the tail most often |
| max | The worst one | Useful for spotting timeouts, too noisy for an SLO |
Requests (ms): 5, 6, 6, 7, 7, 8, 8, 9, 10, 30000
average = 3,009 ms β "our average is 3 seconds!" β nonsense
p50 = 7 ms β the truth for typical users
p99 = 30,000 ms β the truth for the user who is suffering
β οΈ You cannot average percentiles. Host Aβs p99 and host Bβs p99 averaged is not the fleet p99. To aggregate you need mergeable structures β histograms with fixed buckets, HDR histograms, or t-digest β then compute the percentile from the merged distribution. This is a genuinely common interview follow-up, and getting it right is a strong signal.
β οΈ Percentiles donβt average over time either. The mean of 60 one-minute p99s is not the hourβs p99.
Tail latency amplification β the fan-out killer
If one request fans out to N services in parallel and waits for all of them, the slowest one determines your response time.
Each service: 99% fast, 1% slow (p99 tail)
N = 1 β P(no slow call) = 0.99 β 1% of requests are slow
N = 10 β P = 0.99^10 β 0.904 β ~10% of requests are slow
N = 100β P = 0.99^100 β 0.366 β ~63% of requests are slow (!)
A p99 problem in one service becomes a p37 problem for the user at 100-way fan-out. This is why large-scale systems obsess over tail latency rather than average.
Mitigations, in order of how often theyβre the right answer:
| Technique | How it works | Cost |
|---|---|---|
| Reduce fan-out | Denormalize so one call answers the question | Storage, write complexity |
| Timeouts + partial results | Return what arrived in 100ms, degrade the rest | Incomplete responses |
| Hedged requests | At p95, send a duplicate to another replica, take the first answer | ~5% extra load for a big tail win |
| Tied requests | Send to two replicas; the one that starts cancels the other | Slightly more complex |
| Request budgets | Total deadline propagated to every hop; hops that canβt fit fail fast | Requires deadline propagation |
| Avoid sequential hops | 5 sequential 20ms hops = 100ms floor before any work | Architectural |
Say this: βWith fan-out Iβd worry about tail amplification β 10 parallel calls each with a 1% slow tail means roughly 1 in 10 user requests is slow. Iβd propagate a deadline, return partial results past 150ms, and consider hedged requests on the read path where the work is idempotent and cheap.β
Utilization and the queueing wall
The single most counterintuitive fact in performance: queue wait grows non-linearly with utilization and goes vertical near 100%.
For an M/M/1 queue:
Average wait W = S / (1 - Ο) where S = service time, Ο = utilization
Ο = 50% β W = 2 Γ S (wait equals service time)
Ο = 80% β W = 5 Γ S
Ο = 90% β W = 10 Γ S
Ο = 95% β W = 20 Γ S
Ο = 99% β W = 100 Γ S β the wall
latency
β β
β β
β β±
β β±
β____________________β±
ββββββββββββββββββββββββββββββββββββ utilization
0% 50% 80% 90% 95% 99%
Practical consequences:
- Target 60-70% utilization, not 95%. βWeβre only at 70% CPU, why is it slow?β β because at 70% youβre already waiting 2.3Γ service time, and a traffic spike puts you on the vertical part of the curve.
- Variance makes it worse. The formula above assumes exponential service times. Bursty arrivals (which real traffic is) and high service-time variance push the wall left.
- This is why autoscaling on 80% CPU is late. By the time the new instance boots (60-180s), youβre at 95%+ and p99 has already blown the SLO. Scale on queue depth or p99 latency, and keep headroom.
Throughput vocabulary
| Term | Precise meaning |
|---|---|
| QPS / RPS | Queries/requests per second β a rate. Meaningless for capacity without a latency figure. |
| TPS | Transactions per second. One transaction may be several queries β donβt mix the units. |
| Concurrency | Requests in flight simultaneously (= Ξ» Γ W). This is what exhausts resources, not the rate. |
| Bandwidth | Theoretical max capacity of the link |
| Throughput | What you actually achieve |
| Goodput | Useful payload delivered, excluding retransmits, headers, and protocol overhead. Under a retry storm, throughput can be high while goodput collapses β a great phrase to have ready. |
| IOPS | I/O operations per second β the disk metric. 4KB random reads is the conventional unit. |
| Peak vs Average | Peak-to-average is typically 2-5Γ for consumer apps; you provision for peak (plus headroom), you bill on average. Always state which one your number is. |
Availability math
The nines
| Availability | Downtime/year | Downtime/month | Downtime/week |
|---|---|---|---|
| 99% (βtwo ninesβ) | 3.65 days | 7.2 hours | 1.68 hours |
| 99.9% (βthree ninesβ) | 8.77 hours | 43.8 min | 10.1 min |
| 99.95% | 4.38 hours | 21.9 min | 5.04 min |
| 99.99% (βfour ninesβ) | 52.6 min | 4.38 min | 1.01 min |
| 99.999% (βfive ninesβ) | 5.26 min | 26.3 sec | 6.05 sec |
β οΈ Five nines is less than one bad deploy per year. You cannot reach it with a human in the recovery path β it requires automated failover thatβs faster than a page. If you claim five nines in an interview, be ready to justify it.
Serial dependencies multiply
Service chain: gateway β auth β orders β inventory β payments
Each at 99.9%:
0.999^5 = 0.995 β 99.5% β 1.8 DAYS of downtime per year
This is the strongest technical argument against deep synchronous call chains. Five hops of βvery reliableβ services produce a mediocre system.
Redundancy adds nines (if failures are independent)
Two independent instances at 99% each:
P(both down) = 0.01 Γ 0.01 = 0.0001 β availability 99.99%
β οΈ βIndependentβ is doing enormous work in that sentence. Same AZ, same deploy, same config push, same certificate expiry, same poisoned cache entry, same dependency β correlated failure eats the benefit. Say this out loud: βredundancy only adds nines to the extent the failures are actually independent, which is why Iβd spread across AZs and stagger deploys.β
SLI vs SLO vs SLA vs Error budget
| Term | Definition | Example |
|---|---|---|
| SLI β Indicator | The measurement | β% of requests served in <300ms, measured at the LBβ |
| SLO β Objective | Your internal target | β99.9% of requests <300ms over 30 daysβ |
| SLA β Agreement | The external contract, with penalties | β99.5% or you get 10% creditβ |
| Error budget | 100% β SLO β the failure youβre allowed | 0.1% of 30 days = 43 minutes |
Three rules worth knowing:
- SLA is always looser than SLO. You need internal headroom before money is at stake.
- The error budget is a budget β spend it. Untouched budget means youβre over-investing in reliability and under-investing in features. Exhausted budget means freeze deploys until it recovers.
- A good SLI measures the userβs experience, not your serverβs comfort. βCPU < 80%β is not an SLI.
Say this: βIβd define the SLO on the user-facing indicator β p99 latency and success rate at the edge β and use the error budget to govern release velocity. That keeps the reliability conversation quantitative instead of βthe site feels slow.ββ
Failure and recovery metrics
| Metric | Meaning |
|---|---|
| MTBF | Mean Time Between Failures β for repairable systems |
| MTTF | Mean Time To Failure β for non-repairable components (a disk) |
| MTTD | Mean Time To Detect β usually the biggest slice, and the cheapest to improve |
| MTTA | Mean Time To Acknowledge β paging and on-call health |
| MTTR | Mean Time To Repair/Recover |
Availability = MTBF / (MTBF + MTTR)
The insight: halving MTTR improves availability exactly as much as doubling MTBF β and itβs usually 10Γ cheaper. Better alerting, faster rollback, and rehearsed runbooks beat heroic engineering to prevent all failure.
| Β | RPO | RTO |
|---|---|---|
| Question | How much data can I lose? | How long can I be down? |
| Set by | Backup/replication frequency | Recovery process speed |
| RPO = 0 | Synchronous replication (latency cost) | β |
| RTO = 0 | β | Active-active with automatic failover |
Latency numbers to have memorized
| Operation | Time | Relative |
|---|---|---|
| L1 cache reference | 1 ns | 1Γ |
| L2 cache reference | 4 ns | 4Γ |
| Branch mispredict | 3 ns | β |
| Mutex lock/unlock | 17 ns | β |
| Main memory reference | 100 ns | 100Γ |
| Compress 1KB (snappy) | 2 ΞΌs | β |
| Read 1MB sequentially from RAM | 3 ΞΌs | β |
| SSD random read | 16 ΞΌs | 16,000Γ |
| Read 1MB sequentially from SSD | 49 ΞΌs | β |
| Round trip within same datacenter | 500 ΞΌs | β |
| HDD seek | 2 ms | 2,000,000Γ |
| Read 1MB sequentially from HDD | 825 ΞΌs | β |
| Round trip India β US East | ~200 ms | β |
The three ratios that actually matter in an interview:
- Memory is ~100Γ faster than SSD, SSD is ~100Γ faster than a disk seek. Thatβs why caching works.
- Sequential is ~100Γ faster than random on any storage. Thatβs why LSM trees, WALs, and Kafka append.
- Cross-continent round trips are ~200ms and physics wonβt fix it. Thatβs why you need CDNs and regional replicas β no amount of optimization beats the speed of light.
Interview application
Stating non-functional requirements: βReads: p99 < 200ms at 50K rps. Writes: p99 < 500ms at 5K rps. Availability 99.9% for reads, 99.99% for the payment path. Consistency: strong for balances, eventual (< 2s) for feeds. Peak-to-average 3Γ, so Iβll provision for 150K rps read capacity with 30% headroom.β
Defending a design: βI keep the read path to two hops because five 99.9% services in series is 99.5% β a day and a half of downtime a year. Anything not on the critical path moves to async via a queue.β
Sizing: βAt 50K rps with 40ms service time, Littleβs Law gives 2,000 concurrent requests. With 200 concurrent per instance thatβs 10 instances, and Iβd run 15 for headroom and AZ loss tolerance.β
Common Interview Questions
Q: βWhat latency should we target?β A: βDepends on the interaction. Under 100ms feels instant; 100-300ms feels responsive; past 1s users notice and start abandoning. Iβd set p99 < 200ms for the read path and be explicit that Iβm measuring at the edge, not server-side.β
Q: βWhy p99 and not average?β A: βAverages hide the tail β one 30s timeout disappears into thousands of 5ms requests. At 10K rps, p99 is 100 users every second having a bad experience. And with fan-out, a 1% tail per service becomes a 10% tail for the user across 10 parallel calls.β
Q: βHow do you compute p99 across 50 servers?β A: βYou canβt average their p99s. Each host exports a histogram (or t-digest), you merge the distributions centrally, then compute the percentile. Averaging percentiles is a common and wrong shortcut.β
Q: βYour CPU is at 70% and latency is bad. Why?β A: βQueueing. At 70% utilization average wait is already about 2.3Γ service time, and the curve is non-linear β bursty arrivals put you briefly at 95%+ where wait is 20Γ service time. Also check whether youβre actually CPU-bound: lock contention, GC pauses, or a saturated connection pool all produce latency at moderate CPU.β
Q: βLatency vs throughput β can you have both?β A: βCaching gives you both. Almost everything else trades: batching raises throughput and hurts latency, and adding capacity raises throughput without changing per-request latency. The one thing that helps both is removing work β fewer hops, fewer bytes, fewer round trips.β
Q: βHow much capacity do you provision?β A: βPeak, not average β typically 2-5Γ average for consumer traffic β targeting 60-70% utilization at peak so I stay off the vertical part of the queueing curve, plus enough spare that losing an AZ doesnβt push the survivors past that line.β
Q: βWhatβs the difference between an SLO and an SLA?β A: βSLO is the internal target, SLA is the contractual promise with penalties. The SLO is always stricter, so you find out youβre in trouble before your customers get a credit.β
Decision Framework
| Question | Metric to reach for |
|---|---|
| Is the system fast enough for users? | p99 response time at the edge |
| Can it handle the load? | Throughput at target p99 (not max throughput) |
| How many threads / connections? | Littleβs Law: L = Ξ» Γ W |
| How much headroom? | Keep peak utilization β€ 70% |
| How reliable must this be? | SLO + error budget, then check serial-dependency math |
| How fast must we recover? | RTO; and improve MTTD/MTTR before MTBF |
| How much data can we lose? | RPO β replication mode |
| Why is the tail bad? | Fan-out width, queueing, GC pauses, lock contention |
| β Back to Fundamentals | QPS vs TPS vs RPS | Next: Networking Basics β |
Where This Shows Up
This concept is load-bearing in these designs - each link goes straight to the design that leans on it:
- Design a Metrics and Monitoring System - percentiles computed over a stream, and why averages hide the problem
- Design a Twitter Feed - a read-latency budget that forces fan-out on write
- Design a URL Shortener - a redirect latency target that decides the whole storage choice
| All Concepts | All HLD Designs |