Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 58 min read

Designing Spotify - Music Streaming Platform

Difficulty: Intermediate Topics: Audio Streaming, CDN, Offline Sync, Play Counting Asked at: Spotify, Amazon, Google, Apple Prerequisites:CDN and Idempotency


1. Understanding the Problem

Spotify is an audio streaming service. You search a catalog of roughly 100 million tracks, tap one, and sound comes out of the speaker almost immediately. You collect tracks into playlists and reorder them. You download an album before a flight and it plays in airplane mode. Behind all of that, every play that crosses a threshold has to be counted, because a rights holder gets paid for it.

Real examples: Spotify, Apple Music, YouTube Music, Amazon Music, SoundCloud, JioSaavn.

If you have already read the Netflix design, do not paste it here. Video and audio streaming share a delivery story and almost nothing else. The differences are not cosmetic - they move the hard problems to different places.

Dimension Netflix Spotify
Asset size ~3 GB per variant per title ~3.5 MB per track at 128 kbps
Catalog size ~15K titles ~100M tracks
Session shape 1-2 hours, foreground, one asset 2-3 hours, background, 40+ assets back to back
Transition cost Nobody notices a 1s gap between episodes An audible gap between album tracks is a bug report
Offline Nice-to-have A headline feature people pay for
Per-asset accounting None - nobody is paid per view Every qualifying play is a royalty event

Two consequences fall out of that table and they shape everything below. First, the catalog is four orders of magnitude larger than Netflix’s while each file is three orders smaller, which flips the cache problem: you cannot pre-position 100M tracks everywhere, so cache hit rate on the long tail is the thing to worry about. Second, the write path matters here in a way it does not for video. Netflix tracks a play position for convenience. We track plays because money moves.


2. Naive First Cut

flowchart LR
    PHONE["Phone App"]:::client
    API["API Server"]:::service
    DB[("Postgres<br/>tracks and play counts")]:::data
    FS[("File Server<br/>one mp3 per track")]:::data

    PHONE --> API
    API --> DB
    API --> FS

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Color Meaning
🟣 Purple Clients
🟒 Green Services
🟑 Yellow Data stores
πŸ”΅ Blue Edge / CDN

One file per track on a file server. The client downloads the whole file, then plays it. Play counts live in an integer column on the track row, incremented on each play. Playlists are rows with an integer position.

Why this breaks:

The rest of this doc evolves that sketch into a pre-transcoded, edge-delivered catalog with an auditable play ledger.


3. Prior Art We’re Drawing From


4. Functional Requirements

Core (Top 3)

  1. Stream a track on demand - user taps a track, audio starts near-instantly and plays through without a rebuffer
  2. Build and play playlists - create a playlist, add and remove tracks, reorder them, play it start to finish
  3. Download for offline listening - save tracks to the device and play them with no network

Below the Line


5. Non-Functional Requirements

NFR Target
Start-up latency P95 under 200 ms from tap to first audio sample, on a warm connection
Gapless continuity No audible gap between consecutive tracks; zero rebuffers on a stable link
Scale 50M concurrent streams at peak, 600M registered users
Play-count integrity Every qualifying play counted exactly once, auditable, and recomputable from raw events

Below the Line


6. Scale Estimation (Back-of-Envelope)

Assume an average track of 3.5 minutes, so 210 seconds of audio.

The shape to notice: the read path is enormous and almost entirely cacheable, the catalog is tiny in bytes, and the single stateful thing with a hard correctness requirement is the stream of play events.


7. Core Entities


8. API / System Interface

GET /v1/search?q=<query>&type=track,album,artist&limit=20
  Response: { tracks: [{ trackId, name, artist, albumArt, durationMs }], ... }

POST /v1/playback/start
  Body: { trackId, deviceId, networkHint }
  Response: { sessionId, renditions: [{ renditionId, bitrateKbps, baseUrl, totalBytes }],
              licenceToken, playbackIdempotencyKey }
  Note: server picks the default rendition, client may override after measuring throughput

GET /v1/audio/{renditionId}
  Headers: Range: bytes=0-131071, Authorization: Bearer <licenceToken>
  Response: 206 Partial Content, audio chunk bytes
  Note: served by CDN, not by the application tier

PATCH /v1/playlists/{playlistId}/items
  Body: { op: "move", itemId, afterItemId?, beforeItemId? }
  Response: { itemId, rankKey: "a0m", revision: 412 }
  Note: client sends neighbours, server returns the computed ordering key

POST /v1/offline/{trackId}/lease
  Body: { deviceId }
  Response: { encryptedKeyBlob, expiresAt, renditionId }

POST /v1/plays
  Body: { events: [{ eventId, trackId, startedAt, msPlayed, deviceId, offline }] }
  Response: { accepted: 42, duplicates: 3 }
  Note: batched, retry-safe, and the same endpoint replays events queued while offline

Security note: The client never tells us who it is or what it is entitled to. Subscription tier, market, and device count come from the account on the server side. msPlayed arrives from the client and cannot be trusted blindly - it is cross-checked against session duration and CDN byte delivery before it is allowed to move money.


9. High-Level Design

Build it one requirement at a time.

FR1: Stream a Track on Demand

Tap to sound in under 200 ms, from a catalog of 100M tracks. That budget is the whole design constraint. A round trip to a distant origin can eat 150 ms of it on its own, so the audio bytes have to come from somewhere close, and the first bytes have to be useful on their own.

Two decisions follow immediately.

Pre-transcode, do not transcode on demand. One upload fans out into a fixed ladder of 96, 128, 256 and 320 kbps renditions, written once and kept forever. Encoding a 3.5-minute track takes seconds of CPU, not milliseconds - that alone blows a 200 ms budget before a single byte moves. Doing it per request would also mean encoding the same popular track millions of times and producing a byte stream that no cache can share.

Chunk the audio. A rendition is stored and served as small pieces - either fixed byte ranges over one object or short segments as separate objects. The client asks for the first chunk, hands it to the decoder, and keeps fetching while playback runs. πŸ’‘ A 128 kbps stream consumes 16 KB per second of audio. A 128 KB first chunk is 8 seconds of music - enough to start playing and to absorb a network hiccup, and it transfers in well under 200 ms on any usable connection.

New components:

  1. Catalog Service - answers β€œwhat is this track, which renditions exist, is it licensed in this market”. Backed by Postgres, fronted by an aggressive cache, because the read-to-write ratio is 100,000 to 1.
  2. Transcode Pipeline - takes the uploaded master and fans it out into the bitrate ladder, chunks each rendition, and writes them to object storage. Runs out of band; nothing user-facing waits on it.
  3. Audio Object Store - holds every chunk of every rendition. Immutable once written, which is what makes the next component work so well.
  4. CDN - caches chunks at the edge. Immutable, publicly addressable, byte-range-friendly objects are close to the ideal CDN workload.
  5. Playback Service - authorizes the session, picks a starting rendition, returns the chunk base URL and a short-lived token.
flowchart LR
    LABEL["Label Upload"]:::external
    TP["Transcode Pipeline<br/>96 128 256 320 kbps"]:::service
    OBJ[("Audio Object Store<br/>chunked renditions")]:::data
    CAT[("Catalog DB<br/>tracks albums artists")]:::data
    PHONE["Phone App"]:::client
    PS["Playback Service"]:::service
    CDN["CDN<br/>edge chunk cache"]:::edge

    LABEL -->|"1. Upload master"| TP
    TP -->|"2. Write chunked renditions"| OBJ
    TP -->|"3. Mark track playable"| CAT
    PHONE -->|"4. POST playback start"| PS
    PS -->|"5. Read track and renditions"| CAT
    PS -->|"6. Return base url and token"| PHONE
    PHONE -->|"7. Range request first chunk"| CDN
    CDN -->|"8. Fill on miss"| OBJ

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
    classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0

Step-by-step flow:

  1. A label delivers a master file. The Transcode Pipeline validates it and expands it into four renditions.
  2. Each rendition is chunked and written to object storage under an immutable key that includes the rendition id. Once all four land, the track is marked playable in the Catalog DB and becomes searchable.
  3. User taps a track. The client calls POST /v1/playback/start, and Playback Service checks the subscription and the market licence before reading the rendition list from the Catalog Service.
  4. It returns the chunk base URL, the rendition it recommends for this device and network, and a short-lived access token.
  5. The client issues a range request for the first chunk against the CDN and starts decoding as soon as it arrives. On a miss the edge pulls from object storage, serves the client, and keeps the chunk for the next listener in that region.

Why the client picks the rendition and not the server. The server has an account and a device type. The client has an actual measured throughput number from the chunk it just downloaded. We give the client a recommendation so the first chunk is not a guess, then let it switch renditions at any chunk boundary. Switching mid-track is cheap because chunks are independent.

What this leaves broken. Playback works and starts fast on a cache hit. Three gaps:


FR2: Build and Play Playlists

A playlist is an ordered list that users edit constantly, from multiple devices, sometimes offline. The library is the other half: saved tracks and albums, read on practically every app open.

The ordering problem is the interesting part. Store position as an integer and the data model is obvious and wrong. Inserting a track at the top of a 10,000-item playlist means UPDATE playlist_items SET position = position + 1 across 10,000 rows. Dragging one item from the bottom to the top is the same cost. Two devices reordering while offline both renumber from their own stale view, and whichever syncs second overwrites an ordering the user deliberately chose.

The fix is to stop numbering positions and start naming them. Each item carries an ordering key - a value that only needs to sort correctly relative to its neighbours, never to be dense or gap-free.

πŸ’‘ A lexicographic ordering key is just a short string. Between "a" and "b" you can insert "am". Between "am" and "b" you can insert "an". Between "am" and "an" you can insert "amm". There is always room, the key grows one character at a time, and a move writes exactly one row. Fractional numeric keys work the same way but hit floating-point precision limits after a few dozen moves in the same gap, so strings are the safer choice.

New components:

  1. Playlist Service - owns playlists and their items. Computes ordering keys from the neighbours the client sends, and bumps a per-playlist revision on every mutation so clients can detect drift.
  2. Library Store - a wide-column store (Cassandra) holding per-user saved tracks and playlist items, partitioned by owner. This is high-volume, per-user, append-mostly data with no cross-user queries, which is exactly what Cassandra is good at.
flowchart LR
    PHONE["Phone App"]:::client
    PLS["Playlist Service"]:::service
    PS["Playback Service"]:::service
    LIB[("Library Store<br/>playlists and saved tracks")]:::data
    CAT[("Catalog DB<br/>tracks albums artists")]:::data
    CDN["CDN<br/>edge chunk cache"]:::edge

    PHONE -->|"1. Move item with neighbours"| PLS
    PLS -->|"2. Write one row with rank key"| LIB
    PHONE -->|"3. Read playlist page"| PLS
    PLS -->|"4. Hydrate track metadata"| CAT
    PHONE -->|"5. Start next queued track"| PS
    PHONE -->|"6. Fetch chunks"| CDN

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Step-by-step flow:

  1. User drags a track. The client sends op: move with the ids of the items it will sit between.
  2. Playlist Service computes a key strictly between those two neighbours and writes one row. The playlist revision increments.
  3. On open, the client reads a page of items ordered by rank key - a straight clustering-key scan, no sort - and Playlist Service hydrates it with track metadata served from cache almost every time.
  4. Pressing play hands the queue to the playback path from FR1, which authorizes the first track and pre-warms the second. Chunks come from the CDN exactly as before; playlists add no new delivery path.

Why send neighbours instead of an index. An index is a statement about the whole list and goes stale the moment anyone else edits. β€œPut this between X and Y” stays meaningful even if the list changed underneath, and if X or Y has been deleted the server can fall back to the surviving neighbour instead of writing the item somewhere arbitrary.

What this leaves broken. Two devices editing the same playlist offline still converge to an order, just not necessarily a satisfying one - covered honestly in the self-audit rather than hand-waved. And the queue still only works with a live connection, which is FR3.


FR3: Download for Offline Listening

Offline is not a degraded mode here. People pay for the subscription partly because of it, and the design has to assume a device that is offline for hours and then reconnects with a backlog.

Two problems, and they are different. The audio has to be on the device and useless if extracted. The listening that happened while offline has to reach the play pipeline without being counted twice when the client retries.

New components:

  1. Licence Service - issues a device-bound, expiring licence containing the content key for a downloaded rendition, wrapped so only that device can unwrap it. Enforces the per-account device cap.
  2. Encrypted Local Cache - the downloaded chunks on device, stored encrypted. The key never sits next to the audio in plaintext, and the cache is keyed to the device so a copied directory decrypts to noise elsewhere.
  3. Sync Service - reconciles the device on reconnect: replays queued play events to the event bus, pushes playlist mutations, refreshes licences that are near expiry, and drops downloads for tracks that have left the catalog.
flowchart LR
    PHONE["Phone App"]:::client
    LIC["Licence Service"]:::service
    SYNC["Sync Service"]:::service
    CDN["CDN<br/>edge chunk cache"]:::edge
    LOCAL[("Encrypted Local Cache<br/>on device")]:::data
    BUS[["Play Event Bus"]]:::async
    LIB[("Library Store")]:::data

    PHONE -->|"1. Request offline lease"| LIC
    LIC -->|"2. Device bound key and expiry"| PHONE
    PHONE -->|"3. Download all chunks"| CDN
    PHONE -->|"4. Store encrypted"| LOCAL
    PHONE -->|"5. Replay queued events on reconnect"| SYNC
    SYNC -->|"6. Publish with idempotency keys"| BUS
    SYNC -->|"7. Push playlist mutations"| LIB

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Step-by-step flow:

  1. User taps download on an album. The client requests a lease per track, naming its device id.
  2. Licence Service checks the subscription and the device cap, then returns a content key wrapped to that device plus an expiry, typically 30 days.
  3. The client downloads every chunk of the chosen rendition from the CDN, using the same immutable URLs as streaming, so a popular album’s chunks are already warm at the edge.
  4. Chunks are written to the encrypted local cache. Playback decrypts chunk by chunk into the decoder, never to disk.
  5. Offline plays append events to a local durable queue, each with a client-generated eventId. On reconnect the client hands the batch to the Sync Service.
  6. Sync Service publishes those events to the play event bus - the eventId makes a retried batch harmless, so duplicates are dropped downstream rather than paid twice - and pushes offline playlist mutations as the same move and add operations the online path uses, keys and all.

Why an expiring licence rather than a permanent one. A licence with no expiry turns a one-month subscription into a permanent library. Periodic revalidation ties continued offline access to a continued subscription, and it gives us a revocation path when a track leaves the catalog for rights reasons. The cost is real: a device offline past expiry loses its downloads, which is why the window is weeks rather than days and the client nags well in advance.

What this leaves broken. Client-side encryption on a device the user controls is a speed bump, not a wall, and Deep Dive 3 is honest about that. The device also now holds listening data that the royalty system has not seen yet, which widens the window in which a play can be lost. Deep Dive 2 picks which direction we bias when that happens.


10. Technology Choices

Tier What it stores Access pattern Primary pick Alternatives
Audio blob store Chunked renditions, every bitrate Write once, read via CDN S3 GCS, Azure Blob, Ceph
CDN Audio chunks at the edge Byte-range reads, very high volume CloudFront Cloudflare, Fastly, Akamai
Catalog store Tracks, albums, artists, rights, markets Read-heavy, relational, joins and constraints Postgres MySQL, Vitess, Spanner
User library and playlists Saved tracks, playlist items with rank keys Per-user partitioned, write-heavy, no cross-user joins Cassandra ScyllaDB, DynamoDB, Bigtable
Play-event bus Raw play and heartbeat events Append-only, replayable, multiple consumers Kafka Kinesis, Redpanda, Pulsar
Stream processor Windowed aggregation and dedup state Continuous, keyed, stateful Flink Kafka Streams, Spark Structured Streaming
Live counters Approximate play and listener counts Very high write, sub-ms read Redis Memcached, DynamoDB with DAX
Royalty warehouse Deduplicated play ledger and daily rollups Batch scans, re-runnable, auditable BigQuery Snowflake, Redshift, Iceberg on S3
Search index Track, album, artist text and facets Fuzzy full-text, typo tolerant, faceted Elasticsearch OpenSearch, Vespa, Solr
Licence store Device leases, key wrapping, device caps Point reads and writes by device Postgres plus Redis DynamoDB, with a managed KMS for key wrapping

Why pre-transcode instead of transcoding on demand. The whole case rests on the asymmetry between encode cost and play count. Encoding one 3.5-minute track into four renditions is a few seconds of CPU, paid once. Serving it is paid every time anybody listens. Transcoding on demand loses on three separate counts. Latency: a few seconds of CPU inside a 200 ms budget is not a tuning problem, it is a category error. Caching: a transcoded-on-demand response is generated per request, so the CDN has nothing stable to cache and every listener pays the full origin cost - which throws away the biggest single win available in this design. Economics: storage for the extra renditions is the cheapest thing in this design, about 21 MB per track, 2.1 PB across the catalog, call it $50K/month at commodity object-storage rates. That is a rounding error next to the compute bill for re-encoding popular tracks millions of times. Encode once, stream forever, and let immutability do the heavy lifting.

The one case for on-demand transcoding is a rendition so rarely requested that storing it for the whole catalog is wasteful - an unusual codec for a niche device, say. Handle that as a lazily-materialized rendition: transcode on first request, write the result to object storage under a stable key, and serve every later request from cache like any other rendition. The cache story stays intact and the encode cost is paid once rather than per request.

Why Cassandra for the user library but Postgres for the catalog. These two tiers look similar and behave nothing alike.

The catalog is small, highly relational, and overwhelmingly read. 100M tracks with their albums, artists, contributor credits, rights holders and per-market availability is tens of gigabytes - it fits on one well-provisioned Postgres box with replicas, and most of the hot set fits in RAM. The queries are real queries: tracks joined to albums joined to artists, filtered by market availability, with foreign keys that must hold because an orphaned track is a track nobody can pay for. Writes run at about 1/sec. Postgres gives us joins, constraints, and transactions at a volume where none of them cost anything, and the 10⁡-to-1 read ratio means caching absorbs the read load long before sharding becomes interesting.

The user library is the mirror image. 600M users each with saved tracks and playlists is billions of rows, and it grows with the user base rather than the catalog. Every access pattern starts with a user id - show me my playlists, show me this playlist’s items in order - and there is no query that spans users. Writes are frequent, spiky, and per-user. Partition by user_id or playlist_id, cluster by rank key, and Cassandra answers every read with a single-partition scan on a node that owns the data, scaling out linearly by adding nodes. The things Postgres gives us for free are things this tier does not want: no joins to make, no cross-user transactions, and no appetite for a single primary to be the write bottleneck for 600M people’s libraries.

Put plainly: one small relational dataset that everyone reads, one enormous partitionable dataset that each user writes to alone. Different problems, different stores.


11. Data Modeling

Postgres (Catalog - tracks, albums, artists, rights):

CREATE TABLE artists (
    artist_id UUID PRIMARY KEY,
    name      VARCHAR(300) NOT NULL,
    country   CHAR(2)
);

CREATE TABLE albums (
    album_id          UUID PRIMARY KEY,
    title             VARCHAR(500) NOT NULL,
    primary_artist_id UUID NOT NULL REFERENCES artists(artist_id),
    release_date      DATE,
    label             VARCHAR(200),
    artwork_key       VARCHAR(256)
);

CREATE TABLE tracks (
    track_id     UUID PRIMARY KEY,
    isrc         CHAR(12) UNIQUE NOT NULL,  -- global recording id, the royalty join key
    title        VARCHAR(500) NOT NULL,
    album_id     UUID NOT NULL REFERENCES albums(album_id),
    track_number SMALLINT NOT NULL,
    duration_ms  INTEGER NOT NULL,
    is_playable  BOOLEAN NOT NULL DEFAULT false  -- true only once all renditions exist
);

CREATE TABLE renditions (
    rendition_id  UUID PRIMARY KEY,
    track_id      UUID NOT NULL REFERENCES tracks(track_id),
    codec         VARCHAR(20) NOT NULL,      -- ogg_vorbis, aac, flac
    bitrate_kbps  SMALLINT NOT NULL,
    total_bytes   BIGINT NOT NULL,
    chunk_bytes   INTEGER NOT NULL,          -- fixed chunk size for range requests
    object_prefix VARCHAR(256) NOT NULL
);

-- rights, and the reason a track 404s in one country but plays in another
CREATE TABLE track_markets (
    track_id UUID NOT NULL REFERENCES tracks(track_id),
    market   CHAR(2) NOT NULL,
    PRIMARY KEY (track_id, market)
);

Cassandra (Library and playlists - per-user, ordered):

Table: saved_tracks
  Partition key: user_id
  Clustering key: saved_at DESC, track_id
  Columns: album_id, source

Table: playlist_items
  Partition key: playlist_id
  Clustering key: rank_key ASC, item_id
  Columns: track_id, added_by, added_at
  -- rank_key is a short lexicographic string, so reading a page is a clustering-key
  -- scan already in order - no ORDER BY, no sort at read time. A reorder writes
  -- exactly one row: delete the old rank_key, insert the new one.

Table: playlists_by_user
  Partition key: user_id
  Clustering key: updated_at DESC, playlist_id
  Columns: name, item_count, revision

Play events (the royalty-bearing write path):

Kafka topic: play-events
  Partition key: track_id        -- all events for one track land in order on one partition
  Value fields:
    event_id        UUID      client-generated, the idempotency key
    user_id         UUID
    track_id        UUID
    isrc            CHAR(12)  denormalized so the ledger never needs a catalog join
    device_id       UUID
    started_at      TIMESTAMP client clock
    received_at     TIMESTAMP server clock, authoritative for windowing
    ms_played       INTEGER   how much audio actually reached the decoder
    bytes_delivered BIGINT    cross-check against ms_played
    offline         BOOLEAN   true if replayed from a device queue
    market          CHAR(2)   determines the royalty pool

Cassandra (the deduplicated ledger and the daily rollup):

Table: play_ledger
  Partition key: (track_id, play_date, bucket)
  Clustering key: event_id
  Columns: user_id, isrc, ms_played, market, qualified, received_at
  -- event_id as the clustering key makes the insert naturally idempotent:
  -- writing the same event twice produces the same row, not two rows.
  -- bucket = hash(user_id) mod N exists only to break up hot partitions.

Table: daily_track_plays
  Partition key: (play_date, market)
  Clustering key: track_id
  Columns: qualified_plays COUNTER, total_ms_played COUNTER, unique_listeners BIGINT

On partitioning play events by track and day, and the hot partition it creates. Partitioning by (track_id, play_date) is the natural choice because every question the royalty system asks is β€œhow many qualifying plays did this recording get on this day in this market”. One partition, one scan, done.

It also walks straight into the celebrity problem. A track in a viral moment can take a large share of all plays, and everything for it on that day lands on one partition on one set of replicas - a write hotspot during the event, and afterwards a partition far too large to scan, since Cassandra partitions get unhappy well before they reach millions of rows. The fix is a salt: add bucket = hash(user_id) mod N to the partition key, spreading one track-day across N partitions and N replica sets. Reads pay for it, because counting a day for one track now means N queries instead of one, but that trade is acceptable when the read is a once-a-day batch job and the write is happening 240,000 times a second. Pick N per tier rather than globally - the long tail needs N = 1 and gains nothing from splitting - so keep a small set of known-hot tracks on a high N and promote a track into it when the live counter crosses a threshold.

Access Patterns:

Query Data Source How
Play a track CDN then object store Range request for chunk 0, decoder starts, remaining chunks fetched ahead
Track and album metadata Redis then Postgres Cache-aside on track:{trackId}, near-100% hit rate given 1 write/sec
Open or reorder a playlist Cassandra plus Redis Read is a clustering-key scan of playlist_items by rank key with titles hydrated from the metadata cache; a reorder is one delete and one insert at a key computed between the neighbours
Royalty statement for a month BigQuery Batch scan of the deduplicated ledger grouped by ISRC, market and rights holder. The live UI count is a separate GET plays:{trackId} against Redis
Search a track Elasticsearch Fuzzy multi-field match on title, artist and album, filtered by market availability

12. Deep Dives

Deep Dive 1: Hitting 200 ms to First Sample, and Never Leaving a Gap Between Tracks

Problem: FR1 gets a warm track playing fast and does two things badly. A cold track on the long tail resolves from object storage, which eats most of the budget on its own. And when the current track ends, the next one starts from nothing, producing a gap that is instantly audible on a live album or a DJ mix.

Bad: Fetch the whole file, then play. A 7 MB 256 kbps track over a 5 Mbps link is 11 seconds. Even the 2.5 MB 96 kbps rendition is 4 seconds. This is not a slow version of the right answer, it is 20x over budget, and it gets worse on exactly the connections where the user is least patient.

Good: HTTP range requests. Ask for bytes=0-131071, hand those 128 KB to the decoder, and keep fetching the rest while audio plays. πŸ’‘ A 206 Partial Content response is an ordinary cacheable HTTP response. The CDN can cache ranges independently, so the first chunk of a popular track is warm at the edge even if nobody has ever played it through to the end. That gets the first sample out in roughly one round trip plus 128 KB of transfer - comfortably inside 200 ms on a warm cache. Its limit is that it only solves the current track, from a warm edge. Cold long-tail tracks still pay the origin fetch, and the gap between tracks is untouched.

Great: Three mechanisms, each aimed at a different part of the budget.

  1. Predictive prefetch of the next track, during the current one. The queue is known. Roughly 20-30 seconds before the current track ends, fetch the first chunk of the next one, authorize it, and have the licence token in hand. The network is idle at that point anyway - the current track’s chunks were fetched ahead and the buffer is full - so this costs bandwidth that was already paid for.
  2. First-chunk caching on device for likely-next tracks, not just the queue - the rest of the album they are listening to, the top of their recently-played, the first few tracks of their daily playlists. At 128 KB each, 200 cached first chunks is 25 MB, unremarkable for a music app, and it converts a cold network fetch into a local disk read for a large share of taps.
  3. Rendition selection from measured throughput, not from a device profile. The server recommends a starting rendition. Each delivered chunk gives the client an actual throughput sample. If the first chunk at 256 kbps arrived slower than real time, switch down at the next chunk boundary. Chunks are independent, so switching costs nothing but a different URL.

Gapless playback is a decoder problem, not a network problem. Audio decoders work on frames and need the next frames available before the current ones finish rendering. If the next track’s first chunk arrives when the current track ends, you still get a gap - the decoder has to initialize, parse a header, and prime its buffers, and that is tens of milliseconds of silence in a place where silence is conspicuous. Real gapless needs the next track’s chunk decoded and queued in the audio pipeline ahead of the switch, so the hand-off is a buffer concatenation rather than a new playback session. This is also why encoder padding matters: the ladder has to be encoded so that leading and trailing silence introduced by the codec is recorded in the rendition metadata and trimmed at playback, or every album transition gains a few tens of milliseconds of nothing.

In simple terms: Start playing after the first few seconds have arrived instead of waiting for the whole song. While that song plays, quietly fetch the beginning of the next one and get the decoder ready for it, so the switch happens with no silence in between.

Why prefetch beats a bigger buffer: A bigger buffer on the current track protects against network loss, which is a real but different problem. It does nothing for the two cases that actually hurt - the first tap after app open, and the moment between tracks. Prefetch targets exactly those two and costs bandwidth only for content the user is very likely to play.

The lookup order that results: on-device first-chunk cache, then CDN edge, then a CDN regional tier, then object storage. That middle tier is doing real work and it exists because of the catalog size. Netflix can pre-position a meaningful fraction of 15,000 titles at every edge. We cannot pre-position 100M tracks anywhere. A regional cache holding, say, the top few million tracks turns a long-tail edge miss into a ~20 ms regional hit instead of a ~100 ms origin fetch, and it shields origin from the full miss volume of the tail. See CDN for how the tiers interact.


Deep Dive 2: Counting Plays When the Count Is a Payment

Problem: 50M concurrent streams generate roughly 240K track completions/sec at peak. The count is not a vanity metric on an artist page - it determines how royalty pools are divided, which means it has to survive an audit, be reproducible months later, and never pay twice for one listen. Those are different requirements from β€œshow a plausible number fast”, and trying to satisfy both with one mechanism is where this design usually goes wrong.

Bad: UPDATE tracks SET play_count = play_count + 1 WHERE track_id = ?. Two failures, both fatal. The row becomes a contention point - a track in a viral moment serializes hundreds of thousands of writes a second through one row, and the lock queue becomes the system’s bottleneck. And if the application reads the count, adds one, and writes it back, concurrent plays silently overwrite each other. Lost updates are lost payments, and nothing in the system notices.

Good: Fire events into Kafka and aggregate downstream. The client posts play events, the API writes them to a partitioned topic, and a stream processor (Flink / Kafka Streams) maintains windowed counts per track. Writes are append-only and spread across partitions, so the hotspot disappears. πŸ’‘ See Batch vs Stream for why the same event log can feed both a continuous and a periodic consumer. The limit is honesty about delivery semantics. A fire-and-forget client that retries on timeout produces duplicates. A client that crashes before flushing loses events. Windowed stream state is not something you hand an auditor, and a bug in the aggregation job cannot be undone because the aggregate is all you kept.

Great: Recognize that there are two consumers of this data with incompatible requirements, and build both off one log.

The fast counter feeds the UI: the play count on a track page, β€œ12,431 listening now”, the trending row. It should be cheap and current. A Flink job reads the topic, aggregates in short windows, and writes to Redis. Approximation is entirely acceptable here - nobody can tell 12,431 from 12,438, and unique-listener counts can use a probabilistic structure rather than a set. πŸ’‘ HyperLogLog estimates the number of distinct values in a stream using a fixed few kilobytes regardless of how many values it sees, with a small percentage error. Perfect for β€œhow many people listened”, useless for β€œexactly which plays do we pay”. If this tier loses a window to a crash, the number is briefly wrong and then corrects itself. Nobody is harmed.

The auditable ledger feeds royalties and has the opposite priorities:

  1. Idempotent writes keyed by event_id. The client generates the id once, at the moment the play qualifies, and reuses it across every retry. The ledger uses it as part of the primary key, so writing the same event five times produces one row. This is the single most important property in the whole pipeline - without it, a flaky mobile connection that triggers three retries pays an artist three times, and the duplicate is indistinguishable from genuine repeat listening. See Idempotency.
  2. The log is retained long enough to replay, and reconciled in batch. Keep raw events for weeks in Kafka and indefinitely in the warehouse as an immutable landing table. A daily job recomputes qualifying plays per ISRC per market from that table, compares its output to the previous run and to the streaming aggregates, and flags divergence. When - not if - someone finds a bug in the qualification logic, the fix is to re-run over the original events rather than patch aggregates. The warehouse number is authoritative for payment; the streaming number is authoritative for nothing.
  3. Qualification is decided once, server side, and recorded. The event carries ms_played and bytes_delivered. The pipeline marks qualified = true when ms_played >= 30000 and the delivered bytes are consistent with that duration. Spotify publishes a 30-second threshold for a play to count toward royalties, which is why the field exists at all and why a 29-second skip is deliberately worth nothing. (Loud and Clear)

In simple terms: Keep two separate counters. The one on the screen can be slightly off and fast. The one that pays artists keeps every original event forever, tags each with a unique id so a retry cannot be counted twice, and is recalculated from scratch every night so a mistake can always be fixed by running it again.

Why two pipelines beat one accurate pipeline: Making the live counter exact means synchronous, deduplicated, transactional writes on the hot path at 240K/sec - expensive, and it puts the royalty system’s correctness requirement in the path of a number on a web page. Making the royalty number fast means trusting stream state you cannot audit or replay. The two requirements genuinely conflict. One log, two consumers, different guarantees.


Deep Dive 3: Offline Downloads Without Handing Away the Catalog

Problem: Offline playback requires the audio to be on a device the user fully controls, and the rights agreements require that it not be freely copyable. Those two statements are in permanent tension and no client-side mechanism resolves them completely.

Bad: Write the plain Ogg or AAC file into the app’s storage directory. On a rooted or jailbroken device that is a file copy. On desktop it is a file copy with no rooting required. Someone writes a 30-line script, points it at the download folder, and the entire downloaded library is portable. A design that assumes the local filesystem is private is not a security design.

Good: Encrypt each rendition on device with a content key fetched from a licence server. The audio at rest is ciphertext, so copying the file yields noise. This raises the bar meaningfully - casual copying stops working. Its limit is where the key lives. If the key is stored next to the file, or is the same key for every user, or never expires, then extracting it once breaks everything after it. A single static key is one reverse-engineering session away from being published.

Great: Make the licence device-bound, time-bound, and countable.

  1. Device-bound key wrapping. The content key is delivered wrapped to a per-device key established at registration and held in the platform keystore (Keychain, Android Keystore, or the platform DRM module - Widevine, FairPlay, PlayReady). The wrapped blob is useless on any other device because the unwrapping key never leaves that device’s secure storage.
  2. Expiry with periodic revalidation. Licences carry an expiry, on the order of 30 days. The client revalidates while online, which refreshes it. A device that stays offline past expiry loses playback until it reconnects. This is what ties offline access to an active subscription instead of making a one-month signup a permanent library, and it provides the revocation path when a track’s rights lapse.
  3. A device cap per account, enforced server side. Registering device N+1 forces the user to release one. The cap is the thing that stops an account being resold as a family-of-forty plan, and it has to be enforced where the client cannot see it.
  4. Cache keyed so a copy is worthless, and decryption per chunk into the decoder. The local cache filename is derived from the rendition and device and chunks are encrypted individually, so a copied directory decrypts to noise under any other device’s key - there is no partial win from lifting half the cache. Plaintext audio exists only in the buffer handed to the audio pipeline, so there is never a complete plaintext file sitting on storage waiting to be grabbed.

Be honest about what this is. Client-side DRM is defence in depth, not a guarantee. The device is running code the user controls, the decoder has to see plaintext eventually, and anyone determined enough can capture the analog or digital output. What this design actually achieves is raising the cost from β€œcopy a folder” to β€œreverse-engineer a hardened client and extract keys from a platform keystore”, and giving us revocation, expiry, and device limits as real levers. That is the realistic goal, and it is the one the rights agreements are actually written against. Claiming more than that in an interview is a worse answer than admitting the limit.

In simple terms: The downloaded song is scrambled, and the key to unscramble it is locked to that one phone and stops working after about a month unless the app checks in. Copying the files to another device gives you unplayable junk. A determined attacker can still break it, so this buys deterrence and the ability to revoke access, not perfect protection.


Deep Dive 4: Recommendation and Radio Across a 100M Track Catalog

Problem: A user finishes a playlist and radio has to keep going with tracks they will not skip. The catalog is 100M tracks and most of it is unknown to any given listener. Most of it is also barely played at all, which is the part that makes this hard.

Bad: Globally most-popular for everyone. It is one query and it is wrong for nearly every user - it recommends the same forty tracks to a metal listener and a Carnatic listener, and it is a feedback loop that makes popular tracks more popular and buries everything else. For a catalog where the long tail is the product, this actively destroys the thing people subscribe for.

Good: Collaborative filtering over the co-listen matrix. Build a users-by-tracks matrix from listening history, factorize it, and recommend tracks whose latent vectors sit near the user’s. This genuinely works and it captures taste that no attribute tagging would - it can learn that people who love one obscure post-rock record reliably love a particular ambient album that shares no tags with it.

Its failure is cold start, in both directions. A track uploaded this morning has no co-listen data, so it is invisible to the model and stays invisible, because invisibility prevents the plays that would make it visible. With ~100K new tracks a day, that is a permanently growing blind spot. A new user has no history, so their vector is meaningless and the system falls back to popularity - exactly the bad answer, served to the users whose first impression matters most.

Great: A hybrid, split into candidate generation and ranking, served from a precomputed store.

  1. Two signal families, combined. Collaborative signals from the co-listen matrix, plus content features computed from the audio itself - tempo, key, spectral characteristics, instrumentation, learned embeddings from the waveform - and from text around the track such as editorial descriptions and playlist co-occurrence. Content features need no plays, which is precisely what fixes new-track cold start: a track can be recommended on day one because it sounds like things the user already loves.
  2. Candidate generation, then ranking. Candidate generation pulls a few hundred plausible tracks from a 100M catalog cheaply - nearest neighbours in the embedding space, tracks from co-listened playlists, artists adjacent to the ones the user plays. Ranking then scores those few hundred with a heavier model that can afford to look at context: time of day, current session, device, whether the user is skipping a lot right now. Scoring 100M tracks per request is impossible; scoring 300 is routine.
  3. Precompute per user, serve from a key-value store. The expensive half runs offline on a schedule and writes a ranked list per user to Redis or Cassandra. The serving path on app open is a key lookup, which is what keeps the home screen inside its latency budget. A lighter online re-ranker adjusts for in-session signals - three skips in a row should change the next pick without waiting for the nightly job.
  4. A vector index for similarity, plus explicit exploration. Track embeddings live in an approximate nearest neighbour index (Annoy, FAISS, ScaNN, or a vector database), so β€œmore like this” and radio seeding are nearest-neighbour queries at single-digit milliseconds instead of joins over a play history table. Some radio slots are then reserved for candidates the model is unsure about - without that, the system only ever learns about tracks it already recommends and the cold-start hole never closes. The cost is a slightly higher skip rate; the benefit is a model that keeps discovering the tail.
flowchart LR
    HIST[("Play History<br/>co listen signals")]:::data
    AUDIO[("Audio Features<br/>from waveform")]:::data
    TRAIN["Offline Training<br/>embeddings"]:::async
    VEC[("Vector Index<br/>track embeddings")]:::data
    CG["Candidate Generation"]:::service
    RANK["Ranker"]:::service
    STORE[("Per User Reco Store")]:::data
    APP["Player App"]:::client

    HIST -->|"1. Co listen matrix"| TRAIN
    AUDIO -->|"2. Content features"| TRAIN
    TRAIN -->|"3. Write embeddings"| VEC
    VEC -->|"4. Nearest neighbours"| CG
    CG -->|"5. Few hundred candidates"| RANK
    RANK -->|"6. Write ranked lists"| STORE
    APP -->|"7. Key lookup on open"| STORE

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

In simple terms: Learn what goes together from what people actually listen to, and also learn what songs sound like, so a brand-new track can be suggested before anyone has played it. Shortlist a few hundred candidates with a cheap method, then carefully rank just those. Do the slow part overnight and store the answer, so opening the app is one lookup.

Why candidate generation and ranking are separate stages: They have opposite cost profiles. Candidate generation must be sublinear in catalog size, which rules out any model that needs to look at a track’s full feature set. Ranking must be expressive enough to use session context, which makes it too expensive to run at catalog scale. Splitting them lets each use the algorithm that suits it, and it makes the system debuggable - when a bad recommendation shows up, you can ask whether it was a bad candidate or a bad rank.


13. Design Self-Audit

Question Answer
What happens on a CDN miss for a cold long-tail track? Edge misses, regional tier is checked, origin fetch if that misses too. Worst case is roughly 100-150 ms of origin latency inside a 200 ms budget, so a true cold start is tight but not broken - and only the first 128 KB chunk is on the critical path. The mitigations are the regional tier, serving the first chunk of a cold track at 96 kbps and switching up once playing, and the on-device first-chunk cache. With 100M tracks we accept that cold tail starts are slower than warm ones; pretending otherwise would mean pre-positioning a catalog that does not fit anywhere.
Can a play be double-counted or lost, and which way do we bias? Both are possible and we bias toward losing rather than double-paying. event_id as part of the ledger’s primary key makes duplicates idempotent, so the retry path is safe. Loss is the residual risk: a device that is wiped before reconnecting takes its queued offline events with it. We accept a small undercount because an overcount is money paid for listening that did not happen, which is an audit finding, while an undercount is a reconciliation discrepancy we can measure and bound.
Two devices reorder the same playlist while offline. What happens? Both converge to a deterministic order and neither user necessarily gets what they intended. Each move is an independent write of one rank key, so there is no lost-update storm, but two moves into the same gap produce an interleaving neither device chose. We detect it via the playlist revision counter, and on a revision mismatch the client refetches and shows the merged order rather than silently claiming success. Conflict-free collaborative editing is below the line for a reason - doing it properly needs a CRDT, which is a different design.
What does egress actually cost at this scale? Roughly 1 exabyte/month. At commodity CDN rates of a cent or two per GB that is an eight-to-nine-figure annual bill, which is why the Open Connect argument applies to audio too even though each stream is 25x smaller than video. Mitigations in priority order: maximize edge hit rate (immutable chunk URLs, long TTLs), serve the lowest rendition that the user cannot distinguish on their device, negotiate ISP peering, and own appliances for the highest-volume markets.
Can royalty figures be recomputed from scratch after a bug? Yes, and this is a hard requirement rather than a nice property. Raw events land immutably in the warehouse and are retained indefinitely; Kafka retains them for weeks. Every derived number - qualification, per-ISRC daily totals, market pools - is a pure function of the raw landing table. Fixing a qualification bug means changing the job and re-running the affected months, then diffing against what was previously paid. Nothing in the royalty path mutates in place.
Is there a dedicated search index? Yes, Elasticsearch, and it is not optional. Users search by misspelled titles, partial artist names, and lyrics fragments across 100M tracks. Postgres cannot do typo tolerance or relevance ranking at that size. The index is fed from the catalog’s change stream, so a newly ingested track is searchable within seconds of becoming playable. See Search Autocomplete for the prefix-query side of this.

14. Core Flows

Flow 1: Pressing Play

sequenceDiagram
    participant U as User
    participant APP as Player App
    participant PS as Playback Service
    participant LIC as Licence Service
    participant CDN as CDN Edge
    participant OBJ as Object Store

    U->>APP: Tap a track
    APP->>PS: POST playback start with trackId
    PS->>PS: Check subscription and market licence
    PS->>LIC: Mint short lived stream token
    LIC-->>PS: Token and content key reference
    PS-->>APP: Rendition list base url token and idempotency key
    APP->>CDN: GET audio rendition Range bytes 0 to 131071
    alt Edge cache hit
        CDN-->>APP: 206 first chunk from edge
    else Edge cache miss
        CDN->>OBJ: Fetch chunk range
        OBJ-->>CDN: Chunk bytes
        CDN-->>APP: 206 first chunk and cache it
    end
    APP->>APP: Decode first frames and start audio
    Note over U,APP: Target P95 under 200 ms to first sample
    APP->>CDN: Fetch remaining chunks ahead of playback
    APP->>CDN: Prefetch first chunk of next queued track

Walkthrough:

  1. The tap produces exactly one control-plane call. Everything after it is plain cacheable HTTP against the CDN, which is what lets the data plane scale independently of the service tier.
  2. Authorization happens entirely server side - subscription tier and market come from the account, never from a client field - and the returned token is scoped to this rendition and expires in minutes.
  3. The response carries the whole rendition ladder, so the client can switch quality later without another control-plane round trip.
  4. The first range request is the only thing on the critical path. 128 KB is about 8 seconds of audio at 128 kbps, so an edge hit means sound in roughly one round trip; a miss fills the edge from origin and the next listener in that region gets the hit. Remaining chunks are then fetched ahead of the playhead to absorb network variation, and before the track ends the client prefetches and authorizes the next one so the hand-off is gapless.

Non-obvious failure: The stream token expires mid-track - a 5-minute token on a 9-minute track, or a track paused for an hour and then resumed. If the client treats a 403 on chunk 40 as a playback error, the user gets a dead player two-thirds of the way through a song. The client has to refresh the token in the background before expiry and retry the chunk once on 403, and the token lifetime has to exceed the longest plausible track. The same problem appears with a paused session: on resume, revalidate before issuing the next range request rather than discovering the failure at the decoder.

Flow 2: The Play Event Pipeline

sequenceDiagram
    participant APP as Player App
    participant ING as Event Ingest
    participant K as Kafka play events
    participant FL as Flink job
    participant R as Redis counters
    participant LED as Play Ledger
    participant WH as Warehouse

    APP->>APP: Playback crosses 30 seconds
    APP->>APP: Append event with generated eventId
    APP->>ING: POST plays batched
    alt Network available
        ING->>K: Produce events keyed by trackId
        ING-->>APP: 202 accepted with counts
    else Offline
        APP->>APP: Hold batch in durable local queue
        APP->>ING: Replay same eventIds on reconnect
    end
    K->>FL: Consume partition stream
    FL->>R: Increment approximate counters
    FL->>LED: Upsert keyed by trackId playDate and eventId
    K->>WH: Land raw events immutably
    WH->>WH: Nightly recompute qualifying plays per ISRC
    WH->>WH: Diff against previous run and flag divergence

Walkthrough:

  1. The client decides a play has occurred when playback crosses 30 seconds of actual decoded audio, and generates the eventId at that moment. Generating it earlier or later breaks idempotency across retries.
  2. Events are batched rather than sent one per play, which matters for a long background session on a mobile connection, and Ingest produces them to Kafka keyed by track_id for per-track ordering across partitions.
  3. Offline, the batch sits in a durable local queue. On reconnect the same eventIds are replayed, so the retry is free of consequence.
  4. Flink reads the topic once and writes to two places with two different contracts: approximate counters in Redis for the UI, idempotent upserts into the ledger for money.
  5. Raw events also land in the warehouse untouched - that landing table is the system of record, not the ledger and certainly not Redis - and the nightly job recomputes qualifying plays from it, diffs against the previous run, and alerts on unexpected divergence. The warehouse number is what gets paid.

Non-obvious failure: Flink’s checkpoint is lost and the job restarts from an earlier offset, so a window of events is reprocessed. The two outputs behave completely differently, and that is by design. The ledger absorbs it silently because event_id is part of its key - reprocessing writes the same rows. Redis does not: counters are blind increments, so the live count jumps by however much got replayed. The fix is not to make Redis idempotent, it is to accept the drift and have the nightly job overwrite the Redis baseline from the warehouse figure. A visibly wrong play count that self-corrects within a day is a tolerable bug. A double royalty payment is not.


15. Final Architecture

flowchart LR
    APPS["Phone Desktop and Speaker"]:::client
    GW["API Gateway"]:::edge
    CDN["CDN<br/>chunk delivery"]:::edge
    PS["Playback Service"]:::service
    PLS["Playlist Service"]:::service
    LIC["Licence Service"]:::service
    SYNC["Sync Service"]:::service
    SRCH["Search Service"]:::service
    TP["Transcode Pipeline"]:::service
    RECO["Reco Serving"]:::service
    KF[["Kafka<br/>play events"]]:::async
    FL["Flink aggregation"]:::async
    OBJ[("Object Store<br/>audio chunks")]:::data
    CAT[("Postgres<br/>catalog and rights")]:::data
    LIB[("Cassandra<br/>library and playlists")]:::data
    LED[("Cassandra<br/>play ledger")]:::data
    RDS[("Redis<br/>live counters and cache")]:::data
    ES[("Elasticsearch<br/>search index")]:::data
    WH[("Warehouse<br/>royalty reconciliation")]:::data

    APPS --> GW
    APPS -->|"audio chunk ranges"| CDN
    CDN --> OBJ
    GW --> PS
    GW --> PLS
    GW --> LIC
    GW --> SYNC
    GW --> SRCH
    GW --> RECO
    PS --> CAT
    PS --> RDS
    PLS --> LIB
    PLS --> CAT
    SRCH --> ES
    RECO --> RDS
    TP -->|"write renditions"| OBJ
    TP -->|"mark playable"| CAT
    CAT -->|"change stream"| ES
    APPS -->|"play events"| KF
    SYNC -->|"replay offline events"| KF
    KF --> FL
    FL -->|"approximate counts"| RDS
    FL -->|"idempotent upsert"| LED
    KF -->|"immutable landing"| WH
    LED -->|"nightly diff"| WH

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

How it works end-to-end (read path):

  1. Client opens the app β€” recommendation rows come from a precomputed per-user store in Redis, so the home screen is key lookups rather than computation, and Elasticsearch answers any search from an index fed by the catalog’s change stream
  2. User taps a track β€” Playback Service authorizes against subscription and market, then returns the rendition ladder plus a short-lived token; playlists come from a Cassandra clustering-key scan ordered by rank key, hydrated from the metadata cache
  3. Audio arrives from the CDN β€” immutable chunk ranges, edge first, regional tier second, object storage only on a full miss; meanwhile the next track is prepared with its first chunk prefetched, licence obtained and decoder primed, so the transition is gapless

How it works end-to-end (write and ingest path):

  1. Labels deliver masters β€” the Transcode Pipeline fans each one into the 96, 128, 256 and 320 kbps ladder, chunks it, writes to object storage, and marks the track playable only once every rendition exists
  2. Playlist edits write a single row β€” Playlist Service computes a rank key between the neighbours the client named and bumps the playlist revision; downloads take a device-bound lease from the Licence Service and land in the device’s encrypted cache
  3. Plays become events β€” every qualifying play, online or replayed from an offline queue, enters Kafka with a client-generated event_id, and Flink fans it out to approximate Redis counters and idempotent ledger rows
  4. The warehouse is the system of record β€” raw events land immutably and the nightly job recomputes royalty figures from scratch, diffing against the previous run

Key Technologies

Term What it is
HTTP Range Request A GET with a Range header asking for a byte slice. Returns 206 Partial Content. Lets playback start after one chunk instead of a whole file, and stays cacheable at the edge.
Pre-transcoded Bitrate Ladder A fixed set of renditions per track (96, 128, 256, 320 kbps) encoded once at ingest. Avoids per-request CPU and gives the CDN stable objects to cache.
Gapless Playback Having the next track’s audio decoded and queued before the current one ends, plus trimming codec padding, so album transitions contain no silence.
Lexicographic Ordering Key A short sortable string stored per playlist item. A reorder writes one row instead of renumbering the list.
Idempotency Key A client-generated event_id carried on every retry. Used as part of the ledger’s primary key so a duplicate delivery produces one row, not two payments.
Device-bound Licence An expiring grant containing a content key wrapped to one device’s keystore. Enables offline playback while keeping revocation and device caps enforceable.
Approximate Nearest Neighbour Index A vector index (Annoy, FAISS, ScaNN) that finds similar embeddings in single-digit milliseconds. Powers radio seeding and β€œmore like this”.
Flink A stateful stream processor. Reads the play-event log once and fans out to a fast counter and an auditable ledger with different guarantees. (Apache Flink)

What’s Expected at Each Level

This section helps you calibrate your depth. You don’t need to cover everything - just know what’s expected for your level.

Mid-level

Lay out the upload β†’ transcode β†’ object storage β†’ CDN β†’ client path and explain why the audio does not live in the database. Know that playback starts from a range request rather than a whole-file download, and be able to say roughly why 7 MB over a mobile link cannot meet a 200 ms budget. Propose a metadata store for the catalog and a separate store for user libraries. Recognize that incrementing a counter row per play is a hotspot, even before being prompted.

Senior

Argue pre-transcoding versus on-demand transcoding with actual numbers on both sides, including the cache implications. Explain the playlist ordering problem and reach for a fractional or lexicographic key rather than integer positions. Separate the live play counter from the royalty ledger and justify why one event log feeding two consumers beats one pipeline trying to satisfy both. Discuss prefetch and gapless playback as a decoder-buffer problem, not just a network one. Know why partitioning play events by track and day creates a hot partition for a viral track and how a bucket salt trades read cost for write spread.

Staff+

Address the egress economics at exabyte scale and why the Open Connect argument survives the 25x drop in per-stream size. Design the three-tier cache knowing you cannot pre-position a 100M-track catalog, and be specific about what the regional tier buys on the long tail. Be rigorous about royalty correctness: idempotency keys, immutable raw event retention, nightly recomputation, and the ability to re-run a month after a qualification bug. Take a defensible position on offline DRM as deterrence plus revocation rather than protection. Handle recommendation cold start in both directions, and explain why content features from the waveform are the only thing that makes a track recommendable on the day it is uploaded.


🎯 Key Takeaways



Understand the building blocks used in this design:

Discussion

Newest first
You

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access