Designing Netflix / YouTube - Video Streaming Platform
Difficulty: Advanced Topics: Video Encoding Pipeline, CDN, Adaptive Bitrate Streaming, Recommendation Engine, Content Catalog Asked at: Netflix, YouTube, Disney+, Amazon Prime Video, Hotstar Prerequisites:CDN and Message Queues
1. Understanding the Problem
Netflix is a video streaming platform that lets content creators upload video, encodes it into dozens of format/resolution/bitrate combinations, distributes it globally via CDN, and streams it to users with adaptive quality based on their connection speed. The system serves 200M+ concurrent users, each streaming a unique video at a unique bitrate. The hardest parts: encoding pipeline (converting one upload into 100+ playable variants), CDN distribution (getting the right bits to the right edge node before the user needs them), and recommendations (deciding what to show each user from a catalog of 15,000+ titles).
2. Naive First Cut
flowchart LR
Client["Smart TV App"]:::client
API["API Server"]:::service
DB["Postgres DB"]:::data
Files["File Server"]:::data
Client --> API
API --> DB
API --> Files
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
| Color | Meaning |
|---|---|
| π Purple | Client |
| π΅ Blue | Edge / Gateway |
| π’ Green | Service |
| π£ Purple | Async (Workflow / Queue) |
| π‘ Yellow | Data store |
| π΄ Pink | External |
How this breaks:
- Single file server canβt serve 200M concurrent streams - bandwidth alone would be 200 Tbps
- No encoding pipeline - raw 4K master files are too large (50GB per hour of content) to stream directly
- No adaptive quality - users on 3G get buffer wheels, users on fiber get potato quality
- Centralized serving means 500ms+ latency for users far from origin (India streaming from US-West)
- No personalization - everyone sees the same homepage, engagement drops
- Single DB canβt handle 200M concurrent session lookups + play position tracking
The rest of the doc evolves this into a globally distributed streaming platform with intelligent encoding, edge delivery, and personalized recommendations.
3. Prior Art Weβre Drawing From
- Netflix Open Connect - Netflixβs own CDN: custom hardware appliances (OCAs) deployed inside ISP networks. Serves 95%+ of traffic from within the ISP, eliminating internet transit costs. Pre-positions content overnight based on popularity predictions. (Netflix Open Connect overview)
- YouTube Vitess (Video Metadata at Scale) - Sharded MySQL via Vitess for video metadata, serving billions of queries/day. Demonstrates that video catalogs need a horizontally scalable metadata tier. (Vitess.io)
- Netflix Zuul (API Gateway) - Edge gateway handling auth, routing, canary deployments, and load shedding for 200M+ users. Shows why a smart edge layer is critical for streaming. (Netflix Zuul blog)
- Netflix Cosmos (Encoding Pipeline) - Microservice-based media processing platform replacing the monolithic encoder. Uses per-title encoding to optimize bitrate per scene complexity. (Netflix Tech Blog - Cosmos)
- Netflix Recommendations (Two-Stage Ranking) - Candidate generation via collaborative filtering, then re-ranking via deep learning. Row-based homepage layout driven by ML. (Netflix Tech Blog - Recommendations)
In simple terms: Instead of encoding every video the same way, analyze each scene separately. Action scenes get more bits (complex), talking-head scenes get fewer bits (simple). Result: consistent visual quality throughout without wasting bandwidth on easy scenes.
4. Functional Requirements
Core (Top 3)
- Upload and encode video content - content team uploads master files; system encodes into multiple resolutions, bitrates, and codecs
- Stream video with adaptive bitrate - users watch content with quality that adjusts to their bandwidth in real-time
- Personalized recommendations - homepage shows titles ranked by predicted relevance for each user
Below the Line
- User profiles and parental controls
- Offline downloads
- Subtitles and multi-language audio
- Watch party (synchronized playback)
- Content licensing and regional availability
5. Non-Functional Requirements
Core
| NFR | Target |
|---|---|
| Start Playback | Time to first frame < 2 seconds globally |
| Buffer-free | 99% of streams play without rebuffering |
| Scale | 200M+ concurrent streams during peak (Sunday evening) |
| Availability | 99.99% - downtime during prime time is front-page news |
Below the Line
- Encoding pipeline completes within 4 hours of upload (not user-facing latency)
- Recommendation freshness: incorporate new viewing signals within 1 hour
- Multi-region disaster recovery for control plane
6. Scale Estimation (Back-of-Envelope)
- Users: 250M subscribers, ~40M concurrent streams at peak (Sunday evening globally)
- Bandwidth: ~200 Tbps aggregate at peak (40M streams Γ ~5 Mbps average adaptive bitrate)
- Egress volume: ~200PB/day, so roughly 6 exabytes/month of video leaving our infrastructure
- Write QPS: 40M active streams heartbeating every 10s = ~4M position writes/sec
- Read QPS: 250M subscribers Γ ~3 sessions/day = ~750M home screen loads/day, so ~9K/sec average and ~25K/sec at peak
- Storage: 15,000 titles Γ ~30 encoding variants Γ ~3GB per variant β 1.4PB, call it ~2PB once audio tracks and subtitles are counted
Note the shape of this system before designing it: the catalog is small β a couple of petabytes, which is nothing β while the egress is enormous. 6 exabytes a month is the number that decides the architecture, and it is why the interesting problems here are about delivery and encoding efficiency rather than about databases.
7. Core Entities
- Title - movie or series metadata: name, genre, cast, rating, licensing regions
- Video Asset - physical encoding: resolution, bitrate, codec, segment manifest
- Encoding Job - workflow state for converting a master into playable assets
- User Profile - preferences, watch history, ratings, maturity settings
- Playback Session - active stream: title, position, quality level, device info
- Recommendation - pre-computed ranked list of titles for a user
8. API / System Interface
POST /api/v1/content/upload
Body: { titleId, masterFileUrl (S3 pre-signed), metadata }
Response: { encodingJobId, status: "QUEUED", estimatedCompletion }
Auth: Internal service token (content team only)
GET /api/v1/browse/home?profileId=<id>
Response: { rows: [{ rowTitle, titles: [{ titleId, thumbnail, matchScore }] }] }
Note: Personalized homepage with ranked rows
POST /api/v1/playback/start
Body: { titleId, profileId, deviceType, preferredQuality }
Response: { sessionId, manifestUrl (HLS/DASH), licenseUrl (DRM), resumePosition }
Note: Returns manifest URL pointing to CDN
POST /api/v1/playback/heartbeat
Body: { sessionId, currentPosition, currentBitrate, bufferHealth }
Response: { continue: true }
Note: Sent every 10s to track progress and detect abandonment
GET /api/v1/search?q=<query>&filters=genre,year
Response: { results: [{ titleId, title, matchType, thumbnail }] }
9. High-Level Design
FR1: Upload and Encode Video Content
When a studio delivers a new movie to Netflix, it arrives as a massive master file (ProRes 4K, 50-100GB). We need to convert it into playable formats: multiple resolutions (240p to 4K), multiple bitrates per resolution, multiple codecs (H.264, H.265, VP9, AV1), and package them into streamable segments (HLS for Apple, DASH for everything else). This is a multi-hour pipeline.
In simple terms: Netflix receives a 50-100GB raw movie file. It needs to be converted into dozens of formats (different resolutions, codecs, devices). This takes hours of processing.
Build the simplest thing that turns a master file into something a player can stream. An ingest step, a queue, a pool of workers running a fixed bitrate ladder, storage for the output, and one database for the metadata. No workflow engine and no CDN yet β those answer non-functional requirements and belong in deep dives.
New components:
- Ingest Service - Receives the upload notification, validates the master file, and enqueues one encoding task per output variant. The βfront deskβ for new content.
- Encoding Queue - Holds the pending variants. Encoding a 100GB master takes hours, so this cannot live inside a request β the queue is the minimum needed to run it out of band.
- Encoding Workers - GPU machines that do the transcoding. Each worker takes one resolution and codec combination and emits 4-second segments.
- Object Storage (S3) - Stores master files and all encoded segments, as
titles/{titleId}/assets/{resolution}_{bitrate}/segment_000.ts - Content Catalog DB - What exists and who may watch it: available encodings, licensed regions, licensing windows.
flowchart LR
Studio["Content Studio"]:::external
IS["Ingest Service"]:::service
S3M[("S3 Masters")]:::data
Q["Encoding Queue"]:::async
EW["Encoding Workers GPU"]:::service
S3E[("S3 Encoded Segments")]:::data
CAT[("Catalog DB")]:::data
Studio -->|"1. Upload raw master"| IS
IS -->|"2. Store master file"| S3M
IS -->|"3. Enqueue one task per variant"| Q
Q -->|"4. Worker takes a task"| EW
EW -->|"5. Write segments and manifest"| S3E
EW -->|"6. Mark variant complete"| CAT
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0
Step-by-step flow:
- Studio uploads the master to an S3 bucket using a pre-signed, resumable upload, since a 100GB file will not survive a single HTTP request
- An S3 event notification wakes the Ingest Service, which validates the file: expected codec, expected resolution, audio tracks present
- Ingest Service expands one fixed ladder into tasks β 240p through 4K, each at a preset bitrate, across H.264, H.265 and AV1 β and enqueues roughly 30 of them
- Each Encoding Worker takes a task, pulls the master, transcodes with FFmpeg / x265 / SVT-AV1, and writes 4-second segments
- Workers upload segments to S3 and record the completed variant in the Catalog DB
- Once every variant for a title is present, the title is marked playable and manifests are written: HLS
.m3u8for Apple devices, DASH.mpdfor everything else
Why encode ~30 variants instead of one? Because FR2 has to serve a 4K television and a phone on 3G from the same catalog, and there is no single file that suits both. The number of variants is forced by device and network diversity. What bitrate each one runs at is a choice, and that choice is where the next section goes.
What we have deliberately left broken. This produces playable content and it is a reasonable place to start. Two holes:
- The ladder is the same for every title. A hand-drawn cartoon and a handheld action sequence both get 1080p at 5 Mbps, which is far more than one needs and not enough for the other. Since we serve 6 exabytes a month, bits spent unnecessarily are the single largest cost in the system. That is Deep Dive 1.
- A failure late in a multi-hour job loses the whole job. A worker that dies four hours into a 4K AV1 encode leaves a partial variant and a queue entry that has already been consumed. There is no checkpoint and no per-step retry, so recovery means starting over. Deep Dive 1 addresses this alongside the ladder, since both are properties of the encoding pipeline.
And one thing that is not broken yet but will be: step 5 puts the segments in one S3 region. Nothing here moves them closer to viewers, which is Deep Dive 2βs problem, and nothing predicts which titles should be moved before anyone presses play, which is Deep Dive 5βs.
FR2: Stream Video with Adaptive Bitrate
When a user hits βPlay,β we have to check they are allowed to watch it, hand them something they can play, and remember where they got to so they can resume.
In simple terms: You press play, we confirm youβre allowed to watch this in your country, we give your player the file list, and it starts fetching.
The delivery protocol is HLS on Apple devices and DASH everywhere else.
π‘ Both work the same way: the video is a series of 2-4 second segments, and a manifest file lists every available quality level. The player fetches segments one at a time over ordinary HTTP, which means any web server or cache can serve video without understanding video.
New components:
- Playback Service - Authorizes the play: valid subscription, title licensed in this region. Returns the manifest URL and the resume position.
- DRM License Server - Issues the decryption license for the session. Widevine on Android and Chrome, FairPlay on Apple, PlayReady on Windows. This is a licensing obligation, not an optimization, so it is in the base design.
- Playback Position table - One row per profile per title holding the last known position, written from the playerβs periodic heartbeat.
flowchart LR
TV["Smart TV"]:::client
GW["API Gateway"]:::edge
PS["Playback Service"]:::service
DRM["DRM License Server"]:::external
POS[("Playback Positions")]:::data
S3[("S3 Encoded Segments")]:::data
TV -->|"1. POST playback start"| GW
GW -->|"2. Auth and forward"| PS
PS -->|"3. Issue DRM license"| DRM
PS -->|"4. Read resume position"| POS
PS -->|"5. Return manifest URL"| TV
TV -->|"6. Fetch segments in order"| S3
TV -->|"7. Heartbeat position"| POS
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0
Step-by-step flow:
- User hits Play β
POST /playback/startwith the titleId - Playback Service checks the subscription is active and the title is licensed in the userβs region. Region comes from the account, not from a client-supplied field
- It requests a DRM license for this device and session
- It reads the last saved position for this profile and title
- It returns the manifest URL, the DRM license, and the resume position
- The player reads the manifest, picks a quality level, and fetches segments in order from S3 for the rest of the film
- Every 10 seconds the player reports its position so a later session can resume
Why hand the player a manifest instead of streaming from the server? Because segment-over-HTTP puts the delivery on plain, cacheable GETs. The server never holds a per-viewer stream, which is what makes 40M concurrent viewers a bandwidth problem rather than a connection problem β and bandwidth problems can be solved by putting copies closer to people.
What we have deliberately left broken. This plays video correctly and it is the most expensive and most fragile part of the design as written. Three holes:
- Every byte comes from one S3 region. At ~6 exabytes/month of egress, that is both a latency disaster for anyone not near that region and, at list-price egress, a bill that dwarfs everything else in the system. Working out what it actually costs and what to do instead is Deep Dive 2.
- Step 6 picks a quality level and keeps it. When someone walks out of WiFi range mid-scene, the player is still asking for 4K segments it can no longer download in time, so playback stalls and waits rather than degrading. That is Deep Dive 3.
- There is no failure path at all. If a segment request fails, the player has nothing to fall back to β no alternate source, no retry policy, no graceful degradation. Deep Dive 6 deals with what the player should do when the thing serving it stops answering.
FR3: Personalized Recommendations
When users open Netflix, they see a homepage with rows of titles (βBecause you watched X,β βTrending Now,β βTop 10 in Indiaβ). Each userβs homepage is different. Recommendation quality directly drives engagement and retention - Netflix estimates their rec system is worth $1B/year in reduced churn.
In simple terms: Your Netflix homepage shows different movies than your friendβs - based on what youβve watched before. This personalization needs to happen in under 200ms when you open the app.
New components:
- Recommendation Service - Answers
GET /browse/home. Builds the rows for this profile by looking at what they have watched and finding catalog titles that resemble it. - Viewing History table - What each profile has played, how far through it got, and when. FR2βs heartbeat already writes the raw material; this is the same data read the other way round.
No new datastore. The catalog is already in the Catalog DB from FR1 and viewing history comes out of FR2, so a first personalized home screen is a query over things we have.
flowchart LR
TV["Client App"]:::client
GW["Gateway"]:::edge
RS["Recommendation Service"]:::service
HIST[("Viewing History")]:::data
CAT[("Catalog DB")]:::data
TV -->|"1. GET browse home"| GW
GW -->|"2. Forward to reco svc"| RS
RS -->|"3. Read this profile history"| HIST
RS -->|"4. Find similar titles"| CAT
RS -->|"5. Return ranked rows"| TV
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Step-by-step flow:
- User opens the app β
GET /browse/home?profileId=Xhits the Recommendation Service - It reads the profileβs viewing history and derives a rough preference from it: which genres, which actors, which directors show up most
- It queries the Catalog DB for titles matching those attributes that the profile has not already finished, filtered to what is licensed in their region
- It assembles the familiar rows β βBecause you watched X,β βMore in Thrillers,β βTrendingβ from raw play counts β ranks each row by how strongly it matched, and returns the top titles per row
- Every row is computed fresh, on this request, for this profile
Why compute per profile and not per account? Because a household shares one subscription and does not share taste. Recommending based on account-level history produces a home screen that is wrong for everyone in the house, which is worse than no personalization at all.
What we have deliberately left broken. This is genuinely personalized and it is the slowest and shallowest version of that:
- All the work happens while the user waits. Reading a full viewing history and scanning the catalog for attribute matches, per row, is hundreds of milliseconds against a 200ms budget β and it happens ~25K times a second at peak, repeating almost identical work for the same user several times a day.
- Attribute matching is a weak signal. βYou watched a thriller, here are thrillersβ cannot discover that people who liked this particular thriller tend to love a documentary that shares none of its attributes. The interesting recommendations are exactly the ones this approach cannot see.
- A brand-new profile has no history, so step 2 returns nothing and the home screen collapses to raw popularity.
All three are Deep Dive 4: moving the computation offline, replacing attribute matching with learned similarity, and having an answer for a profile we know nothing about.
10. Technology Choices
| Tier | Purpose | Stores | Access Pattern | Primary | Alternatives |
|---|---|---|---|---|---|
| Object Storage | Master files and encoded segments | Raw uploads + HLS/DASH segments | Write once, read via CDN | S3 | GCS, Azure Blob |
| CDN | Video segment delivery | Encoded segments cached at edge | Ultra-high throughput reads | Open Connect (custom) | CloudFront, Akamai, Fastly |
| Content Catalog DB | Title metadata, episodes, licensing | Structured catalog data | Read-heavy, complex queries | Postgres (Vitess-sharded) | CockroachDB, Spanner |
| User Activity Store | Watch history, progress, ratings | Time-series user events | High-write (play events), read for recs | Cassandra | DynamoDB, ScyllaDB |
| Encoding Queue | Encoding job orchestration | Job state, dependencies, retries | Workflow orchestration | Temporal or Cadence | Step Functions, Airflow |
| Recommendation Cache | Pre-computed user recommendations | Ranked title lists per user | High-QPS reads on homepage load | Redis Cluster | Memcached, DynamoDB DAX |
| Event Bus | User events, encoding events | Streaming events | High-throughput append, multiple consumers | Kafka | Redpanda, Kinesis |
| Search Index | Title search, genre browse | Catalog text + facets | Full-text + filtered queries | Elasticsearch | OpenSearch, Meilisearch |
| Session Store | Playback state, DRM tokens | Ephemeral session data | High-QPS read/write | Redis | Memcached |
Why a custom CDN (Open Connect), not CloudFront? At Netflixβs scale (15% of global internet traffic during peak), paying per-GB to a CDN provider is prohibitively expensive. Custom hardware appliances (Open Connect Appliances - OCAs) placed inside ISP networks cost $0.001/GB vs $0.02+/GB for commercial CDNs. Saves $1B+/year.
Why Temporal for encoding, not a simple queue? Video encoding is a multi-step workflow: split β encode each resolution β validate β package into HLS/DASH β DRM encrypt β publish. Steps have dependencies, retries, and can take hours. Temporal handles long-running workflows with checkpointing, retry policies, and visibility - far better than chaining SQS queues.
11. Data Modeling
Postgres / Vitess (Content Catalog β title metadata):
CREATE TABLE titles (
title_id UUID PRIMARY KEY,
name VARCHAR(500) NOT NULL,
type VARCHAR(10), -- movie, series
genres JSONB,
release_year INTEGER,
maturity_rating VARCHAR(10),
description TEXT,
poster_url VARCHAR(500),
avg_rating DECIMAL(2,1)
);
CREATE TABLE episodes (
episode_id UUID PRIMARY KEY,
title_id UUID REFERENCES titles(title_id),
season_number INTEGER,
episode_number INTEGER,
name VARCHAR(500),
duration_seconds INTEGER,
manifest_url VARCHAR(500) -- HLS/DASH manifest in S3
);
Cassandra (User Activity β watch history and progress):
Table: watch_history
PK: user_id
SK: watched_at (DESC)
Columns: title_id, episode_id, progress_seconds, duration_seconds, completed (boolean)
Table: watch_progress (for "continue watching")
PK: user_id
SK: title_id
Columns: episode_id, progress_seconds, last_watched_at, completed
S3 (Object Storage β encoded video segments):
Path: s3://video-assets/{title_id}/{episode_id}/{profile}/{segment_number}.m4s
Profiles: 240p_400kbps, 480p_1mbps, 720p_3mbps, 1080p_5mbps, 4k_15mbps
Manifest: s3://video-assets/{title_id}/{episode_id}/master.m3u8 (HLS adaptive)
Redis (Recommendation Cache β pre-computed per-user):
Key: "recs:{userId}:{row_type}" β List of title_ids (ordered by relevance)
Row types: "continue_watching", "trending", "because_you_watched_{titleId}", "top_picks"
TTL: 6 hours (refreshed by recommendation pipeline)
Access Patterns:
| Query | Data Source | How |
|---|---|---|
| Homepage rows (recommendations) | Redis | LRANGE recs:{userId}:top_picks 0 39 β 40 title_ids, hydrate from catalog cache |
| Play video | CDN (Open Connect) | Client fetches manifest β adaptive bitrate β CDN serves segments |
| Update watch progress | Cassandra | INSERT INTO watch_progress every 30 seconds during playback |
| βContinue Watchingβ row | Cassandra | Query watch_progress WHERE user_id = ? AND completed = false ORDER BY last_watched_at DESC |
| Search titles | Elasticsearch | Full-text on title name + genre facets + maturity filter |
How Adaptive Bitrate Streaming Works at the Data Level:
- Upload pipeline encodes each episode into 5-8 bitrate profiles (240p to 4K) using FFmpeg/custom encoders
- Each profile is split into 2-6 second segments (.m4s files), stored in S3
- A master manifest (
.m3u8) lists all profiles and their segment URLs - Client player downloads manifest β starts with mid-quality β measures throughput per segment
- If bandwidth drops: player switches to lower profile on the next segment (no rebuffer)
- CDN (Open Connect) caches popular segments at ISP-level edge boxes β 95%+ cache hit for popular content
12. Deep Dives
1) Why does one bitrate ladder waste bandwidth on a cartoon and still ruin an action scene?
Problem: FR1 expands one fixed ladder for every title in the catalog. A bitrate is a guess about how hard a frame is to compress, and applying the same guess to a cartoon and to a handheld chase scene is wrong in both directions.
Bad: the fixed ladder FR1 built β 1080p always at 5 Mbps, regardless of content. It fails in both directions at once, which is what makes it worth replacing rather than tuning.
A simple animated title has large flat colour areas and little motion, so it reaches transparent quality well under 5 Mbps. Every bit above that is paid for on egress and buys nothing a viewer can see. A high-motion, film-grain action sequence needs considerably more than 5 Mbps, so the same ladder produces visible blocking on exactly the scenes people care most about.
The reason this matters more here than it would elsewhere is the multiplier. We egress roughly 6 exabytes a month, so a 20% reduction in average bitrate is over an exabyte of traffic a month that simply does not have to happen. Encoding is paid once per title; delivery is paid every single time anyone watches. There is almost no compute budget it would be irrational to spend here.
Good: Per-title encoding β analyze each titleβs complexity once and assign it its own ladder. βMy Little Ponyβ gets 1080p at 2 Mbps, βMad Maxβ gets 8 Mbps. This captures most of the available saving for a modest analysis cost. What it misses is variation within a title: a two-hour film has quiet dialogue scenes and frantic action scenes, and one average ladder over-spends on the former and under-spends on the latter.
Great: Per-shot encoding (borrowing from Netflix Cosmos):
- Scene detection: Split video into shots (scene changes detected via frame difference). Each shot gets its own optimal encoding parameters.
- Convex hull optimization: For each shot, encode at multiple bitrate/quality points. Plot quality (VMAF score) vs bitrate. Find the convex hull - the set of points that gives maximum quality per bit.
- Bitrate allocation: Given a target average bitrate for the stream, allocate more bits to complex shots and fewer to simple shots. Result: consistent visual quality throughout.
- Codec selection: Encode each title in H.264 (compatibility), H.265 (50% more efficient, most devices), and AV1 (30% more efficient than H.265, newer devices). Client picks best codec their device supports.
flowchart LR
Master["Master File"]:::data
SD["Scene Detector"]:::service
CH["Convex Hull Analyzer"]:::service
ENC["Parallel Encoders"]:::service
PKG["Packager HLS and DASH"]:::service
S3["S3 Segments"]:::data
Master -->|"1. Detect scenes"| SD
SD -->|"2. Analyze complexity"| CH
CH -->|"3. Encode parallel"| ENC
ENC -->|"4. Package manifests"| PKG
PKG -->|"5. Store segments"| S3
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Cost consideration: Per-shot encoding costs roughly 3x the compute of a fixed ladder and yields files 20-30% smaller at the same measured quality. Against 6 exabytes/month of egress, a 25% reduction is ~1.5 exabytes/month of traffic avoided, so the extra encoding compute is recovered almost immediately on any title with real viewership. Encode once, stream millions of times β the asymmetry is the whole argument.
And the durability hole from FR1. FR1βs plain queue loses a four-hour encode if the worker dies, because a consumed queue entry is the only record that the work was needed. This pipeline has more steps than that one did β scene detection, then per-shot analysis, then parallel encodes, then packaging β so the exposure is worse, not better. Run it under a workflow engine (Temporal / Cadence / Step Functions).
π‘ A workflow engine persists the position of a long-running job, so a crash resumes from the last completed step instead of restarting. Learn more β
The property that matters: a worker lost during the AV1 encode retries that variant, not the scene detection and analysis that preceded it.
2) At 15% of global internet traffic, what does a rented CDN cost and when do you build one?
Problem: FR2 has every player fetching every segment from one S3 region. That is the single most expensive decision on the page, and the arithmetic is worth doing before reaching for a CDN reflexively.
Bad: serving straight from origin, as FR2 does. Take the egress number seriously: ~6 exabytes/month, which is 6 billion GB. At S3 list-price egress of $0.09/GB that is on the order of $540M/month. Nothing else in this design is within two orders of magnitude of that.
Latency is the other half. A viewer in Jakarta streaming from us-east-1 is ~250ms per segment request, and segments arrive just-in-time by design, so any sustained round-trip above the segment duration means the buffer drains faster than it fills. Playback does not degrade gracefully β it stalls.
Then the obvious fix, which is still bad at this scale: put a commercial CDN (CloudFront / Akamai / Fastly) in front. Latency is solved. Cost is not: even at a heavily committed ~$0.02/GB, 6 billion GB is ~$120M/month, around $1.4B/year, to move bytes we already own. At a few petabytes of catalog and 15% of global internet traffic, we are the largest single flow on many networks, and renting per-GB delivery at that share of traffic is renting at the worst possible volume.
This is the case where βjust use a CDNβ is the wrong answer, and it is wrong for a reason you can only see by computing it.
Good: Build our own CDN β our own PoPs in major metros. The per-GB rent disappears and is replaced by capital and operations we control, which at this volume is straightforwardly cheaper. What it does not remove is the last hop: traffic still crosses from our PoP into each ISPβs network as transit, which both costs money and adds the latency of whatever path the ISP takes to reach us.
Great: Deploy custom appliances (OCAs) directly inside ISP networks:
- Open Connect Appliances (OCAs): Custom servers with 100+ TB SSD storage + 100Gbps NIC. Deployed inside ISP data centers (Comcast, Vodafone, Jio, etc.).
- Pre-positioning (proactive caching): Every night, a popularity prediction model identifies which titles will be watched tomorrow in each region. Those titles are pushed to local OCAs overnight (off-peak bandwidth is nearly free).
- Fill strategy: Cache miss on an OCA β fetch from a parent OCA (regional hub) β only if parent misses β fetch from S3 origin. Three-tier hierarchy minimizes origin load.
- Steering: Control plane selects the best OCA for each client based on: network proximity, current load, content availability. Uses DNS-based and HTTP redirect steering.
Result: 95%+ of bytes served from within the userβs ISP network. Latency < 5ms for segment fetch. ISPs benefit too (no transit costs for Netflix traffic), so they deploy OCAs for free.
Cache eviction: LRU with popularity weighting. A title watched by 1000 users/day stays cached over one watched by 10 users/day, even if the latter was accessed more recently.
3) Bandwidth halves mid-scene. How do we drop quality instead of buffering for 30 seconds?
Problem: FR2βs player picks a quality level at startup and holds it for the rest of the film. Available bandwidth is not a constant β someone walks upstairs, a housemate starts a download, a phone hands off from WiFi to cellular β and the manifest already lists every alternative. We are just not using it.
Bad: the fixed quality level FR2 chose. Work through what happens when a 5 Mbps connection drops to 2 Mbps mid-scene while the player is still requesting 1080p at 5 Mbps.
Each 4-second segment now takes ~10 seconds to arrive, so the buffer drains 6 seconds for every 4 seconds of video delivered. A 30-second buffer is exhausted in about 20 seconds, and then playback stops. It does not stop briefly: the player is still asking for segments it cannot download in time, so it stalls, waits for a segment, plays 4 seconds, and stalls again. The failure is unbounded β it persists for as long as the network is degraded, which could be the rest of the film.
The infuriating part is that a perfectly watchable 1.5 Mbps version of that exact segment already exists in S3, encoded by FR1, listed in the manifest the player is holding. We built the alternatives and then wrote a player that ignores them.
Good: Throughput-based ABR β measure how fast the last segment arrived and pick the highest quality that fits. This solves the stall, and it introduces oscillation: throughput measured over a single 4-second segment is a noisy estimate, so the player chases it, flipping 480p β 1080p β 480p every few seconds. Constant visible quality changes are their own kind of unwatchable, and each switch risks a fresh mis-estimate.
Great: Buffer-based ABR with throughput smoothing (borrowing from Netflixβs practical ABR research):
- Buffer-Based (BBA): Decision is driven by buffer level, not just throughput. If buffer is full (30s ahead), be aggressive (higher quality). If buffer is low (< 5s), be conservative (lower quality). Smooth transitions.
- Throughput estimation: Use harmonic mean of last 5 segment downloads (not arithmetic mean - resists outliers from temporary spikes).
- Startup optimization: During initial buffering, start at lowest quality for the first 2 segments (fast start, reduces time-to-first-frame). Then ramp up aggressively once buffer grows.
- Quality lock: Once at a stable quality for 30+ seconds, donβt drop unless buffer is critically low. Prevents flicker.
Pseudocode for ABR decision:
buffer_level = current_buffer_seconds
throughput = harmonic_mean(last_5_segments_speed)
if buffer_level > 30s:
quality = highest where bitrate < throughput * 0.9
elif buffer_level > 10s:
quality = current_quality (hold steady)
elif buffer_level > 5s:
quality = max(current_quality - 1, lowest)
else:
quality = lowest (emergency drop)
Why not server-side ABR? The client knows its actual buffer state and network conditions in real-time. Server canβt know if the user is on WiFi or just entered a tunnel. Client-side ABR reacts in < 100ms; server-side would add round-trip latency to every quality decision.
4) How do we fill a home screen for each of 250M subscribers without showing everyone the same top ten?
Problem: FR3 builds every row on the request, from viewing history and catalog attributes. It is personalized, it is slow, and the signal it personalizes on is the weakest one available.
Bad: the per-request attribute matching FR3 built. Three failures, and they are independent.
Latency. Reading a full viewing history and scanning the catalog for attribute overlap, for each of ~10 rows, is comfortably into the hundreds of milliseconds against a 200ms budget. At ~25K home screen loads/sec at peak, we are also doing this work ~25K times a second while recomputing a nearly identical answer for the same user several times a day. The work is both too slow and almost entirely redundant.
Signal quality. βYou watched a thriller, here are more thrillersβ can only ever recommend things that look like what you have already seen. It cannot learn that viewers who liked one specific thriller reliably love a nature documentary that shares no genre, cast or director with it β and those non-obvious matches are the ones that make a recommender feel good rather than merely correct.
Cold start. A new profile has no history, so FR3βs step 2 returns nothing and the home screen degrades to raw popularity β which is the undifferentiated Top 10 that personalization was supposed to replace, shown to precisely the users whose first impression matters most.
Good: Collaborative filtering β learn from co-viewing patterns rather than attributes, so βpeople like you also watchedβ can surface titles no attribute match would find. This fixes signal quality and directly addresses the thing attribute matching cannot do. It does not fix latency, since matrix factorization over 250M profiles is not a request-time operation, and it makes cold start worse: a new profile has no interactions, so it has no position in the learned space at all.
Great: Two-stage pipeline with hybrid signals:
Stage 1 - Candidate Generation (offline, runs every 4 hours on Spark):
- Collaborative filtering: Matrix factorization on user-item interaction matrix. Each user and title gets a 128-dimension embedding. Similarity in embedding space = likely interest.
- Content-based: Encode title features (genre, director, cast, keywords, avg shot length) into embeddings. Match against user preference vector.
- Output: 1000 candidate titles per user (from 15K catalog).
Stage 2 - Ranking (near-real-time, per request):
- Features: Candidate title embedding + user context (time of day, device, recent watches, account age).
- Model: Lightweight neural net (2-layer MLP) predicting P(watch > 70% of title).
- Inference: < 10ms for 1000 candidates on CPU. Results sorted by score.
- Row assembly: Group ranked titles into themed rows (βBecause you watched Stranger Things,β βTrending,β βNew Releasesβ). Each row is a retrieval source.
flowchart LR
Events["User Events Kafka"]:::async
Spark["Spark Pipeline"]:::service
CF["Collaborative Filter"]:::service
CB["Content-Based"]:::service
CAND["Candidates 1K per user"]:::data
Ranker["Neural Ranker"]:::service
Cache["Redis Rec Cache"]:::data
Events -->|"1. Stream events"| Spark
Spark -->|"2. Collaborative filter"| CF
Spark -->|"3. Content-based filter"| CB
CF -->|"4. Generate candidates"| CAND
CB -->|"5. Generate candidates"| CAND
CAND -->|"6. Rank"| Ranker
Ranker -->|"7. Store results"| Cache
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Cold start (new user): First session uses onboarding quiz (pick 3 genres you like) + popularity-based defaults. After 5+ interactions, collaborative filtering kicks in.
Freshness vs compute tradeoff: Full pipeline rerun every 4 hours. But: real-time signals (user just finished a horror movie β boost horror in next homepage load) are injected via a lightweight βre-rankerβ that adjusts cached scores using the last 30 minutes of activity.
5) A blockbuster drops at midnight and millions press play. How is it already at the edge?
Problem: Deep Dive 2 put appliances inside ISP networks, which fixes delivery for content the appliance already holds. FR1 never decided how content gets onto them. A cache that fills on demand is fine for a steady catalog and useless for a simultaneous global premiere.
Bad: reactive caching β an ordinary LRU cache, filled by the first viewer in each region. This is the default behaviour of every cache and it is exactly wrong for the traffic shape a premiere produces.
At midnight the episode is on no appliance anywhere. A 2-hour title at 4K across the bitrates people actually use is roughly a 20GB pull per appliance, and with on the order of 10,000 appliances all missing within the same few minutes, that is about 200TB of origin egress inside the first hour β a coordinated stampede on the one component Deep Dive 2 was built to protect.
Worse, it lands on the viewers who care most. The people pressing play at midnight on release night are the most engaged audience the title will ever have, and they are precisely the ones served by a cold cache. Cache hit rate is worst exactly when demand is highest, which is the opposite of what a cache is supposed to do.
Good: Push the new title to every appliance 24 hours before release. The stampede disappears, since day-one viewers hit a warm cache. The problem is that appliance storage is finite β on the order of 100TB against a catalog of several petabytes β so βpush everything everywhereβ does not fit, and much of what you pushed is wrong anyway: a title that dominates in India may get almost no plays in Brazil, and that shelf space had alternatives.
Great: Predictive pre-positioning with regional popularity models:
- Popularity prediction: ML model trained on: historical premiere viewership, pre-release engagement (trailers watched, βremind meβ clicks), genre popularity per region, time of year, marketing spend.
- Regional allocation: Model predicts views-per-region for the first 48 hours. Content pushed to OCAs proportionally - more copies in high-demand regions, fewer in low-demand.
- Fill during off-peak: Transfers happen 2AM-6AM local time when ISP bandwidth is idle. OCAs pull from regional hubs (not origin directly).
- Dynamic rebalancing: After release, actual viewership data feeds back. If Brazil has unexpected demand, nearby OCAs serve while additional copies propagate.
- Tiered encoding push: Push only the most popular bitrates first (1080p H.265, 720p H.264). Niche formats (4K AV1) stay at regional hubs until demanded.
Result: For a major premiere, 99%+ of first-play requests hit warm OCA cache. Origin egress stays flat regardless of demand spike.
6) The CDN node serving an episode dies mid-stream. What does the player do?
Problem: FR2βs player fetches every segment from one place and has no defined behaviour when a fetch fails. Deep Dive 2 changed where that one place is without changing the fact that there is only one, and appliances sitting in thousands of ISP racks are hardware that fails.
Bad: the single-source fetch FR2 built, with no failure path. If the appliance serving a session dies, or the path to it partitions, the playerβs next segment request never completes and playback stops permanently β the player has no second address to try and no rule telling it to look for one.
Note the blast radius that Deep Dive 2 quietly created. An appliance holds tens of thousands of concurrent sessions, so one failure does not inconvenience one viewer, it ends the session for everyone streaming from that appliance at that moment. And it is the most likely component to fail in the whole design, being commodity hardware in a rack we do not operate.
Against a requirement where more than about 2 seconds of interruption loses the viewer, βplayback stopsβ is the worst available outcome, and it is the one the design as written produces.
Good: Retry the failed segment against the same source with exponential backoff. This absorbs transient packet loss and brief congestion, which is a real fraction of failures. It does nothing for the case that matters most: if the source itself is gone, every retry fails, and backoff means the player spends progressively longer not fetching video while it waits to be disappointed again.
Great: Multi-level resilience with fast failover:
- Segment-level retry: If a segment fetch fails or times out (3s), immediately retry from a different CDN PoP (player knows 3+ candidate OCAs from steering).
- Quality downshift on repeated failure: If 2 consecutive segments fail at current quality, drop to lowest bitrate (which has smaller segments, more likely to succeed on degraded network).
- Buffer-based tolerance: ABR algorithm maintains a 30-second buffer when healthy. This βrunwayβ absorbs network glitches - user doesnβt notice a 5-second outage if buffer covers it.
- Session resume: If the player crashes or the device sleeps, resume from the exact position β FR2βs heartbeat already writes it every 10 seconds, so this costs nothing extra. Re-opening the app continues from where the viewer left off rather than restarting the title.
- CDN steering updates: Every 30 seconds, player pings the steering service for updated CDN PoP rankings. If their current OCA is degraded, next segment fetches from a better one.
Backstop: If all CDN PoPs in a region fail (rare, ISP-level outage), players fall back to fetching from a regional hub OCA in a neighboring ISP. Higher latency (50ms vs 5ms) but playback continues.
13. Design Self-Audit
| Question | Answer |
|---|---|
| Dedicated search index? | β Elasticsearch for title/cast/genre search. Covered in FR3 components. |
| Stale reads after writes? | Encoding: title marked PLAYABLE only after all variants confirmed. Recs: 4-hour staleness acceptable for homepage; real-time re-ranker handles short-term signals. |
| Single points of failure? | CDN is massively distributed (10K+ OCAs). Catalog DB is sharded (Vitess). Temporal orchestrator has its own HA (multiple workers, durable state). Playback Service is stateless behind LB. |
| Dead-letter / reconciliation? | Encoding: failed tasks retry 5x in Temporal, then flag for review. Playback: heartbeat timeout after 5 minutes marks session as abandoned (for analytics). |
| Data freshness across caches? | Rec cache refreshed every 4 hours + real-time re-ranker. CDN content is immutable (new encoding = new URL). Session store is real-time (Redis). |
| Cost at scale? | CDN (Open Connect): dominant cost, mitigated by ISP partnerships. Encoding: GPU compute is expensive but one-time per title. Storage: tiered S3 (originals archived to Glacier after encoding). Rec pipeline: Spark cluster - scales with user base. |
14. Core Flows
Flow 1: Video Playback Start-to-Stream
sequenceDiagram
participant User
participant App as Player App
participant GW as API Gateway
participant PS as Playback Service
participant DRM as DRM Server
participant CDN as CDN Edge
participant S3
User->>App: Tap "Play"
App->>GW: POST /playback/start
GW->>PS: Route request
PS->>PS: Check subscription + region + resume position
PS->>DRM: Request license (Widevine or FairPlay)
DRM-->>PS: License token
PS-->>App: manifestUrl + license + resumeAt
App->>CDN: GET manifest.m3u8
CDN-->>App: Manifest (all quality levels)
App->>App: ABR selects initial quality (480p)
App->>CDN: GET segment_001_480p.ts
alt CDN cache hit
CDN-->>App: Segment (from edge cache)
else CDN cache miss
CDN->>S3: Fetch segment
S3-->>CDN: Segment
CDN-->>App: Segment (cached for next user)
end
App->>App: Decode + render first frame
Note over User,App: Time to first frame target < 2s
loop Every segment
App->>App: Measure throughput
App->>CDN: GET next segment at adapted quality
end
Non-obvious failure path: If CDN edge is overloaded (flash crowd for a new release), the playerβs ABR algorithm detects slow segment downloads and drops quality aggressively. If even the lowest quality buffers, the player switches to a different CDN PoP (DNS-based failover). Playback Service also maintains a βsteeringβ endpoint that can redirect players away from saturated edges.
Flow 2: Encoding Pipeline
sequenceDiagram
participant Studio
participant IS as Ingest Service
participant S3
participant Temporal
participant Worker as Encoding Worker
participant CAT as Catalog DB
participant CDN as CDN Pre-position
Studio->>S3: Upload master file
S3->>IS: S3 event notification
IS->>IS: Validate format + metadata
IS->>Temporal: Create encoding workflow
Temporal->>Temporal: Split into parallel tasks
loop For each resolution x codec
Temporal->>Worker: Assign encoding task
Worker->>S3: Download master chunk
Worker->>Worker: Transcode (FFmpeg or SVT-AV1)
Worker->>S3: Upload encoded segments
Worker->>Temporal: Report completion
end
Temporal->>Temporal: All variants complete
Temporal->>Temporal: Generate HLS and DASH manifests
Temporal->>S3: Upload manifests
Temporal->>CAT: Mark title as PLAYABLE
Temporal->>CDN: Trigger pre-positioning for predicted-popular regions
Non-obvious failure path: If an encoding worker crashes mid-transcode (GPU OOM, spot instance reclaimed), Temporal automatically retries the failed task on a different worker. The worker resumes from the last completed segment (checkpointed in Temporal state). Partial segments in S3 are overwritten. The overall workflow only fails if a single task exceeds 5 retries - then itβs flagged for manual review.
Playback Session State Machine
stateDiagram-v2
[*] --> INITIALIZING : Play pressed
INITIALIZING --> BUFFERING : Manifest received
BUFFERING --> PLAYING : Buffer full enough
PLAYING --> BUFFERING : Buffer depleted
PLAYING --> PAUSED : User pauses
PAUSED --> PLAYING : User resumes
PLAYING --> SEEKING : User scrubs timeline
SEEKING --> BUFFERING : Seek position set
PLAYING --> ENDED : Reached end
PLAYING --> ERROR : Network failure
ERROR --> BUFFERING : Retry succeeds
ERROR --> ENDED : Max retries exceeded
15. Final Architecture
flowchart TB
APPS(["TV Mobile and Web"]):::client
DNS["DNS Steering<br>nearest edge"]:::edge
GW["API Gateway"]:::edge
CDN["Open Connect CDN<br>serves the bytes"]:::edge
PS["Playback Service<br>manifest and licence"]:::service
RS["Recommendation Service"]:::service
SS["Search Service"]:::service
ENC["Ingest and Encoding<br>ladder of bitrates"]:::service
KF[["Kafka<br>view events"]]:::async
S3[("Object Store<br>video segments")]:::data
CAT[("Catalog DB<br>titles and metadata")]:::data
CASS[("Cassandra<br>watch history")]:::data
ES[("Elasticsearch<br>search index")]:::data
DRM[/"DRM Widevine and FairPlay"/]:::external
APPS --> DNS
DNS --> CDN
APPS --> GW
GW --> PS
GW --> RS
GW --> SS
SS --> ES
RS --> CASS
PS --> DRM
PS -->|"where are the segments"| CDN
CDN --> S3
ENC -->|"write renditions"| S3
ENC --> CAT
APPS -->|"what was watched"| KF
KF --> CASS
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef async fill:#3b1f5e,stroke:#c084fc,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
classDef external fill:#4a1942,stroke:#f472b6,color:#e2e8f0
How it works end-to-end (playback path):
- Client requests playback β Smart TV/Mobile/Web resolves via DNS Steering to nearest API Gateway (Zuul)
- Playback Service handles request β fetches DRM license from Widevine/FairPlay, checks Redis for session state
- Manifest returned β client receives HLS/DASH manifest with segment URLs pointing to CDN
- Video segments streamed from CDN β Open Connect appliances serve 95%+ of traffic from inside ISPs
- ABR adapts quality β client-side algorithm switches bitrate per segment based on buffer level
How it works end-to-end (content ingestion path):
- New title ingested β Ingest Service kicks off a Temporal workflow
- Encoding Workers transcode β GPU workers produce per-title optimized bitrate ladders, store segments in S3
- Catalog updated β Temporal marks the workflow complete, updates Catalog DB
- Recommendation pipeline refreshes β Kafka streams user activity to ML Pipeline (Spark), updated models cached in Redis
Want a deep dive on DRM (content protection), offline downloads, multi-language audio switching, or live streaming (sports events)? Drop a comment below π
Key Technologies
| Term | What it is |
|---|---|
| HLS / DASH | HTTP-based streaming protocols that split video into small segments and let the player pick quality per segment. |
| ABR (Adaptive Bitrate) | Client-side algorithm that switches video quality in real-time based on available bandwidth and buffer level. |
| CDN / Open Connect | Content Delivery Network; Netflixβs custom CDN places appliances inside ISPs to serve 95%+ of traffic locally. |
| Per-Title Encoding | Analyzing each titleβs visual complexity to assign optimal bitrate per resolution instead of a fixed ladder. |
| Temporal / Cadence | Durable workflow engines that checkpoint multi-step encoding jobs so crashed steps retry without restarting the whole pipeline. |
| Kafka | Distributed event log used here for user activity streaming, encoding events, and decoupling services. |
| Cassandra | Wide-column NoSQL store used for high-write user activity data (watch history, play events). |
| Redis | In-memory cache used for recommendation results, session state, and playback position tracking. |
| Elasticsearch | Search engine powering full-text title search and genre browsing with faceted filtering. |
| DRM (Widevine / FairPlay) | Digital Rights Management systems that encrypt video content and issue device-specific playback licenses. |
Whatβs Expected at Each Level
This section helps you calibrate your depth. You donβt need to cover everything - just know whatβs expected for your level.
Mid-level
Outline the upload β encode β store β stream flow. Understand why a CDN is needed for latency and bandwidth at scale. Describe HLS/DASH at a high level - the idea of splitting video into segments and serving a manifest. With prompting, discuss adaptive bitrate and why users canβt just download a single fixed-quality file.
Senior
Propose per-title encoding optimization (different bitrate ladders for animated vs action content). Explain the CDN cache hierarchy (edge β regional β origin) and why three tiers matter. Discuss the recommendation pipeline (offline batch candidate generation + online re-ranking). Articulate the ABR algorithm trade-offs - buffer-based vs throughput-based - and why client-side ABR beats server-side.
Staff+
Address custom CDN economics (Open Connect vs commercial CDN at Netflix scale - $0.001/GB vs $0.02/GB). Discuss predictive pre-positioning of content based on popularity models and how overnight off-peak bandwidth is leveraged. Explain DRM license serving architecture and the cold-start recommendation problem. Articulate what happens when a new blockbuster drops and millions hit play simultaneously - cache warming, origin shielding, and steering.
π― Key Takeaways
- Per-shot encoding optimizes bitrate per scene complexity - saves 20-30% bandwidth
- Open Connect CDN inside ISPs serves 95% of traffic locally
- Adaptive Bitrate (ABR) driven by buffer level, not just throughput
- Two-stage recommendation: collaborative filtering (offline) β neural ranker (online)
Related Designs
- URL Shortener - CDN caching patterns and redirect optimization
- Chat System - real-time delivery and session management
- Notification System - push notification infrastructure
Related Concepts
Understand the building blocks used in this design:
- CDN β β streams video segments from edge caches (Open Connect-style) for low startup latency
- Object Storage β β holds encoded video assets across every bitrate and resolution
- Caching β β caches manifests and metadata so playback starts fast
- Load Balancing β β spreads playback and API traffic across the fleet
Discussion
Newest first