Designing a Social Media Feed (Twitter / X / Threads)
Difficulty: Intermediate Prerequisites:Fan-Out, Caching, and Message Queues
TL;DR
A social feed shows each user a personalized timeline of posts from people they follow. The core challenge is fan-out - when a user with 10M followers tweets, how do you update 10M timelines quickly?
flowchart LR
POSTER["User posts tweet"]:::client
API["Tweet Service"]:::service
K["Fan-out Service<br/>Kafka"]:::async
CACHE[("Per-user timeline<br/>Redis")]:::data
READER["User opens feed"]:::client
FEED["Feed Service"]:::service
POSTER -->|"1. POST new tweet"| API
API -->|"2. Publish tweet event"| K
K -->|"3. Write to follower feeds"| CACHE
READER -->|"4. GET home timeline"| FEED
FEED -->|"5. Read pre-built feed"| CACHE
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef async fill:#AB47BC,stroke:#4A148C,color:#fff
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
In 3 sentences: When someone tweets, the system either pushes that tweet into every followerβs pre-built timeline cache (fan-out on write) or waits until each follower opens their feed and assembles it on-the-fly (fan-out on read). Most systems use a hybrid: push for regular users, pull for celebrities. The timeline is cached in Redis as a sorted list of tweet IDs per user.
Understanding the Problem
What is a social feed? When you open Twitter/X, Instagram, or LinkedIn, you see a stream of posts from accounts you follow (and maybe recommended content). That stream is your timeline - a personalized, ordered list assembled from thousands of content sources.
Why is it hard?
- User A follows 500 people. Each tweets 5 times/day. Feed must merge 2500 posts/day into a ranked timeline.
- Celebrity with 50M followers tweets once β 50M timelines need updating.
- The feed must load in under 200ms on a slow phone connection.
- βOut of orderβ tweets feel broken - chronological or ranked, but never randomly jumbled.
Real numbers (Twitter/X scale):
- 500M+ DAU
- 500M tweets/day
- Average user follows ~400 accounts
- Median follower count: ~200. Top accounts: 50M+
Naive First Cut
flowchart LR
USER["User opens feed"]:::client
API["Feed API"]:::service
DB[("Tweets table<br/>SELECT WHERE author IN followees ORDER BY time")]:::data
USER --> API
API --> DB
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
On each feed request: SELECT * FROM tweets WHERE author_id IN (SELECT followee_id FROM follows WHERE follower_id = ?) ORDER BY created_at DESC LIMIT 50.
Why this breaks:
- β
IN (500 followee IDs)β massive query scanning millions of rows - β Every feed open = expensive DB query. 500M DAU Γ 10 opens/day = 5B queries/day
- β No caching - same expensive query repeated every few seconds
- β No ranking - just chronological, no relevance
- β Celebrity tweet β 50M users all running this query simultaneously
Prior Art Weβre Drawing From
- Twitter Fan-out Service - The original implementation that coined βfan-out on write.β Pre-computes timelines for users with < 500K followers; assembles on-read for celebrities. Processes 500M tweets/day into 200B+ timeline writes. (Twitter Engineering blog)
- Facebook TAO - Graph-aware caching layer serving the social graph at billions of QPS. Demonstrates that follow relationships must be cached separately from content for performance. (Facebook TAO paper)
- Instagram Feed Ranking - Moved from chronological to ML-ranked feed. Two-stage pipeline: candidate generation (pull from timeline) β ranking model predicting engagement probability. (Instagram Engineering)
- LinkedIn Feed Architecture - Uses a βfeed mixerβ pattern that merges multiple content sources (network updates, sponsored content, recommendations) into a single ranked stream. (LinkedIn Engineering blog)
Functional Requirements
Core (Top 3)
- Users should be able to post a tweet - text plus optional media, visible to their followers
- Users should be able to follow and unfollow other users - the follow graph is what decides whose tweets appear
- Users should be able to open a home timeline - recent tweets from the accounts they follow, newest first
Below the Line
- Replies, quote tweets and threads
- Likes, retweets and bookmarks
- Search and trending topics
- Direct messages
- Notifications
Non-Functional Requirements
Core
| NFR | Target |
|---|---|
| Timeline read latency | Home timeline in under 200ms at p99 - this is the most frequent operation in the product |
| Read-heavy skew | Roughly 100 reads per write, so the design optimises reads even at the cost of write work |
| Availability over consistency | A timeline that is a few seconds stale is fine; a timeline that fails to load is not |
| Fan-out skew | Follower counts span 0 to 100M+, so per-tweet write cost cannot be proportional to follower count |
Below the Line
- Ordering guarantees stronger than newest-first by timestamp
- Deletion propagating instantly to already-built timelines
- Multi-region active-active writes
Scale Estimation (Back-of-Envelope)
- Users: 500M DAU, ~400 average follows per user
- Write QPS: ~6K tweets/sec (500M tweets/day)
- Read QPS: ~100K feed loads/sec at peak (each user opens their feed 10+ times a day)
- Storage: 500M tweets/day Γ ~300 bytes = ~150GB/day, so ~55TB/year of tweet text and metadata. Media lives in S3 and is counted separately
- Bandwidth: ~50 Gbps at peak for feed API responses, plus media served from CDN
The ratio that decides the architecture: reads outnumber writes by about 17 to 1 (100K feed loads/sec against 6K tweets/sec), and each read currently has to touch ~400 authors. Hold onto one more number for the deep dives β if we ever pre-compute timelines, 500M tweets Γ 400 followers is 200B timeline writes a day.
Core Entities
- User - account with profile, follower/following lists, and preferences
- Tweet - text content (280 chars), media attachments, author, timestamp, engagement counts
- Timeline - the ordered set of tweets a given user should see, newest first
- Follow Relationship - directional edge in the social graph (A follows B)
High-Level Design
Letβs build this incrementally, adding components as each requirement demands them.
FR1: User Posts a Tweet
The first interaction: a user types a tweet and hits Post. Read that requirement literally and it asks for one thing β the tweet exists afterwards and their followers can find it. So write it down and return.
New components:
- API Gateway - authenticates, rate-limits, routes. Entry point for all client requests.
- Tweet Service - handles tweet creation: validates content, attaches media references, writes the tweet.
- Tweet Store (Cassandra) - permanent storage for all tweets, partitioned by author so βrecent tweets by this accountβ is a single-partition read.
flowchart LR
POSTER["User"]:::client
GW["API Gateway"]:::edge
TS["Tweet Service"]:::service
TDB[("Tweet Store<br/>tweets by author")]:::data
POSTER -->|"1. POST new tweet"| GW
GW -->|"2. Auth and forward"| TS
TS -->|"3. Persist tweet"| TDB
TS -->|"4. Return tweetId"| POSTER
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Step-by-step:
- User taps Post β request hits API Gateway
- Gateway authenticates the JWT and forwards to Tweet Service.
authorIdcomes from the token, never the request body - Tweet Service writes one row to the Tweet Store, partitioned by
authorIdand clustered by time - It returns the tweetId. The tweet is now live: it is in the store, and FR2βs feed query will find it
Why is there no delivery step? Because nothing needs delivering yet. The tweet is discoverable the moment it is written, since FR2 finds tweets by asking the store for recent rows by author. Posting costs one write regardless of whether the author has 12 followers or 100 million, which is a property worth noticing before we give it up.
What we have deliberately left broken. Very little, and that is the point: at ~6K tweets/sec this is one small write per tweet and it will run on modest hardware for a long time. The cost is paid entirely on the read side, in FR2, and the deep dives are going to move that cost back here β at which point posting stops being cheap and constant, and the 100M-follower case becomes the hardest problem on the page.
FR2: User Opens Their Feed
Now users want to see their feed. Build the obvious thing: when someone asks for their timeline, look up who they follow and fetch those peopleβs recent tweets. Two components, no new infrastructure.
New components:
- Social Graph - stores who-follows-whom. Answers βwho does this user follow?β
- Feed Service - handles βshow me my feedβ. Asks the Social Graph for the followee list, fetches recent tweets for those authors from the Tweet Store, merges them by timestamp and returns the top 50.
flowchart LR
READER(["User"]):::client
GW["API Gateway"]:::edge
FEED["Feed Service<br>merges followee tweets"]:::service
GRAPH[("Social Graph<br>who follows whom")]:::data
TDB[("Tweet Store<br>tweets by author")]:::data
READER -->|"1. GET home timeline"| GW
GW -->|"2. Forward to feed svc"| FEED
FEED -->|"3. Who do they follow"| GRAPH
FEED -->|"4. Recent tweets per author"| TDB
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
classDef async fill:#AB47BC,stroke:#4A148C,color:#fff
Step-by-step for feed read:
- User opens app β
GET /feed - Feed Service asks the Social Graph for the userβs followee list
- For each followee, fetch their recent tweets from the Tweet Store
- Merge by timestamp, take the top 50, return them
What we have deliberately left broken. This is correct and it is genuinely fine for an account following a few dozen people. It stops being fine at the numbers in the requirements. A user following ~400 accounts means ~400 partition reads on every single feed open, and feed opens are the most frequent operation in the product β so the work per request scales with how social the user is, which is exactly backwards. Nothing here is a missing feature; it is a latency failure, so it is earned back in the deep dives: the pre-built timeline in Deep Dive 1, what that pre-building does to accounts with 100M followers in Deep Dive 2, and the fact that a timestamp is a poor way to choose 50 tweets out of a thousand in Deep Dive 3.
FR3: User Follows and Unfollows
The follow graph is the input to FR2βs feed query, so this is the requirement that makes the other two mean anything. It is also the smallest: a follow is a directed edge.
New components: none. FR2 already introduced the Social Graph; this is the write side of it.
The edge is a row in the Social Graph indexed both ways β (follower_id, followee_id) to
answer βwho do I followβ for the feed query, and (followee_id, follower_id) to answer βwho
follows meβ for the profile screen.
flowchart LR
USER["User"]:::client
GW["API Gateway"]:::edge
SG["Social Graph Service"]:::service
GRAPH[("Social Graph<br/>follow edges")]:::data
FEED["Feed Service"]:::service
USER -->|"1. POST follow account B"| GW
GW -->|"2. Auth and forward"| SG
SG -->|"3. Insert follow edge"| GRAPH
SG -->|"4. Return 200"| USER
FEED -->|"5. Later reads followees here"| GRAPH
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Step-by-step:
- User taps Follow on account B β
POST /users/{B}/follow - Gateway authenticates. The follower is taken from the token, so nobody can make someone else follow an account
- Social Graph Service inserts the edge. A unique constraint on the pair makes a double-tap idempotent instead of creating two edges
- It returns
200. Unfollow is the same operation in reverse: delete the row - Nothing else happens. The next feed request reads the new edge and Bβs tweets are in the results
What we have deliberately left broken. Almost nothing, and this is the one place the naive design is genuinely better than the production one. Because FR2 computes the timeline at read time, a follow takes effect on the very next refresh, and an unfollow stops showing that account immediately β no backfill, no cleanup, no window where the feed disagrees with the graph. Deep Dive 1 takes that away: once timelines are pre-built, a new follow changes nothing until someone backfills Bβs recent tweets, and an unfollow leaves stale tweets in the cached timeline until something removes them. That is a real regression and Deep Dive 1 has to pay for it.
The one thing that does break at scale is the follower-side read. followee_id = B for a
100M-follower account is a 100M-row scan, which nothing on this page needs yet β and which
Deep Dive 2 will need on every single tweet that account posts.
Technology Choices
| Tier | Purpose | Primary Pick | Alternatives | Why Primary Wins Here |
|---|---|---|---|---|
| Timeline Cache | Pre-built per-user feed (sorted tweet IDs) | Redis Sorted Sets | Memcached, DynamoDB, Cassandra | O(log N) insert for fan-out writes; O(1) range read for feed loads |
| Tweet Store | Durable tweet content + metadata | Cassandra | DynamoDB, ScyllaDB, Postgres | Append-heavy at 6K tweets/sec; partition by user_id for efficient fan-out-on-read |
| Event Bus | Fan-out events from tweet publish | Kafka | Kinesis, Pulsar, RabbitMQ | Ordered per-partition; replayable for backfill; handles 200B+ events/day |
| Social Graph | Follower/following edges | Redis (adjacency) + Postgres (durable) | Neo4j, DynamoDB, Cassandra | Redis for hot reads (get followers for fan-out); Postgres for graph mutations |
| CDN | Media attachments (images/video) | CloudFront | Cloudflare, Fastly, Akamai | Offloads media from API servers; immutable media = high cache hit ratio |
| Search | Tweet text and hashtag search | Elasticsearch | Algolia, Meilisearch, Typesense | Near-real-time indexing of 500M tweets/day with relevance ranking |
Why Cassandra over Postgres for tweets? At 500M tweets/day, append-only writes and partition-by-user access patterns map perfectly to Cassandraβs LSM storage. Postgres would struggle with write amplification and vacuuming at this scale.
Data Modeling
Redis (Timeline Cache β per-user pre-built feed):
Key: "timeline:{userId}" β Sorted Set (score = timestamp, member = tweetId)
Max entries: 800 per user
TTL: none (evicted by ZREMRANGEBYRANK when exceeding 800)
Cassandra (Tweet Store β permanent tweet content):
Table: tweets
PK: tweet_id (UUID, Snowflake-generated)
Columns: author_id, content, media_urls (list), created_at, like_count, retweet_count, reply_count
Table: user_tweets (for fan-out-on-read path)
PK: author_id
SK: created_at (DESC)
Columns: tweet_id, content_preview
Redis (Social Graph β hot follower lookups):
Key: "followers:{userId}" β Set of follower userIds
Key: "following:{userId}" β Set of followee userIds
Key: "celebrity_followees:{userId}" β Set of celebrity userIds this user follows
Access Patterns:
| Query | Data Source | How |
|---|---|---|
| Read home timeline | Redis | ZREVRANGE timeline:{userId} 0 49 β 50 tweet IDs |
| Hydrate tweet details | Cassandra | Multi-get by tweet_id batch |
| Fan-out on write | Redis | ZADD timeline:{followerId} <ts> <tweetId> per follower |
| Celebrity tweets at read time | Cassandra | Query user_tweets for each celebrity followee |
| Check if A follows B | Redis | SISMEMBER followers:{B} A |
How Fan-Out Works for a User with 10M Followers (Hybrid):
- Tweet arrives β Fan-out Service checks follower count
- If β€ 10K followers: push path. Workers read follower list from Redis
followers:{authorId}and ZADD tweet to each followerβs timeline sorted set - If > 10K followers (celebrity): skip push entirely. Tweet stays only in
user_tweetstable - At read time: Feed Service reads pre-built timeline from Redis, THEN fetches latest 5 tweets from each celebrity the user follows (via
celebrity_followees:{userId}set), merges by timestamp, ranks, returns top 50 - Trim: after each ZADD,
ZREMRANGEBYRANK timeline:{userId} 0 -801keeps the set bounded at 800 entries
Deep Dives
1) How do we return fifty tweets in about a millisecond?
Problem: FR2 builds the timeline at read time by asking the Social Graph who you follow and then reading each of those authorsβ recent tweets. Correct, and the work it does per request scales with how many accounts you follow, on the most frequent operation in the product.
Bad: the fan-out-on-read from FR2. The arithmetic is unforgiving: 100K feed loads/sec at peak Γ ~400 followees each is roughly 40M partition reads/sec, and our budget for the whole response is 200ms at p99.
Batching helps the count and not the latency. Even issued as one multi-partition query, the response cannot return until the slowest of ~400 partition reads comes back, so p99 feed latency tracks roughly p99.99 partition latency. Tail amplification is the real killer here: any storage layer with a 1-in-10,000 slow read will produce a slow feed for almost every user, because every user is rolling that dice 400 times per open.
It is also wasted work. The same 400 partitions get re-read on every open by every follower, and the answer barely changes between opens β we recompute a nearly identical timeline tens of times a day per user.
| Β | Fan-out on read (what FR2 built) | Fan-out on write |
|---|---|---|
| Feed read | Merge ~400 sources per request | One range read of one pre-built list |
| Tweet post | One write, regardless of followers | One write per follower |
| Freshness | Always current | Lags by fan-out delay |
| Breaks when | Users follow many accounts | Accounts have many followers |
Good: Cache the assembled timeline per user after the first read. The second open within the TTL is cheap. But the first open after any followee tweets is still a 400-way merge, and with 500M DAU opening the app throughout the day, cache misses dominate β you have added a layer without removing the expensive path.
Great: Invert it. Pre-build every userβs timeline at write time and keep it in a Redis sorted set, so a feed read is one range query.
π‘ Fan-out on write = doing the merge work once when a tweet is posted, instead of repeatedly every time a follower opens the app. Learn more β
Why a Redis sorted set? The access pattern is βordered list with cheap insert and cheap range read,β which is exactly what a sorted set is:
ZADD timeline:{userId} <timestamp> <tweetId>β O(log N) insert on fan-outZREVRANGE timeline:{userId} 0 50β newest 50 in O(log N + 50)ZREMRANGEBYRANK timeline:{userId} 0 -201β trim so timelines do not grow without bound
Mechanism:
- Tweet Service writes the tweet, then publishes a
TweetCreatedevent to an event bus (Kafka / Redpanda / Kinesis). The bus is what makes fan-out survive a worker crash, and it is why the poster does not wait for it - Fan-out Workers consume the event, read the authorβs follower list, and
ZADDthe tweetId into each followerβs timeline - Feed Service now answers a read with one
ZREVRANGEagainst one key, then hydrates the tweetIds with content from the Tweet Store - Timelines are trimmed to ~200 entries. Beyond that the client is deep-scrolling, which is rare enough to serve from FR2βs original read path
Memory math: a sorted-set entry holding an 8-byte tweetId costs roughly 64 bytes once skiplist and dictionary overhead are counted β not 8. At 800 entries per user that is 500M Γ 800 Γ 64 β 25TB, which is why the trim matters: at 200 entries it is 500M Γ 200 Γ 64 β 6.4TB, which fits on ~100 nodes with 64GB usable each.
What this costs. Two things got worse, and both were flagged earlier on this page:
- Posting is no longer cheap or constant. 500M tweets/day Γ ~400 followers is 200B timeline writes/day, about 2.3M writes/sec sustained. We moved the cost from the read path to the write path, which is the right trade only because reads outnumber posts by orders of magnitude here.
- Follows are no longer instant. FR3 got that for free. Now a new follow changes nothing until a backfill copies Bβs recent tweets into the followerβs timeline, and an unfollow leaves stale tweets until a cleanup job removes them. Seconds, usually, and a genuine regression we accepted knowingly.
Cold cache: a user who has not opened the app in weeks may have been trimmed or evicted. Fall back to FR2βs read path, build the timeline, cache it, and serve. Lazy population means we only hold timelines for users who actually show up.
2) What happens when someone with 100M followers tweets?
Problem: Deep Dive 1 bought a fast read path by paying per follower at write time. That trade was priced for the average account, which has 400 followers. It was not priced for 100 million.
In simple terms: A celebrity with 100M followers posts. If we push that tweet into 100M timelines one at a time, it takes minutes β and for most of those minutes, most followers cannot see a tweet that has already been published.
Bad: the uniform fan-out-on-write from Deep Dive 1, applied to every account equally. One tweet becomes 100M ZADD calls. Against the ~2.3M writes/sec the fan-out tier is sized for, that single tweet consumes roughly 43 seconds of the entire fleetβs fan-out capacity.
Three separate failures come out of that one number:
- Visibility. The tweet is live in the Tweet Store but absent from most timelines for tens of seconds. Followers refreshing during that window see nothing, then see it appear, which reads as a bug.
- Fairness. Everyone elseβs fan-out is queued behind it. A user with 40 followers posts and waits minutes, because a celebrity posted first. The systemβs latency for ordinary users is now a function of celebrity activity.
- Waste. Most of those 100M followers will not open the app before the tweet is stale, so most of the 100M writes buy nothing.
Worth naming what is not broken: FR2βs original read path handled this case perfectly, because posting cost one write no matter the follower count. We introduced this problem ourselves in Deep Dive 1, deliberately, to fix a worse one.
Good: Skip fan-out entirely for large accounts and fetch their tweets at read time β go back to FR2βs approach, but only for them. Correct, and it puts a scatter-gather back into the read path we just spent Deep Dive 1 removing.
The saving grace is the asymmetry: a user follows ~400 accounts but only a handful of mega-accounts, so the merge is over 5-20 sources rather than 400. That is a 20x smaller version of the problem we rejected, which is what makes it acceptable here and not there.
Great: Tier by follower count, so each account gets the strategy its size warrants.
- Regular (< 10K followers): push immediately. Bounded work, and reads stay a single range query
- Mid-tier (10K-1M): push, but on a lower-priority queue so it cannot starve the regular tier. Followers see a few seconds of extra delay, which nobody notices
- Mega (1M+): never push. Always pulled at read time
- Read-time merge: Feed Service keeps a small per-user list of mega-account followees, fetches their recent tweets from the Tweet Store, and merges with the pre-built timeline before ranking
Why 10K as the threshold? Pushing to 10K followers is a few thousand ZADDs β tens of milliseconds, invisible. Pushing to 50M takes minutes. The threshold is where fan-out cost stops being noise, and it is worth tuning against your actual fan-out capacity rather than treating 10K as a magic number.
3) Out of everything a user could see today, how do we choose what they do see?
Problem: FR2 sorted by timestamp and Deep Dive 1 preserved that ordering in the sorted-set score. A user following 400 accounts has far more eligible tweets per session than they will ever read.
In simple terms: Showing tweets purely by time means you miss the important ones that happened while you slept, and the accounts that tweet most often crowd out the ones you care about.
Bad: the pure chronological ordering the design currently ships. 400 followees posting even a few times a day each is well over a thousand eligible tweets, against roughly 50 a user actually reads per session. So most of what a user explicitly subscribed to never reaches them, and the selection is made entirely by posting frequency β an account that tweets 30 times a day will fill the timeline ahead of a close friend who tweets twice a week, regardless of which one the user cares about.
Good: A weighted score over a few obvious signals β recency, engagement counts β applied to the candidate set before returning it. Better than chronological, and identical for every user, so it optimizes for what is popular rather than what is relevant to this reader.
Great: Score candidates per reader on signals that include the relationship between reader and author.
flowchart LR
TWEET["Tweet candidate"]:::client
FRESH["Freshness<br/>newer is higher"]:::service
ENGAGE["Engagement<br/>likes retweets replies"]:::service
SOCIAL["Social closeness<br/>do you interact"]:::service
CONTENT["Content type<br/>media beats text"]:::service
SCORE["Weighted sum"]:::data
TWEET -->|"1. Freshness"| FRESH
TWEET -->|"2. Engagement"| ENGAGE
TWEET -->|"3. Closeness"| SOCIAL
TWEET -->|"4. Content type"| CONTENT
FRESH -->|"5. Weight"| SCORE
ENGAGE -->|"6. Weight"| SCORE
SOCIAL -->|"7. Weight"| SCORE
CONTENT -->|"8. Weight"| SCORE
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
score = w1 Γ freshness + w2 Γ engagement + w3 Γ social_closeness + w4 Γ content_type
Twitterβs production ranker is an ML model, originally βEarlybird,β but the weighted sum captures the idea and is defensible in an interview.
Keep the online model cheap. Ranking runs on every feed load, so at 100K loads/sec a 50ms model adds 50ms to a 200ms budget and needs thousands of cores. Score a few hundred candidates in well under a millisecond on CPU; heavy models belong in offline training, producing the weights and embeddings the cheap online path reads.
4) A followee tweets while the user is mid-scroll. Should it appear, and how does it get there?
Problem: Deep Dive 1 writes the tweet into the followerβs timeline within seconds of it being posted, but nothing tells a client that is already open. The user is looking at a list that is now out of date and has no way to know.
In simple terms: Youβre scrolling your feed. Someone you follow tweets. Should it pop in immediately and move the thing you were reading, or should you be told and left to decide?
Note that the product question comes first here, and it settles the engineering question. Injecting tweets into a list someone is actively reading moves their scroll position, which is hostile. What Twitter does is show a βnew tweets availableβ banner and load on tap. So the requirement is not βdeliver the tweet,β it is βdeliver a countβ β a much smaller thing, and it changes which transport is appropriate.
Bad: The client polls GET /feed every 30 seconds. At 500M DAU with even a fraction of sessions open, that is a continuous flood of requests whose answer is almost always βnothing new,β and each one re-runs the full timeline read and hydration to discover that. It also means the banner is up to 30 seconds late.
Good: Long polling. The client holds a request open and the server responds when something arrives. The wasted round trips disappear and the banner is prompt. The cost is a held connection per open session plus a request cycle after every response, which at this scale is a lot of connection churn for a payload that is usually a single integer.
Great: Push the count over a persistent channel, and keep it separate from the timeline read.
- When a client opens a session it subscribes to a channel for its own userId, over WebSocket or SSE.
π‘ SSE is a good fit when traffic is one-directional β the server pushes, the client never sends. It is plain HTTP, so it survives proxies that mishandle WebSocket upgrades. Learn more β - Fan-out Workers, having done the
ZADDin Deep Dive 1, also publish a lightweight βyou have new tweetsβ signal to that userβs channel - The client increments its banner count. It does not fetch anything yet
- On tap, the client issues a normal
GET /feedβ the same range read as always, no special path - If the channel drops, the client falls back to polling on a slow interval. Missing a banner is a cosmetic degradation, not a correctness one, which is what makes this safe to run on a best-effort transport
The design point worth keeping: we push the notification and pull the content. The push path carries a few bytes and can be lossy; the pull path is the one already built, tested and cached.
Core Flows
Flow 1: User posts a tweet
sequenceDiagram
autonumber
participant U as User
participant TS as Tweet Service
participant DB as Tweet Store
participant K as Kafka
participant FAN as Fan-out Workers
participant G as Social Graph
participant R as Redis Timelines
U->>TS: POST tweet
TS->>DB: store tweet
TS->>K: publish TweetCreated event
K->>FAN: consume
FAN->>G: get followers of poster
G-->>FAN: follower list
FAN->>FAN: filter out celebrities from push
FAN->>R: ZADD timeline:{followerId} tweetId for each follower
FAN->>R: ZREMRANGEBYRANK trim to latest 800
- Tweet stored permanently in Cassandra/DynamoDB.
- Event published to Kafka for async fan-out.
- Fan-out workers get the posterβs follower list from the social graph.
- For each non-celebrity follower, push the tweet ID into their Redis sorted set (scored by timestamp).
- Trim each timeline to 800 entries (older ones fall off; user can fetch from DB if they scroll far enough).
Flow 2: User opens their feed
sequenceDiagram
autonumber
participant U as User
participant F as Feed Service
participant R as Redis
participant DB as Tweet Store
participant RANK as Ranking Service
U->>F: GET /feed
F->>R: ZREVRANGE timeline:{userId} 0 50
R-->>F: cached tweet IDs
F->>DB: multi-get tweet details by IDs
DB-->>F: tweet objects
F->>F: fetch celebrity tweets on read (merge)
F->>RANK: rank and filter
RANK-->>F: sorted feed
F-->>U: feed response
- Read the userβs pre-built timeline from Redis (just tweet IDs, sorted by time).
- Hydrate: fetch full tweet objects from the tweet store.
- Merge in recent tweets from celebrities the user follows (fan-out on read for these).
- Apply ranking (relevance score, engagement signals, freshness decay).
- Return the ranked feed.
Final Architecture
flowchart TD
USERS["Users"]:::client
GW["API Gateway<br/>auth rate-limit"]:::edge
TS["Tweet Service"]:::service
FEED["Feed Service"]:::service
RANK["Ranking Service<br/>ML model"]:::service
FANOUT["Fan-out Workers"]:::service
TDB[("Tweet Store<br/>Cassandra")]:::data
GRAPH[("Social Graph<br/>who follows whom")]:::data
CACHE[("Timeline Cache<br/>Redis sorted sets")]:::data
K["Kafka<br/>tweet events"]:::async
MEDIA[("Media<br/>S3 plus CDN")]:::data
USERS -->|"POST or GET"| GW
GW -->|"Forward to tweet svc"| TS
GW -->|"Forward to feed svc"| FEED
TS -->|"Persist tweet"| TDB
TS -->|"Store media file"| MEDIA
TS -->|"Publish tweet event"| K
K -->|"Process tweet event"| FANOUT
FANOUT -->|"Lookup followers"| GRAPH
FANOUT -->|"Prepend to follower feeds"| CACHE
FEED -->|"Read pre-built feed"| CACHE
FEED -->|"Hydrate tweet details"| TDB
FEED -->|"Rank by relevance"| RANK
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef async fill:#AB47BC,stroke:#4A148C,color:#fff
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
How it works end-to-end (write path β posting a tweet):
- User posts tweet β request hits API Gateway (auth + rate limit applied)
- Tweet Service persists β writes tweet to Cassandra (Tweet Store), uploads media to S3/CDN
- Event emitted to Kafka β tweet creation event published for async fan-out
- Fan-out Workers distribute β reads Social Graph for follower list, pushes tweetId into each followerβs Redis sorted set (Timeline Cache)
- Celebrity exception β users with >500K followers skip fan-out; their tweets merged at read time
How it works end-to-end (read path β viewing feed):
- User opens feed β Feed Service checks Timeline Cache (Redis sorted set) for pre-built timeline
- Hydration β tweet IDs fetched from cache, hydrated with full tweet content from Cassandra
- Ranking Service scores β ML model re-ranks by freshness, engagement, and social closeness
- Response returned β ranked feed served to the user in <200ms P99
Interview Cheat Sheet
| Question | Answer |
|---|---|
| βHow do you build the feed?β | Hybrid fan-out: push for regular users, pull for celebrities |
| βWhereβs the timeline stored?β | Redis sorted set per user (tweet IDs scored by timestamp) |
| βHow do you handle celebrities?β | Donβt push to 50M followers. Merge their tweets at read time. |
| βHow do you rank?β | Weighted score: freshness + engagement + social closeness |
| βWhat about real-time?β | WebSocket for βnew tweets availableβ banner, not auto-inject |
| βStorage for tweets?β | Cassandra or DynamoDB - partition by tweetId, immutable, replicated |
| βSocial graph storage?β | Adjacency list in Redis or dedicated graph DB. followers:{userId} β Set<userId> |
| βWhatβs the read latency?β | P99 < 200ms. Pre-built cache β hydrate β rank. |
Key Technologies
| Term | What it is |
|---|---|
| Fan-out | Taking one event (a tweet) and delivering it to many recipients (followers). βFan-out on writeβ = push at creation time. βFan-out on readβ = pull at view time. |
| Redis Sorted Set | A Redis data structure that stores elements with a score. Lets you get the top-N elements efficiently (perfect for βlatest 50 tweetsβ). |
| Social Graph | The network of who-follows-whom. Stored as adjacency lists. Queried as βgive me all followers of user X.β |
| Kafka | Event streaming platform. Tweet creation events go here for async fan-out workers to consume. |
| Cassandra | Wide-column NoSQL database. Stores tweets durably. Good for high write volume and partition-per-user access patterns. |
| CDN | Content Delivery Network. Serves media (images, videos) from edge servers close to users. |
| Hydration | Converting a list of IDs into full objects. βHydrate tweet IDs β fetch full tweet with text, likes, media URLs.β |
Whatβs Expected at Each Level
This section helps you calibrate your depth. You donβt need to cover everything - just know whatβs expected for your level.
Mid-level
Produce a working design with tweet storage and basic feed assembly. Recognize that JOIN-based feed queries donβt scale. With prompting, propose pre-computing timelines (fan-out on write) so that feed reads are a simple cache lookup rather than a complex multi-table query.
Senior
Articulate the fan-out-on-write vs fan-out-on-read tradeoff without prompting. Propose the hybrid approach for celebrities (>10K followers skip fan-out, merged at read time). Discuss Redis sorted sets or Cassandra for timeline cache. Explain how to handle the celebrity problem (50M followers) and why naive fan-out would generate 50M writes per tweet.
Staff+
Address feed ranking vs chronological ordering trade-offs and the ML pipeline needed for relevance scoring. Discuss real-time feed injection (new tweets appearing without refresh via WebSocket/SSE), tweet deletion propagation across cached timelines, and the operational cost of fan-out at Twitter scale (500M users Γ 400 followers = 200B cache writes/day). Cover cache eviction strategies for inactive users.
π― Key Takeaways
- Fan-out on write pre-computes feeds - reads are instant
- Celebrity exception skips fan-out for >500K followers - merged at read time
- Kafka decouples posting from feed distribution
- Redis caches hot timelines for active users
Related Designs
- Chat System - real-time message delivery
- Notification System - push to users
- Leaderboard - real-time ranking updates
Related Concepts
Understand the building blocks used in this design:
- Fan-Out Patterns β β pushes each new tweet into follower timelines (fan-out-on-write) with a pull path for celebrities
- Caching β β precomputed timelines live in Redis for millisecond reads
- Database Sharding β β partitions tweets and timelines across nodes to handle write volume
- CDN β β serves attached images and video close to the viewer
Discussion
Newest first