News feed system design is the problem of building a service that shows each user a fresh, ordered list of posts from the accounts they follow, at the scale of Facebook, Instagram, or X (Twitter). A strong interview answer centers on one trade-off: precompute timelines when posts are written (fan-out on write), build them when feeds are read (fan-out on read), or use a hybrid that pushes posts from normal users and pulls posts from celebrities. Add cursor pagination, a layered cache, and a basic ranking pipeline, and you have a complete answer.
This guide walks through that answer the way you would deliver it in a 45 to 60 minute round, with real numbers at every step.
Key Takeaways
- Feeds are read-heavy. Twitter has publicly described read traffic far above write traffic, on the order of 300K timeline reads per second against a few thousand tweets per second in its 2012-era talks, so design for cheap reads first.
- Fan-out on write gives one-lookup reads, but a single post from an account with 100 million followers becomes 100 million cache writes.
- The hybrid model (push for most users, pull for accounts above a follower threshold) is the answer interviewers expect you to reach, and you should be able to justify the threshold with math.
- Use cursor pagination, never offsets. New posts arriving at the top will corrupt offset-based paging.
- Senior candidates should sketch ranking as candidate generation, scoring with predicted engagement, then post-ranking rules.
What Requirements Should You Clarify First?
Start by pinning down scope. The feed problem is broad, and interviewers want to see you narrow it before drawing boxes. This step matters in every round, which our system design interview guide covers in more detail.
Functional requirements:
- Users can create posts with text and optional images or video.
- Users can follow and unfollow other users.
- Users see a home feed of posts from accounts they follow, newest or most relevant first.
- The feed supports infinite scroll (pagination).
Non-functional requirements:
- Feed load latency under roughly 200 ms at p99 for the feed service itself.
- High availability over strict consistency. A post can take a few seconds to appear in followers' feeds.
- Durability for posts. Losing a cached timeline is acceptable; losing a post is not.
Out of scope unless asked: comments, direct messages, search, ads, and notifications. If the interviewer wants notifications, the notification system design is its own full problem.
Scale Estimates: How Big Is This System?
Back-of-the-envelope math decides your architecture, so do it out loud. Use round numbers and state your assumptions.
| Assumption | Value |
|---|---|
| Daily active users (DAU) | 500 million |
| Feed loads per user per day | 10 |
| Posts per user per day | 0.2 |
| Average followers per user | 200 |
| Peak to average ratio | 3x |
From these:
- Feed reads: 500M x 10 = 5 billion per day, about 58K QPS on average and roughly 175K QPS at peak.
- Post writes: 500M x 0.2 = 100 million per day, about 1.2K QPS average and 3.5K QPS at peak.
- Read to write ratio: roughly 50:1. This is the single most important number in the design.
- Fan-out inserts (if pushing everything): 100M posts x 200 followers = 20 billion timeline inserts per day, about 230K inserts per second on average.
- Post storage: at about 1 KB of metadata per post, 100 GB per day, or about 36 TB per year before replication. Media lives separately in object storage.
- Timeline cache: storing the newest 800 post IDs per user at 8 bytes each is about 6.4 KB per user. For 500M users, that is roughly 3.2 TB of memory, before replication and data structure overhead.
These numbers line up with the shape of real systems. In a 2012 talk on timelines at scale, Twitter's then VP of Engineering Raffi Krikorian described roughly 300K QPS of timeline reads against a few thousand tweets per second, with tweet IDs fanned out into Redis-backed home timelines. Treat those figures as approximate, from memory of a 2012 talk. The ratio is what matters, and it tells you to make reads cheap.
API and Data Model
Keep the API small. Two endpoints carry the whole design.
POST /v1/posts
Body: { "text": "...", "media_ids": ["m_123"] }
Returns: { "post_id": "1849204567382016000", "created_at": "..." }
GET /v1/feed?cursor=<opaque_cursor>&limit=20
Returns: { "posts": [...], "next_cursor": "..." }
The data model has three core entities:
posts (
post_id BIGINT PRIMARY KEY, -- time-sortable 64-bit ID
author_id BIGINT,
text VARCHAR(2000),
media_keys TEXT,
created_at TIMESTAMP
)
follows (
follower_id BIGINT,
followee_id BIGINT,
created_at TIMESTAMP,
PRIMARY KEY (follower_id, followee_id)
)
-- second index or table keyed by followee_id for "who follows X?"
Use time-sortable IDs such as the Snowflake format Twitter popularized: a 64-bit integer that combines a millisecond timestamp, a machine ID, and a sequence number. Sorting by ID then means sorting by time, which simplifies merging and pagination.
Store posts in a horizontally scalable database (Cassandra or sharded MySQL are both defensible) partitioned by post_id or author_id. The follow graph needs both directions: "who do I follow" for fan-out on read and "who follows me" for fan-out on write. Media goes to object storage and is served through a CDN.
How Does Post Creation Work?
Post creation should return fast and push the expensive work to the background. The write path looks like this:
- The client uploads media directly to object storage using a pre-signed URL and gets back media keys.
- The client calls
POST /v1/posts. The post service generates an ID, writes the post to the posts database, and writes it to the post cache. - The post service publishes a
PostCreatedevent to a message queue such as Kafka. - Fan-out workers consume the event and decide how to distribute the post (covered next).
- The API returns success after step 3. Followers see the post seconds later.
Returning before fan-out completes is a deliberate choice. It keeps write latency low and accepts eventual consistency, which you already agreed to in the requirements. Put a rate limiter in front of the post endpoint so spam bursts do not flood the fan-out pipeline.
Fan-Out on Write vs Fan-Out on Read
This is the heart of the interview. Name both approaches, compare them with numbers, then propose the hybrid.
Fan-out on write (push) means that when a user posts, workers look up all of the user's followers and insert the new post ID into each follower's precomputed timeline in cache. Reading a feed becomes one cache lookup.
Fan-out on read (pull) means posts are stored once. When a user opens the feed, the system fetches recent posts from every account they follow, merges them, and returns the top results.
| Factor | Fan-out on write (push) | Fan-out on read (pull) |
|---|---|---|
| Feed read cost | 1 cache lookup | N lookups (N = accounts followed) plus a merge |
| Post write cost | F inserts (F = follower count) | 1 insert |
| Read latency | Very low | Higher, grows with follow count |
| Celebrity posts | Catastrophic (millions of inserts) | Fine |
| Inactive followers | Wasted work | No waste |
| Freshness | Seconds of delay during fan-out | Always current |
| Memory | High (timeline per user) | Low |
With a 50:1 read to write ratio, push looks obvious. It wins for the typical user. The celebrity problem is what breaks it.
The Celebrity Problem: A Worked Example
The celebrity problem (also called the hot-user problem) is the cost explosion that happens when an account with a huge follower count posts under fan-out on write. Here is the math you should be able to produce on the whiteboard.
Pure push, one celebrity post: take an account with 100 million followers. One post means 100 million timeline inserts. If your fan-out fleet sustains 1 million inserts per second in total, that single post occupies the entire fleet for 100 seconds. Every other user's posts queue behind it. If the account posts 10 times a day, that is 1 billion inserts, five percent of your entire daily fan-out budget from one account. Worse, a large share of those followers will not open the app before the post scrolls out of their cached timeline, so much of the work is wasted.
Pure pull, every feed load: if the average user follows 300 accounts, each feed load needs about 300 lookups plus a merge. At 175K peak QPS, that is about 52 million lookups per second. The cost lands on the latency-sensitive read path, which is exactly where you cannot afford it.
Hybrid: push for accounts below a follower threshold, pull for accounts above it. Suppose the threshold is 1 million followers and the average user follows 5 such accounts.
- Write side: celebrity posts are written once to a hot-post cache, not fanned out. The fan-out fleet only handles normal accounts.
- Read side: each feed load does 1 timeline lookup plus about 5 celebrity lookups. At 175K QPS, that is roughly 875K extra lookups per second, all against a tiny, extremely hot dataset of recent celebrity posts that you can replicate across many cache nodes.
The threshold is a tunable, not a magic number. Lower it and reads get more expensive; raise it and fan-out spikes return. Saying this out loud signals that you understand the trade-off rather than reciting a pattern.
The hybrid design
WRITE PATH
Client --> API Gateway --> Post Service --> Posts DB + Post Cache
|
v
Kafka: PostCreated
|
v
Fan-out Workers
/ \
followers < threshold? followers >= threshold?
push post_id into write to Hot Post Cache
each active follower's (replicated, keyed by
timeline (Redis) author_id)
READ PATH
Client --> API Gateway --> Feed Service
|
+--------------------+---------------------+
| | |
Timeline Cache Celebrity list for Hot Post Cache
(pushed post IDs) this user (graph) (recent celeb posts)
\ | /
+----------> merge + dedupe <------------+
|
Ranking Service
|
hydrate from Post Cache
|
page + next_cursor --> Client
Two refinements make the hybrid stronger:
- Skip inactive users. Do not push into timelines of users who have not logged in for, say, 30 days. When they return, rebuild their timeline with a one-time pull. This cuts a large share of wasted inserts.
- Cap timeline length. Keep only the newest several hundred IDs per user. Deeper scrolling falls back to querying the posts database.
Fan-out math is easy to lose when an interviewer pushes back mid-whiteboard. TechScreen is an invisible AI interview assistant that helps you keep capacity estimates and trade-offs straight in real time, and it stays hidden during screen shares on Zoom, Google Meet, and Teams. Start with 3 free tokens, no credit card.
How Does Feed Ranking Work?
Feed ranking is the step that orders candidate posts by predicted relevance instead of time. Interviewers at companies with ranked feeds, including Meta, increasingly ask senior candidates to sketch it. You do not need model internals. You need the pipeline.
- Candidate generation. Collect posts from the pushed timeline, pulled celebrity posts, and optionally out-of-network sources such as recommended accounts. When X open-sourced its recommendation code in 2023, its engineering blog described pulling about 1,500 candidates per request, split roughly evenly between in-network and out-of-network posts.
- Feature fetching. For each candidate, gather features: author affinity (how often the viewer interacts with the author), post age, engagement so far, media type, and the viewer's recent behavior.
- Scoring. A model predicts probabilities such as P(like), P(reply), P(share), and P(hide or report). The final score is a weighted sum, with negative weights on negative signals.
- Post-ranking rules. Apply filters (blocked users, deleted posts, policy violations) and diversity rules, such as no more than two consecutive posts from the same author.
A simple scoring function you can write on the whiteboard:
def score(post, viewer):
p = model.predict(post, viewer)
raw = (1.0 * p["like"] + 5.0 * p["reply"] + 3.0 * p["share"]
- 20.0 * p["hide"])
age_hours = (now() - post.created_at).total_seconds() / 3600
return raw / (1 + age_hours) ** 1.5
The weights are illustrative. The point is to show that ranking combines predicted engagement with a freshness decay, and that negative feedback carries heavy weight. If the role is ML-focused, expect to go deeper; our machine learning engineer interview guide covers recommendation system rounds.
Pagination: Why Cursors Beat Offsets
Offset pagination (?page=3&limit=20) breaks on feeds. If ten new posts arrive while the user reads page one, page two's offset now points at posts they already saw, and they get duplicates.
Cursor pagination fixes this. Two common approaches:
- Chronological feeds: the cursor is the last post ID returned. The next request asks for posts with IDs lower than that cursor. Time-sortable IDs make this a simple range query.
- Ranked feeds: scores change between requests, so ranking again on every page can reorder posts and cause duplicates. Instead, rank once per session, store the ordered list of post IDs in cache for a short TTL, and make the cursor a reference to that snapshot plus a position.
New posts that arrive while the user is scrolling should not be inserted mid-scroll. Show a "new posts" indicator at the top instead, and load them when the user taps it or pulls to refresh.
Caching Strategy
A news feed is mostly a caching problem. Use separate caches for separate access patterns, which is the same thinking behind a distributed cache design.
| Cache | Contents | Key | Notes |
|---|---|---|---|
| Timeline cache | Post IDs per user | user_id | Redis sorted set or list, capped length |
| Post cache | Post objects | post_id | Hydrates IDs into full posts |
| Hot post cache | Recent celebrity posts | author_id | Heavily replicated to spread load |
| Social graph cache | Follower and followee lists | user_id | Used by fan-out and pull |
| Counter cache | Likes, replies, shares | post_id | Updated asynchronously, eventually consistent |
| CDN | Images and video | media key | Offloads nearly all bandwidth |
Three failure modes to raise before the interviewer does:
- Hot keys. A viral post can overload one cache shard. Replicate hot keys across nodes, or add a small in-process cache on feed servers with a TTL of a few seconds.
- Thundering herd. When a hot key expires, thousands of requests miss at once. Use request coalescing so only one request rebuilds the value.
- Cache loss. Timelines are derived data. If a Redis node dies, rebuild affected timelines from the posts database and follow graph on the next read. This is why the posts database, not the cache, is the source of truth.
Common Follow-Up Questions
Once the core design is on the board, interviewers probe edges. Have a short answer ready for each.
- What happens when a user follows someone new? Backfill: pull that account's recent posts and merge them into the follower's timeline asynchronously.
- What about unfollows and blocks? Filter at read time. Removing IDs from every affected timeline is expensive and unnecessary; the read path already checks the viewer's block and follow lists.
- How do you delete or edit a post? Mark the post deleted in the posts database and post cache. Timelines keep the ID, and hydration drops it. Edits update only the post object, since timelines store IDs, not content.
- How do you support multiple regions? Keep each user's timeline in their home region, replicate posts across regions asynchronously, and route reads to the nearest replica.
- How do you push new posts in real time? Use a persistent connection such as WebSockets or server-sent events to send a lightweight "new posts available" signal, not the posts themselves. The chat app design covers connection management at scale.
- How do you insert ads? Treat ads as another candidate source with its own ranking and frequency caps, blended after organic ranking.
- How would you measure success? Feed load latency, fan-out lag (time from post to appearance in followers' feeds), and engagement metrics evaluated through A/B tests.
At the senior and staff levels, the follow-ups are where the round is decided. Interviewers want to see you weigh cost, latency, and operational complexity without being prompted, a skill the staff engineer interview guide treats in depth.
How to Practice This Question
Practice the news feed design out loud, end to end, with a timer. Aim for this pacing in a 45 to 60 minute round:
- Requirements and estimates: 5 minutes.
- API, data model, and high-level design: 10 minutes.
- Fan-out strategy, celebrity problem, and caching: 20 minutes.
- Ranking, pagination, and follow-ups: the remaining time.
Draw the hybrid diagram from memory until you can do it in under three minutes. Then rehearse the celebrity math until the numbers come out without hesitation. For a structured routine, see how to practice system design interviews.
Want a safety net for the real round? TechScreen gives you invisible, real-time AI help with fan-out trade-offs, capacity math, and follow-up questions during your news feed system design interview, hidden from screen shares on Zoom, Google Meet, and Teams. Try it with 3 free tokens, no credit card required.
Frequently Asked Questions
What is the difference between fan-out on write and fan-out on read?
Fan-out on write (push) copies a new post's ID into every follower's precomputed timeline when the post is created, so reading a feed is a single cache lookup. Fan-out on read (pull) stores posts only once and builds the feed at request time by querying every account the reader follows and merging the results. Push makes reads cheap and writes expensive. Pull makes writes cheap and reads expensive. Most production feeds combine the two.
How do you handle celebrities with millions of followers in a news feed?
Exclude them from fan-out on write. Accounts above a follower threshold have their posts stored once in a hot, heavily replicated cache. When a user loads their feed, the system reads their precomputed timeline and then pulls recent posts from the handful of celebrities they follow, merging both lists before ranking. This turns a write of millions of timeline inserts into a few cheap reads per feed request.
Should a news feed use offset or cursor pagination?
Use cursor pagination. Offset pagination breaks on a feed because new posts arrive at the top while the user scrolls, which shifts every offset and causes duplicates or skipped items. A cursor encodes a stable position, such as the last post's score and ID or a reference to a cached ranked snapshot, so the next page continues exactly where the previous one ended regardless of new writes.
Do I need to talk about machine learning ranking in a news feed interview?
For mid-level roles, a chronological feed with a brief note on ranking is usually enough. Senior candidates are increasingly expected to sketch a ranking pipeline: candidate generation, feature fetching, a model that predicts engagement probabilities, a weighted score, and post-ranking rules such as author diversity and filtering. You do not need model architecture details unless you are interviewing for an ML role.
What database should I use for a news feed?
There is no single correct answer, but a common split is a horizontally scalable store such as Cassandra or sharded MySQL for posts, a separate store for the follow graph indexed in both directions, Redis or a similar in-memory store for precomputed timelines, and object storage behind a CDN for images and video. Justify each choice by its access pattern rather than naming technologies.
How long should a news feed system design answer take?
In a typical 45 to 60 minute round, spend about 5 minutes on requirements and estimates, 10 minutes on the API, data model, and high-level design, 20 minutes on deep dives such as fan-out strategy, the celebrity problem, and caching, and the remaining time on ranking, failure modes, and the interviewer's follow-up questions.
Ready to use AI assistance in your next interview?
TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.
Ace your next interview →