Chat app system design is the interview question where you design a messenger like WhatsApp: persistent connections for real-time delivery, durable per-conversation message storage, an offline inbox, group fan-out, presence, and sent, delivered, and read receipts. A strong answer centers on one idea. Messages are delivered at least once and deduplicated, ordered per conversation, and acknowledged at every hop so that each delivery state maps to a concrete event in the system.
Key Takeaways
- Scope first. Agree on 1:1 chat, groups (cap the size, WhatsApp allows up to 1,024 members), presence, and receipts before drawing anything. Media, calls, and search are usually out of scope.
- WebSockets win for chat because one full-duplex connection carries messages, acks, and receipts in both directions. Treat long polling as a fallback, not the design.
- Order per conversation, never globally. A per-conversation sequence number gives clients a sort key and a way to detect gaps.
- Each delivery state is a separate ack. Sent means the server persisted it, delivered means the recipient device stored it, and read means the user opened the chat.
- Groups need a fan-out decision. Small groups fan out on write to per-user inboxes; very large channels are better served by fan-out on read.
- E2EE changes the server's job. With Signal-style encryption, the server routes and stores ciphertext and never sees content, which rules out server-side search and moderation of message bodies.
What Requirements Should You Clarify First?
Your first five minutes decide the shape of the interview. If you need a refresher on pacing the full 45 minutes, the framework in how to ace a system design interview applies directly here.
Functional requirements
- 1:1 messaging between users, text first.
- Group chats with a size cap (ask; 256 to 1,024 is a reasonable range).
- Message delivery states: sent, delivered, read.
- Online presence and last-seen timestamp.
- Offline users receive everything they missed when they reconnect, on every device they own.
Non-functional requirements
- Low latency: a message should reach an online recipient in well under a second.
- Durability: once the sender sees "sent," the message must never be lost.
- Per-conversation ordering. Global ordering is not required.
- High availability over strict consistency for presence. A stale "online" dot is fine; a lost message is not.
- End-to-end encryption, at least at a high level.
Out of scope unless the interviewer pushes: calls, search, stories, payments, and spam detection. Naming what you skip is a senior signal.
Back-of-the-envelope estimates
Use round numbers and state them as assumptions.
| Assumption | Value | Implication |
|---|---|---|
| Daily active users | 500 million | Size of the session registry |
| Peak concurrent connections | ~200 million | Gateway fleet size |
| Connections per gateway server | ~500K to 1M (tuned) | Roughly 200 to 400 gateway hosts at peak |
| Messages per user per day | 40 | 20 billion messages per day |
| Average write rate | 20B / 86,400 s | ~230K messages per second, plan for 3x at peak |
| Average message size with metadata | ~200 bytes | ~4 TB of new message data per day before replication |
Chat is a connection-holding problem first and a storage problem second.
High-Level Architecture
The core design has six components. Draw them in this order and explain the arrows as you go.
- Clients (phone, desktop, web) hold one persistent connection each.
- Load balancer terminates TLS and routes new connections to a gateway using least-connections, not round robin.
- Gateway (connection) service holds WebSocket connections. It is stateless apart from the sockets it owns.
- Session registry maps
user_id + device_idto the gateway currently holding that device's socket. A Redis cluster with TTL heartbeats is a common choice. - Chat service validates, assigns sequence numbers, persists messages, and routes them to recipient gateways.
- Storage: a message store (wide-column), an inbox or sync queue per device, and a relational store for users, groups, and memberships.
Add a push notification service (APNs and FCM) for devices with no live connection. The notification system design guide covers that component in depth, so you can treat it as a black box here.
Connection Layer: WebSockets vs Long Polling vs SSE
A WebSocket is a single TCP connection, upgraded from HTTP, that both sides can write to at any time. The protocol is defined in RFC 6455. For chat, that bidirectional channel is the deciding factor, but interviewers like to hear you compare the alternatives.
| Criterion | WebSockets | Long polling | Server-sent events (SSE) |
|---|---|---|---|
| Direction | Full duplex | Client pulls, server holds request open | Server to client only |
| Latency for new message | Lowest, pushed immediately | Low, but a new request is needed after every response | Low for server push |
| Client sending messages | Same socket | Separate HTTP request | Separate HTTP request |
| Overhead per message | Small frame header | Full HTTP headers per poll cycle | Small, but sends need full HTTP requests |
| Proxy and firewall friendliness | Mostly fine today, occasional corporate proxy issues | Works almost everywhere | Works over plain HTTP |
| Server state | Long-lived socket per device | Short-lived, easier to load balance | Long-lived stream per device |
| Typing indicators and receipts | Natural fit | Clunky and chatty | Needs a second channel for upstream |
| Best use in a chat design | Primary transport | Fallback for hostile networks | Read-only feeds, live notifications |
Verdict to say in the interview: WebSockets as the primary transport, long polling as a degraded fallback, and SSE only if the product were read-heavy and one-directional. WhatsApp itself predates browser WebSockets on mobile and historically ran a custom protocol over persistent TCP connections. The principle is the same: one long-lived, bidirectional connection per device.
Scaling the gateway tier
Gateways are where most interview follow-ups land. Cover these four points:
- Routing. When the chat service needs to deliver to user B, it looks up B's devices in the session registry, then forwards to those gateways over an internal RPC or pub/sub channel.
- Heartbeats. Clients ping every 20 to 60 seconds. Missed heartbeats close the socket and expire the registry entry, which also feeds presence.
- Deploys and failures. A gateway restart drops hundreds of thousands of sockets at once. Clients reconnect with jittered exponential backoff so the fleet does not get hit by a thundering herd. Rate-limit reconnects at the load balancer; the rate limiter design guide explains the token bucket you would use.
- Memory per connection. Memory and file descriptors run out before CPU does, which is why lightweight concurrency models (Erlang processes, goroutines, event loops) suit gateways.
How Are Messages Stored and Ordered?
Message storage is a write-heavy, append-mostly workload whose dominant read is "give me the latest N messages in this conversation." That access pattern points to a wide-column store such as Cassandra or ScyllaDB, partitioned by conversation.
CREATE TABLE messages (
conversation_id bigint,
bucket int,
seq bigint,
message_id bigint,
sender_id bigint,
client_msg_id uuid,
body blob,
created_at timestamp,
PRIMARY KEY ((conversation_id, bucket), seq)
) WITH CLUSTERING ORDER BY (seq DESC);
The bucket (for example, a time window or a block of sequence numbers) caps partition size so a very active group does not grow one partition forever. Discord's engineering team describes almost exactly this layout, with channel ID plus a static time bucket as the partition key, in How Discord Stores Trillions of Messages.
Sequence numbers, not timestamps
Ordering is per conversation. Wall-clock timestamps from different servers drift, so two messages sent milliseconds apart can sort incorrectly. Use a per-conversation sequence number instead:
- Route every write for a conversation to one owner (a partition of the chat service chosen by hashing
conversation_id), which increments a counter. This is the simplest correct answer. - Alternatively, use an atomic
INCRon a counter in a strongly consistent store. This is simpler to operate but becomes a hot key for busy groups. - Keep a separate globally unique, time-sortable
message_id(Snowflake-style: timestamp, worker ID, sequence) for references, replies, and deduplication across systems.
Clients sort by seq. If a client holds seq 41 and receives 43, it knows 42 is missing and fetches the gap. That gap detection is what makes the sync protocol self-healing.
Message Delivery States, End to End
This is the section most candidates hand-wave. Each checkmark in the UI corresponds to a specific durable event, and you should be able to draw all three.
Sent (one check): the server has durably stored the message and acknowledged the sender.
Sender A Gateway A Chat Service Message Store Inbox(B)
| SEND(cmid=u1, conv=7, body) | | |
|--------------->| | | |
| |--------------->| | |
| | |-- dedupe on cmid -->| |
| | |-- assign seq=42 --->| write msg |
| | |---------------------------------->| append(42)
| |<-- ACK(cmid=u1, seq=42, msg_id) -----| |
|<---------------| | | |
| UI: one check (SENT) | | |
Delivered (two checks): the recipient's device has received and stored the message locally.
Chat Service Session Registry Gateway B Recipient B
|-- where is B? -->| | |
|<-- gw-B-17 ------| | |
|-- PUSH(conv=7, seq=42) -------------->| |
| | |----------------->|
| | | | store locally
| | |<-- DELIVERED(conv=7, seq=42)
|<----------------------------------------| |
| mark delivered for B; remove 42 from Inbox(B) |
|-- RECEIPT(delivered, seq<=42) --> Gateway A --> Sender A |
| UI: two checks (DELIVERED) |
If B is offline, the registry has no entry. The message stays in B's inbox, the chat service triggers a push notification, and the delivered receipt waits until B reconnects and syncs.
Read (blue checks): the user opened the conversation and the message was on screen.
Recipient B Gateway B Chat Service Gateway A Sender A
| opens conv 7 | | | |
|-- READ(conv=7, up_to_seq=42) ----->| | |
| |---------------->| update read_seq(B,7)=42 |
| | |-- RECEIPT(read, seq<=42) ------->|
| | | |-------------->|
| | | | UI: blue checks
Two design choices make this efficient:
- High-water marks. Receipts carry
up_to_seq, not individual message IDs. One read receipt covers every earlier message in the conversation. - Receipts are messages too. They ride the same transport and the same at-least-once machinery, but they are idempotent (setting
read_seq = max(read_seq, 42)), so duplicates are harmless.
| State | Trigger event | Who acks | Stored where |
|---|---|---|---|
| Pending (clock icon) | User taps send | Nobody yet | Sender's local outbox |
| Sent | Server persisted the message | Chat service to sender | Message store and recipient inbox |
| Delivered | Recipient device stored it | Recipient device to server | Per-recipient delivery cursor |
| Read | Recipient opened the chat | Recipient device to server | Per-recipient read cursor |
For groups, "delivered" and "read" become per-member cursors, and the sender's UI shows the aggregate (for example, all members have read it).
Can you draw these three sequence diagrams from memory while an interviewer asks follow-ups about duplicate retries? TechScreen is an invisible AI interview assistant that stays hidden during Zoom, Google Meet, and Teams screen shares and gives you real-time prompts for system design rounds like this one. Start with 3 free tokens, no credit card required.
Delivery Guarantees and Offline Sync
No distributed chat system delivers exactly once over an unreliable network. What you build is at-least-once delivery plus idempotent processing, which the user experiences as exactly once.
On the sending side:
- The client generates
client_msg_id(a UUID) and stores the message in a local outbox. - It retries with backoff until it gets an ack carrying that ID.
- The server checks
client_msg_idagainst a short-lived dedupe cache (or a unique index) before assigning a sequence number. A retry of an already-stored message returns the original ack.
On the receiving side:
- Every message is written to the recipient's per-device inbox before the push attempt.
- The inbox entry is removed only after the device acks. If the ack is lost, the server pushes again, and the client drops the duplicate by
seq.
Offline sync on reconnect:
client -> SYNC { device_id, cursors: { conv_7: 41, conv_9: 1203 } }
server -> for each conversation with seq > cursor:
return messages in pages of 100, oldest first
client -> ACK highest seq per conversation
server -> trim inbox entries <= acked seq
Mention multi-device explicitly: each device keeps its own cursor and inbox entry, so a phone offline for a week catches up without affecting the desktop client. Set a retention limit on undelivered items and say what happens when it expires.
Most reads hit the last few dozen messages, so cache recent history; the distributed cache design guide covers eviction and invalidation.
Group Fan-Out
Group messaging is where the design branches. The question is whether the server copies the message reference into every member's inbox at write time or members pull from the shared conversation at read time.
| Approach | How it works | Strength | Weakness | Use for |
|---|---|---|---|---|
| Fan-out on write | Store once, append a pointer to each member's inbox, push to each online member | Fast reads, simple sync | Write cost grows with group size | Groups up to ~1,000 members |
| Fan-out on read | Store once in the conversation; members fetch since their cursor | Constant write cost | Reads are heavier, harder to push | Broadcast channels with very large audiences |
| Hybrid | Push to online members, others pull on reconnect | Balances both | More code paths | Large groups with many inactive members |
For a WhatsApp-style product with capped group sizes, fan-out on write is the right default. With the 1,024-member cap, one message creates at most about a thousand inbox appends, which a queue (Kafka or similar) absorbs easily. If you have seen the celebrity problem in news feed system design, point out that the same trade-off appears here for broadcast channels.
Keep group membership in a relational store and cache the member list. Membership changes (adds, removes, admin changes) are themselves ordered events in the conversation, so a removed member stops receiving messages from a specific sequence number onward.
Presence and Last-Seen
Presence is a high-write, low-value signal. Design it to be cheap and eventually consistent.
- Source of truth: the gateway heartbeat. A live socket with recent heartbeats means online. Store
user_id -> {status, last_seen}in an in-memory store with a TTL slightly longer than the heartbeat interval. - Do not broadcast every change to every contact. A user with 500 contacts going online and offline repeatedly would generate enormous traffic. Instead, push presence only to users who currently have a conversation with that person open, and let others fetch it lazily.
- Debounce flapping. Mobile networks drop constantly. Wait a few seconds before marking someone offline so a brief reconnect does not produce an offline-online blip.
- Respect privacy settings. Last-seen visibility (everyone, contacts, nobody) is checked at read time, not stored per viewer.
- Typing indicators are ephemeral: send them over the socket, never persist them, and drop them if the recipient is offline.
End-to-End Encryption at a High Level
End-to-end encryption means only the sender and recipient devices hold the keys needed to decrypt a message; the server relays ciphertext it cannot read. WhatsApp uses the Signal Protocol, and its security whitepaper describes the design. In an interview, three or four sentences are enough:
- Each device publishes public identity keys and a batch of one-time prekeys to a key server. A sender fetches the recipient's prekey bundle to start a session without the recipient being online.
- The Double Ratchet algorithm derives a new key for every message, which gives forward secrecy.
- For groups, the whitepaper describes Sender Keys: each member shares a sender key over pairwise sessions, then encrypts each group message once, and the server fans out a single ciphertext.
- Multi-device means each device has its own keys, so the sender encrypts for every device of every recipient.
Then state the consequences: the server cannot index bodies for search, cannot moderate content, and cannot restore history unless the client uploads an encrypted backup.
How Do Interviewers Grade This Question?
Most interviewers score the chat question on a handful of signals. This rubric reflects common patterns in system design loops at companies such as Meta (see the Meta interview process guide), not any one company's official scorecard.
| Signal | Mid-level bar | Senior bar |
|---|---|---|
| Requirements | Lists core features | Caps group size, names non-goals, sets durability vs presence consistency |
| Connection layer | Picks WebSockets | Justifies vs long polling and SSE, handles reconnect storms and gateway routing |
| Storage | Picks a NoSQL store | Partition key with buckets, per-conversation sequence numbers, gap detection |
| Delivery | Mentions acks | Walks sent, delivered, read end to end with idempotent retries and per-device cursors |
| Groups | Mentions fan-out | Chooses write vs read fan-out by group size, handles membership events |
| Depth on demand | Answers follow-ups | Proactively flags hot partitions, multi-region, and E2EE consequences |
Senior and staff candidates are expected to drive the trade-off discussion without prompting. The staff engineer interview guide explains how that expectation changes the conversation.
Common Follow-Up Questions
Prepare one or two sentences for each of these:
- "What if a gateway dies mid-delivery?" The message is already in the inbox, the device reconnects to a new gateway, syncs from its cursor, and receives it. Nothing is lost because persistence happens before push.
- "How do you handle a hot group?" Shard the sequence counter owner, batch inbox appends, and move very large audiences to fan-out on read.
- "Multi-region?" Home each conversation in one region for ordering; route cross-region messages over a replicated queue. Users connect to the nearest gateway.
- "How would you add media?" Upload media directly to object storage with a pre-signed URL, then send a normal message containing the media reference and (with E2EE) the decryption key.
How to Practice This Design
Practice out loud against a timer. Draw the architecture in under five minutes, then spend most of your time on delivery states and fan-out, where follow-ups go. The drills in how to practice system design interviews work well here, and the backend engineer interview guide shows where this question fits in a typical loop.
A good final check: can you explain, without notes, why a message is never lost if a gateway crashes after the sender sees one check? If yes, you understand the design.
System design rounds reward candidates who can recall the right trade-off table under pressure. TechScreen runs invisibly during your screen share and gives real-time structure for questions like "design WhatsApp," from capacity estimates to delivery guarantees. Try it on a mock interview with 3 free tokens.
Frequently Asked Questions
How do you design a chat app like WhatsApp in a system design interview?
Start by scoping requirements: 1:1 chat, group chat, online presence, and sent, delivered, and read receipts. Then estimate scale, and propose persistent WebSocket connections through a stateless gateway tier, a session registry that maps users to gateways, a chat service that assigns per-conversation sequence numbers, a wide-column message store partitioned by conversation, and a per-user inbox for offline delivery. Finish with group fan-out, presence, and a short end-to-end encryption overview.
Should a chat app use WebSockets, long polling, or server-sent events?
WebSockets are the default for chat because the connection is full-duplex: the client sends messages and acks on the same socket that the server uses to push new messages. Long polling works through any proxy but adds latency and reconnect overhead on every message. Server-sent events are simple for server-to-client streams but only flow one way, so the client needs separate HTTP calls to send. Mention long polling as a fallback for restrictive networks.
How does a chat system guarantee message delivery?
Chat systems provide at-least-once delivery plus deduplication, which looks exactly-once to the user. The client attaches a unique client message ID and retries until the server acks. The server writes the message durably before acknowledging, then pushes it to the recipient and keeps it in the recipient's inbox until the recipient device acks. Duplicates from retries are dropped by checking the client message ID, usually with a unique constraint or idempotency cache.
How do you keep chat messages in order?
Order messages within a conversation, not globally. The chat service assigns a monotonically increasing sequence number per conversation, typically by routing each conversation to a single owner partition or using an atomic counter. Clients sort by that sequence number and detect gaps when a number is missing, then fetch the missing range. Global ordering across all conversations is unnecessary and expensive, and interviewers expect you to say so.
How do read receipts work in a chat app?
Read receipts are small control messages that travel the same path as chat messages in reverse. When the recipient device stores a message it sends a delivered ack; when the user opens the conversation it sends a read ack containing the highest sequence number seen. The server updates per-recipient state and forwards the receipt to the sender. Sending a single high-water-mark sequence number instead of one receipt per message keeps the traffic small.
How does WhatsApp handle group messages with end-to-end encryption?
According to WhatsApp's security whitepaper, groups use the Sender Keys component of the Signal Protocol. Each member distributes a sender key to the other members over existing pairwise encrypted sessions. After that, the sender encrypts each group message once, uploads a single ciphertext, and the server fans it out to every member. The server routes and stores ciphertext but cannot read message contents.
How do you store billions of chat messages?
Use a wide-column or key-value store such as Cassandra or ScyllaDB, partitioned by conversation ID plus a time bucket, with the message sequence or time-sortable ID as the clustering key. That layout makes the most common query, fetching the latest messages in one conversation, a single-partition range scan. Discord's engineering blog describes this pattern for its own message storage, using channel ID and a time bucket as the partition key.
Ready to use AI assistance in your next interview?
TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.
Ace your next interview →