← All articles
13 min read

Design WhatsApp: Chat App System Design Interview Guide 2026

A complete walkthrough of the 'design WhatsApp' interview question, with sequence diagrams for every delivery state and a clear trade-off table for WebSockets, long polling, and SSE.

Chat app system design is the interview question where you design a messenger like WhatsApp: persistent connections for real-time delivery, durable per-conversation message storage, an offline inbox, group fan-out, presence, and sent, delivered, and read receipts. A strong answer centers on one idea. Messages are delivered at least once and deduplicated, ordered per conversation, and acknowledged at every hop so that each delivery state maps to a concrete event in the system.

Key Takeaways

  • Scope first. Agree on 1:1 chat, groups (cap the size, WhatsApp allows up to 1,024 members), presence, and receipts before drawing anything. Media, calls, and search are usually out of scope.
  • WebSockets win for chat because one full-duplex connection carries messages, acks, and receipts in both directions. Treat long polling as a fallback, not the design.
  • Order per conversation, never globally. A per-conversation sequence number gives clients a sort key and a way to detect gaps.
  • Each delivery state is a separate ack. Sent means the server persisted it, delivered means the recipient device stored it, and read means the user opened the chat.
  • Groups need a fan-out decision. Small groups fan out on write to per-user inboxes; very large channels are better served by fan-out on read.
  • E2EE changes the server's job. With Signal-style encryption, the server routes and stores ciphertext and never sees content, which rules out server-side search and moderation of message bodies.

What Requirements Should You Clarify First?

Your first five minutes decide the shape of the interview. If you need a refresher on pacing the full 45 minutes, the framework in how to ace a system design interview applies directly here.

Functional requirements

  1. 1:1 messaging between users, text first.
  2. Group chats with a size cap (ask; 256 to 1,024 is a reasonable range).
  3. Message delivery states: sent, delivered, read.
  4. Online presence and last-seen timestamp.
  5. Offline users receive everything they missed when they reconnect, on every device they own.

Non-functional requirements

  • Low latency: a message should reach an online recipient in well under a second.
  • Durability: once the sender sees "sent," the message must never be lost.
  • Per-conversation ordering. Global ordering is not required.
  • High availability over strict consistency for presence. A stale "online" dot is fine; a lost message is not.
  • End-to-end encryption, at least at a high level.

Out of scope unless the interviewer pushes: calls, search, stories, payments, and spam detection. Naming what you skip is a senior signal.

Back-of-the-envelope estimates

Use round numbers and state them as assumptions.

AssumptionValueImplication
Daily active users500 millionSize of the session registry
Peak concurrent connections~200 millionGateway fleet size
Connections per gateway server~500K to 1M (tuned)Roughly 200 to 400 gateway hosts at peak
Messages per user per day4020 billion messages per day
Average write rate20B / 86,400 s~230K messages per second, plan for 3x at peak
Average message size with metadata~200 bytes~4 TB of new message data per day before replication

Chat is a connection-holding problem first and a storage problem second.

High-Level Architecture

The core design has six components. Draw them in this order and explain the arrows as you go.

  1. Clients (phone, desktop, web) hold one persistent connection each.
  2. Load balancer terminates TLS and routes new connections to a gateway using least-connections, not round robin.
  3. Gateway (connection) service holds WebSocket connections. It is stateless apart from the sockets it owns.
  4. Session registry maps user_id + device_id to the gateway currently holding that device's socket. A Redis cluster with TTL heartbeats is a common choice.
  5. Chat service validates, assigns sequence numbers, persists messages, and routes them to recipient gateways.
  6. Storage: a message store (wide-column), an inbox or sync queue per device, and a relational store for users, groups, and memberships.

Add a push notification service (APNs and FCM) for devices with no live connection. The notification system design guide covers that component in depth, so you can treat it as a black box here.

Connection Layer: WebSockets vs Long Polling vs SSE

A WebSocket is a single TCP connection, upgraded from HTTP, that both sides can write to at any time. The protocol is defined in RFC 6455. For chat, that bidirectional channel is the deciding factor, but interviewers like to hear you compare the alternatives.

CriterionWebSocketsLong pollingServer-sent events (SSE)
DirectionFull duplexClient pulls, server holds request openServer to client only
Latency for new messageLowest, pushed immediatelyLow, but a new request is needed after every responseLow for server push
Client sending messagesSame socketSeparate HTTP requestSeparate HTTP request
Overhead per messageSmall frame headerFull HTTP headers per poll cycleSmall, but sends need full HTTP requests
Proxy and firewall friendlinessMostly fine today, occasional corporate proxy issuesWorks almost everywhereWorks over plain HTTP
Server stateLong-lived socket per deviceShort-lived, easier to load balanceLong-lived stream per device
Typing indicators and receiptsNatural fitClunky and chattyNeeds a second channel for upstream
Best use in a chat designPrimary transportFallback for hostile networksRead-only feeds, live notifications

Verdict to say in the interview: WebSockets as the primary transport, long polling as a degraded fallback, and SSE only if the product were read-heavy and one-directional. WhatsApp itself predates browser WebSockets on mobile and historically ran a custom protocol over persistent TCP connections. The principle is the same: one long-lived, bidirectional connection per device.

Scaling the gateway tier

Gateways are where most interview follow-ups land. Cover these four points:

  • Routing. When the chat service needs to deliver to user B, it looks up B's devices in the session registry, then forwards to those gateways over an internal RPC or pub/sub channel.
  • Heartbeats. Clients ping every 20 to 60 seconds. Missed heartbeats close the socket and expire the registry entry, which also feeds presence.
  • Deploys and failures. A gateway restart drops hundreds of thousands of sockets at once. Clients reconnect with jittered exponential backoff so the fleet does not get hit by a thundering herd. Rate-limit reconnects at the load balancer; the rate limiter design guide explains the token bucket you would use.
  • Memory per connection. Memory and file descriptors run out before CPU does, which is why lightweight concurrency models (Erlang processes, goroutines, event loops) suit gateways.

How Are Messages Stored and Ordered?

Message storage is a write-heavy, append-mostly workload whose dominant read is "give me the latest N messages in this conversation." That access pattern points to a wide-column store such as Cassandra or ScyllaDB, partitioned by conversation.

CREATE TABLE messages (
  conversation_id  bigint,
  bucket           int,
  seq              bigint,
  message_id       bigint,
  sender_id        bigint,
  client_msg_id    uuid,
  body             blob,
  created_at       timestamp,
  PRIMARY KEY ((conversation_id, bucket), seq)
) WITH CLUSTERING ORDER BY (seq DESC);

The bucket (for example, a time window or a block of sequence numbers) caps partition size so a very active group does not grow one partition forever. Discord's engineering team describes almost exactly this layout, with channel ID plus a static time bucket as the partition key, in How Discord Stores Trillions of Messages.

Sequence numbers, not timestamps

Ordering is per conversation. Wall-clock timestamps from different servers drift, so two messages sent milliseconds apart can sort incorrectly. Use a per-conversation sequence number instead:

  • Route every write for a conversation to one owner (a partition of the chat service chosen by hashing conversation_id), which increments a counter. This is the simplest correct answer.
  • Alternatively, use an atomic INCR on a counter in a strongly consistent store. This is simpler to operate but becomes a hot key for busy groups.
  • Keep a separate globally unique, time-sortable message_id (Snowflake-style: timestamp, worker ID, sequence) for references, replies, and deduplication across systems.

Clients sort by seq. If a client holds seq 41 and receives 43, it knows 42 is missing and fetches the gap. That gap detection is what makes the sync protocol self-healing.

Message Delivery States, End to End

This is the section most candidates hand-wave. Each checkmark in the UI corresponds to a specific durable event, and you should be able to draw all three.

Sent (one check): the server has durably stored the message and acknowledged the sender.

Sender A        Gateway A        Chat Service        Message Store      Inbox(B)
   |  SEND(cmid=u1, conv=7, body)    |                     |                |
   |--------------->|                |                     |                |
   |                |--------------->|                     |                |
   |                |                |-- dedupe on cmid -->|                |
   |                |                |-- assign seq=42 --->|  write msg     |
   |                |                |---------------------------------->|  append(42)
   |                |<-- ACK(cmid=u1, seq=42, msg_id) -----|                |
   |<---------------|                |                     |                |
   |  UI: one check (SENT)           |                     |                |

Delivered (two checks): the recipient's device has received and stored the message locally.

Chat Service     Session Registry     Gateway B        Recipient B
   |-- where is B? -->|                    |                  |
   |<-- gw-B-17 ------|                    |                  |
   |-- PUSH(conv=7, seq=42) -------------->|                  |
   |                  |                    |----------------->|
   |                  |                    |                  | store locally
   |                  |                    |<-- DELIVERED(conv=7, seq=42)
   |<----------------------------------------|                |
   | mark delivered for B; remove 42 from Inbox(B)             |
   |-- RECEIPT(delivered, seq<=42) --> Gateway A --> Sender A  |
   |                          UI: two checks (DELIVERED)       |

If B is offline, the registry has no entry. The message stays in B's inbox, the chat service triggers a push notification, and the delivered receipt waits until B reconnects and syncs.

Read (blue checks): the user opened the conversation and the message was on screen.

Recipient B       Gateway B        Chat Service        Gateway A       Sender A
   | opens conv 7     |                 |                   |               |
   |-- READ(conv=7, up_to_seq=42) ----->|                   |               |
   |                  |---------------->| update read_seq(B,7)=42          |
   |                  |                 |-- RECEIPT(read, seq<=42) ------->|
   |                  |                 |                   |-------------->|
   |                  |                 |                   |  UI: blue checks

Two design choices make this efficient:

  • High-water marks. Receipts carry up_to_seq, not individual message IDs. One read receipt covers every earlier message in the conversation.
  • Receipts are messages too. They ride the same transport and the same at-least-once machinery, but they are idempotent (setting read_seq = max(read_seq, 42)), so duplicates are harmless.
StateTrigger eventWho acksStored where
Pending (clock icon)User taps sendNobody yetSender's local outbox
SentServer persisted the messageChat service to senderMessage store and recipient inbox
DeliveredRecipient device stored itRecipient device to serverPer-recipient delivery cursor
ReadRecipient opened the chatRecipient device to serverPer-recipient read cursor

For groups, "delivered" and "read" become per-member cursors, and the sender's UI shows the aggregate (for example, all members have read it).

Can you draw these three sequence diagrams from memory while an interviewer asks follow-ups about duplicate retries? TechScreen is an invisible AI interview assistant that stays hidden during Zoom, Google Meet, and Teams screen shares and gives you real-time prompts for system design rounds like this one. Start with 3 free tokens, no credit card required.

Get started free →

Delivery Guarantees and Offline Sync

No distributed chat system delivers exactly once over an unreliable network. What you build is at-least-once delivery plus idempotent processing, which the user experiences as exactly once.

On the sending side:

  • The client generates client_msg_id (a UUID) and stores the message in a local outbox.
  • It retries with backoff until it gets an ack carrying that ID.
  • The server checks client_msg_id against a short-lived dedupe cache (or a unique index) before assigning a sequence number. A retry of an already-stored message returns the original ack.

On the receiving side:

  • Every message is written to the recipient's per-device inbox before the push attempt.
  • The inbox entry is removed only after the device acks. If the ack is lost, the server pushes again, and the client drops the duplicate by seq.

Offline sync on reconnect:

client -> SYNC { device_id, cursors: { conv_7: 41, conv_9: 1203 } }
server -> for each conversation with seq > cursor:
            return messages in pages of 100, oldest first
client -> ACK highest seq per conversation
server -> trim inbox entries <= acked seq

Mention multi-device explicitly: each device keeps its own cursor and inbox entry, so a phone offline for a week catches up without affecting the desktop client. Set a retention limit on undelivered items and say what happens when it expires.

Most reads hit the last few dozen messages, so cache recent history; the distributed cache design guide covers eviction and invalidation.

Group Fan-Out

Group messaging is where the design branches. The question is whether the server copies the message reference into every member's inbox at write time or members pull from the shared conversation at read time.

ApproachHow it worksStrengthWeaknessUse for
Fan-out on writeStore once, append a pointer to each member's inbox, push to each online memberFast reads, simple syncWrite cost grows with group sizeGroups up to ~1,000 members
Fan-out on readStore once in the conversation; members fetch since their cursorConstant write costReads are heavier, harder to pushBroadcast channels with very large audiences
HybridPush to online members, others pull on reconnectBalances bothMore code pathsLarge groups with many inactive members

For a WhatsApp-style product with capped group sizes, fan-out on write is the right default. With the 1,024-member cap, one message creates at most about a thousand inbox appends, which a queue (Kafka or similar) absorbs easily. If you have seen the celebrity problem in news feed system design, point out that the same trade-off appears here for broadcast channels.

Keep group membership in a relational store and cache the member list. Membership changes (adds, removes, admin changes) are themselves ordered events in the conversation, so a removed member stops receiving messages from a specific sequence number onward.

Presence and Last-Seen

Presence is a high-write, low-value signal. Design it to be cheap and eventually consistent.

  • Source of truth: the gateway heartbeat. A live socket with recent heartbeats means online. Store user_id -> {status, last_seen} in an in-memory store with a TTL slightly longer than the heartbeat interval.
  • Do not broadcast every change to every contact. A user with 500 contacts going online and offline repeatedly would generate enormous traffic. Instead, push presence only to users who currently have a conversation with that person open, and let others fetch it lazily.
  • Debounce flapping. Mobile networks drop constantly. Wait a few seconds before marking someone offline so a brief reconnect does not produce an offline-online blip.
  • Respect privacy settings. Last-seen visibility (everyone, contacts, nobody) is checked at read time, not stored per viewer.
  • Typing indicators are ephemeral: send them over the socket, never persist them, and drop them if the recipient is offline.

End-to-End Encryption at a High Level

End-to-end encryption means only the sender and recipient devices hold the keys needed to decrypt a message; the server relays ciphertext it cannot read. WhatsApp uses the Signal Protocol, and its security whitepaper describes the design. In an interview, three or four sentences are enough:

  1. Each device publishes public identity keys and a batch of one-time prekeys to a key server. A sender fetches the recipient's prekey bundle to start a session without the recipient being online.
  2. The Double Ratchet algorithm derives a new key for every message, which gives forward secrecy.
  3. For groups, the whitepaper describes Sender Keys: each member shares a sender key over pairwise sessions, then encrypts each group message once, and the server fans out a single ciphertext.
  4. Multi-device means each device has its own keys, so the sender encrypts for every device of every recipient.

Then state the consequences: the server cannot index bodies for search, cannot moderate content, and cannot restore history unless the client uploads an encrypted backup.

How Do Interviewers Grade This Question?

Most interviewers score the chat question on a handful of signals. This rubric reflects common patterns in system design loops at companies such as Meta (see the Meta interview process guide), not any one company's official scorecard.

SignalMid-level barSenior bar
RequirementsLists core featuresCaps group size, names non-goals, sets durability vs presence consistency
Connection layerPicks WebSocketsJustifies vs long polling and SSE, handles reconnect storms and gateway routing
StoragePicks a NoSQL storePartition key with buckets, per-conversation sequence numbers, gap detection
DeliveryMentions acksWalks sent, delivered, read end to end with idempotent retries and per-device cursors
GroupsMentions fan-outChooses write vs read fan-out by group size, handles membership events
Depth on demandAnswers follow-upsProactively flags hot partitions, multi-region, and E2EE consequences

Senior and staff candidates are expected to drive the trade-off discussion without prompting. The staff engineer interview guide explains how that expectation changes the conversation.

Common Follow-Up Questions

Prepare one or two sentences for each of these:

  • "What if a gateway dies mid-delivery?" The message is already in the inbox, the device reconnects to a new gateway, syncs from its cursor, and receives it. Nothing is lost because persistence happens before push.
  • "How do you handle a hot group?" Shard the sequence counter owner, batch inbox appends, and move very large audiences to fan-out on read.
  • "Multi-region?" Home each conversation in one region for ordering; route cross-region messages over a replicated queue. Users connect to the nearest gateway.
  • "How would you add media?" Upload media directly to object storage with a pre-signed URL, then send a normal message containing the media reference and (with E2EE) the decryption key.

How to Practice This Design

Practice out loud against a timer. Draw the architecture in under five minutes, then spend most of your time on delivery states and fan-out, where follow-ups go. The drills in how to practice system design interviews work well here, and the backend engineer interview guide shows where this question fits in a typical loop.

A good final check: can you explain, without notes, why a message is never lost if a gateway crashes after the sender sees one check? If yes, you understand the design.

System design rounds reward candidates who can recall the right trade-off table under pressure. TechScreen runs invisibly during your screen share and gives real-time structure for questions like "design WhatsApp," from capacity estimates to delivery guarantees. Try it on a mock interview with 3 free tokens.

Get started free →

Frequently Asked Questions

How do you design a chat app like WhatsApp in a system design interview?

Start by scoping requirements: 1:1 chat, group chat, online presence, and sent, delivered, and read receipts. Then estimate scale, and propose persistent WebSocket connections through a stateless gateway tier, a session registry that maps users to gateways, a chat service that assigns per-conversation sequence numbers, a wide-column message store partitioned by conversation, and a per-user inbox for offline delivery. Finish with group fan-out, presence, and a short end-to-end encryption overview.

Should a chat app use WebSockets, long polling, or server-sent events?

WebSockets are the default for chat because the connection is full-duplex: the client sends messages and acks on the same socket that the server uses to push new messages. Long polling works through any proxy but adds latency and reconnect overhead on every message. Server-sent events are simple for server-to-client streams but only flow one way, so the client needs separate HTTP calls to send. Mention long polling as a fallback for restrictive networks.

How does a chat system guarantee message delivery?

Chat systems provide at-least-once delivery plus deduplication, which looks exactly-once to the user. The client attaches a unique client message ID and retries until the server acks. The server writes the message durably before acknowledging, then pushes it to the recipient and keeps it in the recipient's inbox until the recipient device acks. Duplicates from retries are dropped by checking the client message ID, usually with a unique constraint or idempotency cache.

How do you keep chat messages in order?

Order messages within a conversation, not globally. The chat service assigns a monotonically increasing sequence number per conversation, typically by routing each conversation to a single owner partition or using an atomic counter. Clients sort by that sequence number and detect gaps when a number is missing, then fetch the missing range. Global ordering across all conversations is unnecessary and expensive, and interviewers expect you to say so.

How do read receipts work in a chat app?

Read receipts are small control messages that travel the same path as chat messages in reverse. When the recipient device stores a message it sends a delivered ack; when the user opens the conversation it sends a read ack containing the highest sequence number seen. The server updates per-recipient state and forwards the receipt to the sender. Sending a single high-water-mark sequence number instead of one receipt per message keeps the traffic small.

How does WhatsApp handle group messages with end-to-end encryption?

According to WhatsApp's security whitepaper, groups use the Sender Keys component of the Signal Protocol. Each member distributes a sender key to the other members over existing pairwise encrypted sessions. After that, the sender encrypts each group message once, uploads a single ciphertext, and the server fans it out to every member. The server routes and stores ciphertext but cannot read message contents.

How do you store billions of chat messages?

Use a wide-column or key-value store such as Cassandra or ScyllaDB, partitioned by conversation ID plus a time bucket, with the message sequence or time-sortable ID as the clustering key. That layout makes the most common query, fetching the latest messages in one conversation, a single-partition range scan. Discord's engineering blog describes this pattern for its own message storage, using channel ID and a time bucket as the partition key.

Ready to use AI assistance in your next interview?

TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.

Ace your next interview →