Notification system design is a classic interview question where you build a service that takes events from other systems and reliably delivers them to users over push, email, SMS and in-app channels. A strong answer is channel-agnostic: one pipeline handles routing, user preferences, templating, deduplication and retries, and thin per-channel workers talk to providers like APNs, FCM, an email service and an SMS gateway. Interviewers mostly grade how you handle failure, duplicates and spikes, not how many boxes you draw.
Key Takeaways
- Separate the core from the channels. A single notification service owns preferences, templates and idempotency; channel workers only adapt to provider APIs.
- Queues are the backbone. Use per-channel, per-priority queues so a marketing blast never delays a password reset code.
- At-least-once plus idempotency beats "exactly once." Derive a deterministic key from event, user and channel, and check it before every send.
- Classify errors before retrying. Retry 429s and 5xx with exponential backoff and jitter; never retry an unregistered token or a hard bounce.
- Protect the user, not just the system. Per-user frequency caps, quiet hours and digest batching matter as much as throughput.
- Close the loop with delivery tracking. Provider callbacks, a status state machine and per-channel success rates are what an interviewer expects in the monitoring section.
What is the notification system design question really testing?
The question tests whether you can design an asynchronous, multi-provider pipeline that stays correct when things fail. Unlike a URL shortener, most of the hard work happens after the API returns 202 Accepted. The interviewer wants to see queues, retries, idempotency and backpressure handled deliberately.
It also tests product judgment: good candidates bring up preferences, quiet hours and frequency caps without being asked. If you want a refresher on how these rounds are scored overall, the system design interview guide covers the general rubric.
Requirements and channels
Start by pinning down scope. Spend about five minutes here and write the answers on the board.
Functional requirements
- Other internal services can trigger a notification for one user, a list of users or a segment.
- Supported channels: mobile push (iOS and Android), email, SMS and an in-app inbox.
- Users can opt in or out per channel and per category (security, transactional, social, marketing).
- Messages are rendered from templates with variables and localization.
- Support both immediate sends and scheduled sends.
- Track status per message: queued, sent, delivered, opened, failed.
Non-functional requirements
- Reliability: no lost transactional notifications; duplicates kept to a minimum.
- Latency: high-priority messages such as one-time passcodes go out within seconds. Marketing can take minutes.
- Scalability: handle bursts many times the average rate during campaigns.
- Availability: the ingest API stays up even if a provider is down.
Channel constraints that shape the design
Each channel has hard limits that you should mention, because they drive templating and batching decisions.
| Channel | Provider examples | Key constraint | Failure signal to handle |
|---|---|---|---|
| iOS push | APNs | 4 KB payload limit | HTTP 410 for an unregistered device token |
| Android and web push | FCM | 4,096-byte payload for most messages, TTL up to 28 days | UNREGISTERED (404) means the token is dead |
| Amazon SES, SendGrid, Mailgun | Authentication and bulk-sender rules at Gmail and Yahoo | Hard bounces and spam complaints | |
| SMS | Twilio, Vonage, Amazon SNS | 160 GSM-7 or 70 UCS-2 characters per segment, cost per segment | Carrier errors, invalid numbers, opt-out keywords |
| In-app | Your own service | Must persist even if the user is offline | None from a third party; you own storage |
According to Firebase's FCM error code reference, UNREGISTERED means the app instance is no longer valid and the token should be discarded. On SMS, Twilio's documentation notes that a single emoji switches the whole message to UCS-2 and cuts the per-segment limit from 160 to 70 characters, which directly raises cost. For email, Google's sender guidelines require bulk senders (over 5,000 messages a day to Gmail) to authenticate with SPF, DKIM and DMARC, support one-click unsubscribe, and keep user-reported spam rates below 0.10%, never reaching 0.30%.
Back-of-envelope estimation
If the interviewer gives no numbers, propose some and confirm them. Assume 50 million daily active users and an average of 5 notifications per user per day.
Daily sends = 50M users * 5 = 250M notifications/day
Average rate = 250M / 86,400 s ~ 2,900/s
Peak (10x burst) ~ 29,000/s
Status events ~ 3-4 per message ~ 1B status rows/day
Average throughput is modest; bursts and status writes dominate the sizing.
High-level architecture with queues
The architecture is a pipeline: ingest, enrich, route, render, send, track. Each stage is connected by durable queues so a slow stage never blocks a fast one.
Producers (orders, auth, social, marketing)
|
v
[Notification API] --> validate, assign notification_id, return 202
|
v
[Ingest queue]
|
v
[Notification Service / Router]
- fan out to recipients
- load user preferences + contact info
- idempotency check
- pick channels, render templates
|
+--> [push.high] [push.low] --> Push workers --> APNs / FCM
+--> [email.high] [email.low] --> Email workers --> SES / SendGrid
+--> [sms.high] --> SMS workers --> Twilio / Vonage
+--> [inapp] --> Inbox writer --> Inbox DB + WebSocket
|
v
[Status store] <-- provider webhooks (delivered, bounced, opened)
[Retry scheduler] / [Dead-letter queues]
Why queues sit between every stage
Queues decouple producers from the slow, failure-prone provider calls. They absorb spikes, keep messages durable when APNs or an email provider has an outage, and let you scale push workers separately from SMS workers. Kafka fits when you want replay and high throughput; SQS or RabbitMQ are simpler when you want per-message acknowledgment and built-in dead-letter queues.
Split queues by channel and by priority. A marketing campaign to 20 million users should never sit in front of a login code.
Component responsibilities
- Notification API: authenticates the calling service, validates the payload, assigns an ID and enqueues. It does no provider work, so it stays fast and available.
- Router: resolves recipients (a segment may expand to millions of users, so do that fan-out in batches), checks preferences, applies idempotency and frequency caps, then publishes per-channel jobs.
- Channel workers: stateless consumers that format the provider request, call the provider, classify the response and update status.
- Contact store: device tokens, email addresses and phone numbers per user. A user may have several devices, so push fans out again here.
- Inbox service: stores in-app notifications and pushes them to connected clients over WebSockets. The real-time delivery piece overlaps heavily with a chat app design like WhatsApp.
Data model
CREATE TABLE notifications (
notification_id UUID PRIMARY KEY,
idempotency_key TEXT UNIQUE NOT NULL,
user_id BIGINT NOT NULL,
category TEXT NOT NULL,
channel TEXT NOT NULL,
template_id TEXT NOT NULL,
payload JSONB,
priority SMALLINT NOT NULL,
status TEXT NOT NULL,
attempts SMALLINT DEFAULT 0,
scheduled_at TIMESTAMPTZ,
created_at TIMESTAMPTZ NOT NULL,
updated_at TIMESTAMPTZ NOT NULL
);
CREATE TABLE user_preferences (
user_id BIGINT,
category TEXT,
channel TEXT,
enabled BOOLEAN NOT NULL,
quiet_start TIME,
quiet_end TIME,
timezone TEXT,
PRIMARY KEY (user_id, category, channel)
);
The notification log grows fast (roughly a billion status changes a day in the estimate above), so store it in a write-heavy store such as Cassandra or a partitioned table with a retention policy. Preferences are small and read-heavy, so cache them aggressively. A distributed cache in front of the preference store removes a database read from every send.
Notification design is the round where interviewers push hardest on follow-ups like "what if the provider times out after accepting the message?" If you blank on those under pressure, TechScreen runs invisibly during your Zoom, Google Meet or Teams screen share and surfaces trade-offs and failure modes in real time. Try it with 3 free tokens, no credit card.
Templates and user preferences
Templates and preferences are what make the core channel-agnostic. Producers send an event type and variables, never finished text. The router decides which channels apply and renders the right template for each one.
Templates
A template is keyed by (template_id, channel, locale, version). The same order_shipped event renders as a short push title, a full HTML email and a sub-160-character SMS. Versioning matters: if marketing edits a template mid-campaign, in-flight messages should render with the version they were enqueued with, so store the version on the notification row.
Validate templates at publish time, not send time. Check that every variable exists, the push payload stays under 4 KB, and the SMS body stays within GSM-7 if you want to avoid UCS-2 costs.
Preference evaluation order
The order of checks is something interviewers listen for. A clean version:
- Is the category mandatory? Security alerts and legal notices may bypass opt-outs (but still respect channel availability).
- Is the channel enabled for this user and category?
- Is the contact valid? Skip suppressed emails, dead tokens and numbers that replied STOP.
- Quiet hours: if inside the window in the user's timezone, delay non-urgent messages rather than drop them.
- Frequency cap: has the user exceeded N notifications in this category in the last hour or day?
- Channel fallback: if push fails or the user has no device, should this category fall back to email or SMS?
Compliance belongs here too. SMS opt-outs (STOP replies) and email one-click unsubscribe must flow back into the preference store quickly, since Google's guidelines expect unsubscribes to be honored within two days.
Retries, idempotency and deduplication
This is the deep dive most interviewers choose, and the one where answers separate. The core point: queues give at-least-once delivery, providers can time out after accepting a message, so duplicates are a design input, not an edge case.
How do you make sends idempotent?
Build a deterministic idempotency key, for example hash(source_event_id, user_id, channel, template_id). The producer supplies the event ID, so if it retries its API call, the key is identical.
def process(job):
key = f"notif:{job.event_id}:{job.user_id}:{job.channel}"
if not redis.set(key, "processing", nx=True, ex=86400):
return # already handled or in flight
try:
result = provider.send(job)
except TransientError:
redis.delete(key)
schedule_retry(job)
return
except PermanentError as e:
mark_failed(job, reason=e.code)
handle_permanent(job, e)
return
redis.set(key, "sent", ex=86400)
mark_sent(job, provider_id=result.id)
Two layers are worth naming. The UNIQUE constraint on idempotency_key in the notifications table catches duplicate API calls at ingest. The Redis set-if-absent check catches queue redeliveries at send time. Where providers accept an idempotency key of their own, pass it through as a third layer.
Be honest about the limit. If a provider call times out, you do not know whether the message went out. Retrying risks a duplicate; not retrying risks a loss. For a login code, a duplicate is harmless, so retry. For a marketing email, dropping is fine. That kind of per-category reasoning is exactly what interviewers look for.
Retry policy
| Error type | Examples | Action |
|---|---|---|
| Transient | Timeout, HTTP 429, 5xx, connection reset | Retry with exponential backoff and jitter, cap attempts |
| Throttled by provider | 429 with Retry-After | Honor the header, slow the whole worker pool for that provider |
| Permanent, recipient | FCM UNREGISTERED, APNs 410, hard bounce, invalid number | Do not retry; delete token or suppress address |
| Permanent, request | Payload too large, bad template | Do not retry; alert the owning team |
| Expired | Login code older than its validity window | Drop silently |
After the maximum attempts, move the job to a dead-letter queue with the last error attached, and alert on its depth.
Add a circuit breaker per provider. If error rates cross a threshold, stop hammering the provider, let queues buffer, and optionally fail over to a secondary SMS or email vendor. Multi-provider failover is a strong senior-level signal, the same pattern you would use in a payment system design.
Rate limiting and batching
Rate limiting runs in two directions: protecting providers and protecting users.
Outbound limits. APNs, FCM, email providers and SMS senders all enforce throughput limits, and email providers also watch your sending reputation. Workers should pull from a token bucket per provider and per sending identity. If you need a refresher on token bucket versus sliding window trade-offs, see the rate limiter design walkthrough.
Per-user limits. A user who gets 40 "someone liked your post" pushes in an hour will disable notifications entirely. Enforce caps per user and category with a counter in Redis keyed by user:category:window.
Batching and digests. Batching shows up in three forms:
- Collapsing: FCM's collapse key and the APNs collapse identifier let a newer message replace an older undelivered one, which suits "you have 12 unread messages" style updates.
- Digests: hold low-priority social events in a per-user buffer for a window, then send one summary ("Ana and 9 others liked your post"). This is a natural fit for feeds; it pairs well with the news feed design question.
- Provider batching: many email and push APIs accept multiple recipients per request, which cuts overhead during campaigns.
Campaign smoothing. For a 20-million-user blast, the router should expand the segment in pages and trickle jobs into the low-priority queue at a controlled rate, rather than dumping everything at once. This keeps provider limits intact and leaves headroom for transactional traffic.
Monitoring and delivery tracking
Treat each message as a state machine so every send is observable.
CREATED -> QUEUED -> SENT_TO_PROVIDER -> DELIVERED -> OPENED/CLICKED
\-> SUPPRESSED (preference, cap, quiet hours)
\-> FAILED_TRANSIENT -> retry -> ...
\-> FAILED_PERMANENT
\-> DEAD_LETTERED
Providers report outcomes asynchronously. Email services send webhooks for delivery, bounces and complaints; SMS providers send delivery receipts; push providers confirm acceptance, and opens are tracked by the client app. Ingest these webhooks through their own queue, since they arrive in bursts and can be retried by the provider (so they need idempotent handling too).
Metrics and alerts
- Queue depth and consumer lag per channel and priority
- End-to-end latency from API call to provider acceptance, at p50 and p99, split by priority
- Success, transient failure and permanent failure rates per provider
- DLQ depth and age of the oldest message
- Email bounce and complaint rates (Google asks senders to stay below 0.10% user-reported spam and never reach 0.30%)
- SMS cost per day and segments per message
- Suppression counts by reason, to catch a preference bug that silently blocks everything
The reliability checklist interviewers use
Use this checklist as a final pass. It maps to what senior interviewers typically probe by the end of the round, so walk through it before you say "I'm done."
| Check | What a strong answer includes | Common miss |
|---|---|---|
| Durability | Message persisted before the API returns 202 | Writing to an in-memory buffer |
| Priority isolation | Separate queues and worker pools for transactional vs marketing | One shared queue |
| Idempotency | Deterministic key, checked at ingest and at send | Relying on the queue for exactly-once |
| Error classification | Transient vs permanent handled differently | Retrying dead tokens forever |
| Backpressure | Circuit breakers, provider token buckets, campaign smoothing | Unlimited worker concurrency |
| User protection | Preferences, quiet hours, frequency caps, digests | Treating every event as a send |
| Compliance | Unsubscribe and STOP processing, suppression lists | Not mentioned |
| Contact hygiene | Token cleanup on 404/410, bounce suppression | Stale tokens inflating failure rates |
| Observability | Status state machine, webhooks, DLQ alerts, per-provider dashboards | "We'd add logging" |
| Failover | Secondary provider for SMS and email | Single vendor dependency |
Naming a concrete mechanism for most rows, not just the keyword, is what reads as senior. The backend engineer interview guide covers how the same reliability themes show up in other rounds.
Follow-up questions to expect
Interviewers use follow-ups to probe depth. Prepare a two- or three-sentence answer for each.
- "How do you send to 50 million users at once?" Expand the segment in pages, enqueue at a controlled rate into low-priority queues, and rely on provider batching. Precompute rendered content once per locale instead of once per user.
- "How would you guarantee ordering?" You generally don't across channels. Within a user and channel, partition the queue by user ID so messages land on the same consumer in order. Push providers do not guarantee ordering, so include a sequence number for the client.
- "What happens if the preference service is down?" Fail closed for marketing (don't send) and fail open for security-critical messages using cached preferences.
- "How do you handle time zones for scheduled sends?" Store
scheduled_atin UTC, compute it from the user's timezone at schedule time, and use a scheduler that scans a time-bucketed table or a delay queue. - "How would you add a new channel, like WhatsApp?" Add a channel adapter and a queue. If the core is channel-agnostic, nothing else changes. This is the payoff of the design.
If you want to rehearse these out loud before the real round, the guide on how to practice system design interviews has a structured plan, and the site reliability engineer interview guide goes deeper on alerting and incident questions that often follow this prompt.
The candidates who do best keep the core small and spend their time on failure handling. Draw the pipeline once, then let every question land on a specific box.
Want a second brain in the room for your next system design round? TechScreen stays invisible on screen shares across Zoom, Google Meet, Teams and CoderPad, and can suggest queue designs, retry policies and follow-up answers as the conversation moves. Start with 3 free tokens, no credit card required.
Frequently Asked Questions
What is a notification system in system design?
A notification system is a service that accepts events from other services and turns them into messages delivered to users over channels such as mobile push, email, SMS and in-app inboxes. It resolves who should be notified, checks their preferences, renders a template per channel, and hands the result to third-party providers like APNs, FCM, an email service or an SMS gateway. It also tracks delivery status and retries failures without sending duplicates.
Why use a message queue in a notification system?
A queue decouples the services that produce events from the slow, failure-prone work of calling external providers. It absorbs traffic spikes such as a marketing blast, lets each channel scale its workers independently, and keeps messages durable while a provider is down. Separate queues per channel and priority also stop a backlog of low-priority email from delaying an urgent one-time passcode sent over SMS.
How do you prevent duplicate notifications?
Give every notification a deterministic idempotency key, usually derived from the source event ID, the user ID and the channel. Before sending, a worker records that key in a store with a uniqueness constraint or an atomic set-if-absent operation in Redis with a TTL. If the key already exists, the worker skips the send. Queues typically guarantee at-least-once delivery, so this check is what makes the end-to-end behavior effectively once.
Can a notification system guarantee exactly-once delivery?
Not end to end. Once a message is handed to APNs, FCM, an email provider or an SMS carrier, you cannot control whether a timeout means it was delivered or lost. What you can build is at-least-once processing with idempotent sends, which removes most duplicates caused by your own retries and redeliveries. In an interview, say this explicitly. Claiming exactly-once delivery to a phone is a common red flag.
How should retries work for failed notifications?
Classify errors first. Transient failures such as timeouts, HTTP 429 or 5xx responses get retried with exponential backoff and jitter, up to a capped number of attempts, then go to a dead-letter queue. Permanent failures such as an invalid device token or a hard email bounce are never retried; instead the token is deleted or the address is suppressed. Time-sensitive messages like login codes should have a short expiry so stale retries are dropped.
How long should a notification system design answer take in an interview?
Most system design rounds run 45 to 60 minutes. A good split is about 5 minutes on requirements and scale, 10 minutes on the high-level architecture, 20 minutes on two or three deep dives such as idempotency, preferences or rate limiting, and the remaining time on monitoring, failure modes and follow-up questions. Let the interviewer steer which deep dives matter most.
What scale numbers should I assume for a notification system?
If the interviewer gives none, propose round numbers and confirm them. A common assumption is tens of millions of daily active users receiving a few notifications per day, which works out to hundreds of millions of sends daily and a few thousand per second on average, with peaks several times higher during campaigns. The point is to show that bursts, not averages, drive queue and worker sizing.
Ready to use AI assistance in your next interview?
TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.
Ace your next interview →