← All articles
13 min read

Notification System Design Interview: Push, Email, SMS 2026

A notification system design answer that scores well is channel-agnostic, queue-backed and idempotent. Here is the full architecture, the trade-offs, and the reliability checklist interviewers grade against.

Notification system design is a classic interview question where you build a service that takes events from other systems and reliably delivers them to users over push, email, SMS and in-app channels. A strong answer is channel-agnostic: one pipeline handles routing, user preferences, templating, deduplication and retries, and thin per-channel workers talk to providers like APNs, FCM, an email service and an SMS gateway. Interviewers mostly grade how you handle failure, duplicates and spikes, not how many boxes you draw.

Key Takeaways

  • Separate the core from the channels. A single notification service owns preferences, templates and idempotency; channel workers only adapt to provider APIs.
  • Queues are the backbone. Use per-channel, per-priority queues so a marketing blast never delays a password reset code.
  • At-least-once plus idempotency beats "exactly once." Derive a deterministic key from event, user and channel, and check it before every send.
  • Classify errors before retrying. Retry 429s and 5xx with exponential backoff and jitter; never retry an unregistered token or a hard bounce.
  • Protect the user, not just the system. Per-user frequency caps, quiet hours and digest batching matter as much as throughput.
  • Close the loop with delivery tracking. Provider callbacks, a status state machine and per-channel success rates are what an interviewer expects in the monitoring section.

What is the notification system design question really testing?

The question tests whether you can design an asynchronous, multi-provider pipeline that stays correct when things fail. Unlike a URL shortener, most of the hard work happens after the API returns 202 Accepted. The interviewer wants to see queues, retries, idempotency and backpressure handled deliberately.

It also tests product judgment: good candidates bring up preferences, quiet hours and frequency caps without being asked. If you want a refresher on how these rounds are scored overall, the system design interview guide covers the general rubric.

Requirements and channels

Start by pinning down scope. Spend about five minutes here and write the answers on the board.

Functional requirements

  • Other internal services can trigger a notification for one user, a list of users or a segment.
  • Supported channels: mobile push (iOS and Android), email, SMS and an in-app inbox.
  • Users can opt in or out per channel and per category (security, transactional, social, marketing).
  • Messages are rendered from templates with variables and localization.
  • Support both immediate sends and scheduled sends.
  • Track status per message: queued, sent, delivered, opened, failed.

Non-functional requirements

  • Reliability: no lost transactional notifications; duplicates kept to a minimum.
  • Latency: high-priority messages such as one-time passcodes go out within seconds. Marketing can take minutes.
  • Scalability: handle bursts many times the average rate during campaigns.
  • Availability: the ingest API stays up even if a provider is down.

Channel constraints that shape the design

Each channel has hard limits that you should mention, because they drive templating and batching decisions.

ChannelProvider examplesKey constraintFailure signal to handle
iOS pushAPNs4 KB payload limitHTTP 410 for an unregistered device token
Android and web pushFCM4,096-byte payload for most messages, TTL up to 28 daysUNREGISTERED (404) means the token is dead
EmailAmazon SES, SendGrid, MailgunAuthentication and bulk-sender rules at Gmail and YahooHard bounces and spam complaints
SMSTwilio, Vonage, Amazon SNS160 GSM-7 or 70 UCS-2 characters per segment, cost per segmentCarrier errors, invalid numbers, opt-out keywords
In-appYour own serviceMust persist even if the user is offlineNone from a third party; you own storage

According to Firebase's FCM error code reference, UNREGISTERED means the app instance is no longer valid and the token should be discarded. On SMS, Twilio's documentation notes that a single emoji switches the whole message to UCS-2 and cuts the per-segment limit from 160 to 70 characters, which directly raises cost. For email, Google's sender guidelines require bulk senders (over 5,000 messages a day to Gmail) to authenticate with SPF, DKIM and DMARC, support one-click unsubscribe, and keep user-reported spam rates below 0.10%, never reaching 0.30%.

Back-of-envelope estimation

If the interviewer gives no numbers, propose some and confirm them. Assume 50 million daily active users and an average of 5 notifications per user per day.

Daily sends      = 50M users * 5      = 250M notifications/day
Average rate     = 250M / 86,400 s    ~ 2,900/s
Peak (10x burst) ~ 29,000/s
Status events    ~ 3-4 per message    ~ 1B status rows/day

Average throughput is modest; bursts and status writes dominate the sizing.

High-level architecture with queues

The architecture is a pipeline: ingest, enrich, route, render, send, track. Each stage is connected by durable queues so a slow stage never blocks a fast one.

Producers (orders, auth, social, marketing)
        |
        v
[Notification API] --> validate, assign notification_id, return 202
        |
        v
  [Ingest queue]
        |
        v
[Notification Service / Router]
   - fan out to recipients
   - load user preferences + contact info
   - idempotency check
   - pick channels, render templates
        |
        +--> [push.high]  [push.low]  --> Push workers  --> APNs / FCM
        +--> [email.high] [email.low] --> Email workers --> SES / SendGrid
        +--> [sms.high]               --> SMS workers   --> Twilio / Vonage
        +--> [inapp]                  --> Inbox writer  --> Inbox DB + WebSocket
        |
        v
[Status store] <-- provider webhooks (delivered, bounced, opened)
[Retry scheduler] / [Dead-letter queues]

Why queues sit between every stage

Queues decouple producers from the slow, failure-prone provider calls. They absorb spikes, keep messages durable when APNs or an email provider has an outage, and let you scale push workers separately from SMS workers. Kafka fits when you want replay and high throughput; SQS or RabbitMQ are simpler when you want per-message acknowledgment and built-in dead-letter queues.

Split queues by channel and by priority. A marketing campaign to 20 million users should never sit in front of a login code.

Component responsibilities

  • Notification API: authenticates the calling service, validates the payload, assigns an ID and enqueues. It does no provider work, so it stays fast and available.
  • Router: resolves recipients (a segment may expand to millions of users, so do that fan-out in batches), checks preferences, applies idempotency and frequency caps, then publishes per-channel jobs.
  • Channel workers: stateless consumers that format the provider request, call the provider, classify the response and update status.
  • Contact store: device tokens, email addresses and phone numbers per user. A user may have several devices, so push fans out again here.
  • Inbox service: stores in-app notifications and pushes them to connected clients over WebSockets. The real-time delivery piece overlaps heavily with a chat app design like WhatsApp.

Data model

CREATE TABLE notifications (
  notification_id   UUID PRIMARY KEY,
  idempotency_key   TEXT UNIQUE NOT NULL,
  user_id           BIGINT NOT NULL,
  category          TEXT NOT NULL,
  channel           TEXT NOT NULL,
  template_id       TEXT NOT NULL,
  payload           JSONB,
  priority          SMALLINT NOT NULL,
  status            TEXT NOT NULL,
  attempts          SMALLINT DEFAULT 0,
  scheduled_at      TIMESTAMPTZ,
  created_at        TIMESTAMPTZ NOT NULL,
  updated_at        TIMESTAMPTZ NOT NULL
);

CREATE TABLE user_preferences (
  user_id      BIGINT,
  category     TEXT,
  channel      TEXT,
  enabled      BOOLEAN NOT NULL,
  quiet_start  TIME,
  quiet_end    TIME,
  timezone     TEXT,
  PRIMARY KEY (user_id, category, channel)
);

The notification log grows fast (roughly a billion status changes a day in the estimate above), so store it in a write-heavy store such as Cassandra or a partitioned table with a retention policy. Preferences are small and read-heavy, so cache them aggressively. A distributed cache in front of the preference store removes a database read from every send.

Notification design is the round where interviewers push hardest on follow-ups like "what if the provider times out after accepting the message?" If you blank on those under pressure, TechScreen runs invisibly during your Zoom, Google Meet or Teams screen share and surfaces trade-offs and failure modes in real time. Try it with 3 free tokens, no credit card.

Get started free →

Templates and user preferences

Templates and preferences are what make the core channel-agnostic. Producers send an event type and variables, never finished text. The router decides which channels apply and renders the right template for each one.

Templates

A template is keyed by (template_id, channel, locale, version). The same order_shipped event renders as a short push title, a full HTML email and a sub-160-character SMS. Versioning matters: if marketing edits a template mid-campaign, in-flight messages should render with the version they were enqueued with, so store the version on the notification row.

Validate templates at publish time, not send time. Check that every variable exists, the push payload stays under 4 KB, and the SMS body stays within GSM-7 if you want to avoid UCS-2 costs.

Preference evaluation order

The order of checks is something interviewers listen for. A clean version:

  1. Is the category mandatory? Security alerts and legal notices may bypass opt-outs (but still respect channel availability).
  2. Is the channel enabled for this user and category?
  3. Is the contact valid? Skip suppressed emails, dead tokens and numbers that replied STOP.
  4. Quiet hours: if inside the window in the user's timezone, delay non-urgent messages rather than drop them.
  5. Frequency cap: has the user exceeded N notifications in this category in the last hour or day?
  6. Channel fallback: if push fails or the user has no device, should this category fall back to email or SMS?

Compliance belongs here too. SMS opt-outs (STOP replies) and email one-click unsubscribe must flow back into the preference store quickly, since Google's guidelines expect unsubscribes to be honored within two days.

Retries, idempotency and deduplication

This is the deep dive most interviewers choose, and the one where answers separate. The core point: queues give at-least-once delivery, providers can time out after accepting a message, so duplicates are a design input, not an edge case.

How do you make sends idempotent?

Build a deterministic idempotency key, for example hash(source_event_id, user_id, channel, template_id). The producer supplies the event ID, so if it retries its API call, the key is identical.

def process(job):
    key = f"notif:{job.event_id}:{job.user_id}:{job.channel}"
    if not redis.set(key, "processing", nx=True, ex=86400):
        return  # already handled or in flight

    try:
        result = provider.send(job)
    except TransientError:
        redis.delete(key)
        schedule_retry(job)
        return
    except PermanentError as e:
        mark_failed(job, reason=e.code)
        handle_permanent(job, e)
        return

    redis.set(key, "sent", ex=86400)
    mark_sent(job, provider_id=result.id)

Two layers are worth naming. The UNIQUE constraint on idempotency_key in the notifications table catches duplicate API calls at ingest. The Redis set-if-absent check catches queue redeliveries at send time. Where providers accept an idempotency key of their own, pass it through as a third layer.

Be honest about the limit. If a provider call times out, you do not know whether the message went out. Retrying risks a duplicate; not retrying risks a loss. For a login code, a duplicate is harmless, so retry. For a marketing email, dropping is fine. That kind of per-category reasoning is exactly what interviewers look for.

Retry policy

Error typeExamplesAction
TransientTimeout, HTTP 429, 5xx, connection resetRetry with exponential backoff and jitter, cap attempts
Throttled by provider429 with Retry-AfterHonor the header, slow the whole worker pool for that provider
Permanent, recipientFCM UNREGISTERED, APNs 410, hard bounce, invalid numberDo not retry; delete token or suppress address
Permanent, requestPayload too large, bad templateDo not retry; alert the owning team
ExpiredLogin code older than its validity windowDrop silently

After the maximum attempts, move the job to a dead-letter queue with the last error attached, and alert on its depth.

Add a circuit breaker per provider. If error rates cross a threshold, stop hammering the provider, let queues buffer, and optionally fail over to a secondary SMS or email vendor. Multi-provider failover is a strong senior-level signal, the same pattern you would use in a payment system design.

Rate limiting and batching

Rate limiting runs in two directions: protecting providers and protecting users.

Outbound limits. APNs, FCM, email providers and SMS senders all enforce throughput limits, and email providers also watch your sending reputation. Workers should pull from a token bucket per provider and per sending identity. If you need a refresher on token bucket versus sliding window trade-offs, see the rate limiter design walkthrough.

Per-user limits. A user who gets 40 "someone liked your post" pushes in an hour will disable notifications entirely. Enforce caps per user and category with a counter in Redis keyed by user:category:window.

Batching and digests. Batching shows up in three forms:

  • Collapsing: FCM's collapse key and the APNs collapse identifier let a newer message replace an older undelivered one, which suits "you have 12 unread messages" style updates.
  • Digests: hold low-priority social events in a per-user buffer for a window, then send one summary ("Ana and 9 others liked your post"). This is a natural fit for feeds; it pairs well with the news feed design question.
  • Provider batching: many email and push APIs accept multiple recipients per request, which cuts overhead during campaigns.

Campaign smoothing. For a 20-million-user blast, the router should expand the segment in pages and trickle jobs into the low-priority queue at a controlled rate, rather than dumping everything at once. This keeps provider limits intact and leaves headroom for transactional traffic.

Monitoring and delivery tracking

Treat each message as a state machine so every send is observable.

CREATED -> QUEUED -> SENT_TO_PROVIDER -> DELIVERED -> OPENED/CLICKED
                \-> SUPPRESSED (preference, cap, quiet hours)
                \-> FAILED_TRANSIENT -> retry -> ...
                \-> FAILED_PERMANENT
                \-> DEAD_LETTERED

Providers report outcomes asynchronously. Email services send webhooks for delivery, bounces and complaints; SMS providers send delivery receipts; push providers confirm acceptance, and opens are tracked by the client app. Ingest these webhooks through their own queue, since they arrive in bursts and can be retried by the provider (so they need idempotent handling too).

Metrics and alerts

  • Queue depth and consumer lag per channel and priority
  • End-to-end latency from API call to provider acceptance, at p50 and p99, split by priority
  • Success, transient failure and permanent failure rates per provider
  • DLQ depth and age of the oldest message
  • Email bounce and complaint rates (Google asks senders to stay below 0.10% user-reported spam and never reach 0.30%)
  • SMS cost per day and segments per message
  • Suppression counts by reason, to catch a preference bug that silently blocks everything

The reliability checklist interviewers use

Use this checklist as a final pass. It maps to what senior interviewers typically probe by the end of the round, so walk through it before you say "I'm done."

CheckWhat a strong answer includesCommon miss
DurabilityMessage persisted before the API returns 202Writing to an in-memory buffer
Priority isolationSeparate queues and worker pools for transactional vs marketingOne shared queue
IdempotencyDeterministic key, checked at ingest and at sendRelying on the queue for exactly-once
Error classificationTransient vs permanent handled differentlyRetrying dead tokens forever
BackpressureCircuit breakers, provider token buckets, campaign smoothingUnlimited worker concurrency
User protectionPreferences, quiet hours, frequency caps, digestsTreating every event as a send
ComplianceUnsubscribe and STOP processing, suppression listsNot mentioned
Contact hygieneToken cleanup on 404/410, bounce suppressionStale tokens inflating failure rates
ObservabilityStatus state machine, webhooks, DLQ alerts, per-provider dashboards"We'd add logging"
FailoverSecondary provider for SMS and emailSingle vendor dependency

Naming a concrete mechanism for most rows, not just the keyword, is what reads as senior. The backend engineer interview guide covers how the same reliability themes show up in other rounds.

Follow-up questions to expect

Interviewers use follow-ups to probe depth. Prepare a two- or three-sentence answer for each.

  • "How do you send to 50 million users at once?" Expand the segment in pages, enqueue at a controlled rate into low-priority queues, and rely on provider batching. Precompute rendered content once per locale instead of once per user.
  • "How would you guarantee ordering?" You generally don't across channels. Within a user and channel, partition the queue by user ID so messages land on the same consumer in order. Push providers do not guarantee ordering, so include a sequence number for the client.
  • "What happens if the preference service is down?" Fail closed for marketing (don't send) and fail open for security-critical messages using cached preferences.
  • "How do you handle time zones for scheduled sends?" Store scheduled_at in UTC, compute it from the user's timezone at schedule time, and use a scheduler that scans a time-bucketed table or a delay queue.
  • "How would you add a new channel, like WhatsApp?" Add a channel adapter and a queue. If the core is channel-agnostic, nothing else changes. This is the payoff of the design.

If you want to rehearse these out loud before the real round, the guide on how to practice system design interviews has a structured plan, and the site reliability engineer interview guide goes deeper on alerting and incident questions that often follow this prompt.

The candidates who do best keep the core small and spend their time on failure handling. Draw the pipeline once, then let every question land on a specific box.

Want a second brain in the room for your next system design round? TechScreen stays invisible on screen shares across Zoom, Google Meet, Teams and CoderPad, and can suggest queue designs, retry policies and follow-up answers as the conversation moves. Start with 3 free tokens, no credit card required.

Get started free →

Frequently Asked Questions

What is a notification system in system design?

A notification system is a service that accepts events from other services and turns them into messages delivered to users over channels such as mobile push, email, SMS and in-app inboxes. It resolves who should be notified, checks their preferences, renders a template per channel, and hands the result to third-party providers like APNs, FCM, an email service or an SMS gateway. It also tracks delivery status and retries failures without sending duplicates.

Why use a message queue in a notification system?

A queue decouples the services that produce events from the slow, failure-prone work of calling external providers. It absorbs traffic spikes such as a marketing blast, lets each channel scale its workers independently, and keeps messages durable while a provider is down. Separate queues per channel and priority also stop a backlog of low-priority email from delaying an urgent one-time passcode sent over SMS.

How do you prevent duplicate notifications?

Give every notification a deterministic idempotency key, usually derived from the source event ID, the user ID and the channel. Before sending, a worker records that key in a store with a uniqueness constraint or an atomic set-if-absent operation in Redis with a TTL. If the key already exists, the worker skips the send. Queues typically guarantee at-least-once delivery, so this check is what makes the end-to-end behavior effectively once.

Can a notification system guarantee exactly-once delivery?

Not end to end. Once a message is handed to APNs, FCM, an email provider or an SMS carrier, you cannot control whether a timeout means it was delivered or lost. What you can build is at-least-once processing with idempotent sends, which removes most duplicates caused by your own retries and redeliveries. In an interview, say this explicitly. Claiming exactly-once delivery to a phone is a common red flag.

How should retries work for failed notifications?

Classify errors first. Transient failures such as timeouts, HTTP 429 or 5xx responses get retried with exponential backoff and jitter, up to a capped number of attempts, then go to a dead-letter queue. Permanent failures such as an invalid device token or a hard email bounce are never retried; instead the token is deleted or the address is suppressed. Time-sensitive messages like login codes should have a short expiry so stale retries are dropped.

How long should a notification system design answer take in an interview?

Most system design rounds run 45 to 60 minutes. A good split is about 5 minutes on requirements and scale, 10 minutes on the high-level architecture, 20 minutes on two or three deep dives such as idempotency, preferences or rate limiting, and the remaining time on monitoring, failure modes and follow-up questions. Let the interviewer steer which deep dives matter most.

What scale numbers should I assume for a notification system?

If the interviewer gives none, propose round numbers and confirm them. A common assumption is tens of millions of daily active users receiving a few notifications per day, which works out to hundreds of millions of sends daily and a few thousand per second on average, with peaks several times higher during campaigns. The point is to show that bursts, not averages, drive queue and worker sizing.

Ready to use AI assistance in your next interview?

TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.

Ace your next interview →