← All articles
13 min read

Design a Rate Limiter: System Design Interview Guide (2026)

A rate limiter is one of the most common system design interview questions, and one of the easiest to answer shallowly. This guide gives you the algorithm decision table, working code, and a distributed Redis design that holds up to senior-level follow-ups.

Rate limiter system design is the process of building a component that caps how many requests a client can make in a time window, rejecting the excess with HTTP 429. In an interview, a strong answer picks an algorithm (usually token bucket or sliding window counter), places the limiter at the API gateway, and stores counters in Redis using atomic Lua scripts so many servers share one accurate limit. The questions that separate levels are about distributed race conditions, failure modes, and multi-tier limits, not the algorithm itself.

Key Takeaways

  • Clarify the limit key (user, API key, IP, endpoint) and the scale first. They decide most of the design.
  • Token bucket is the safe default: two stored values per key, burst-friendly, and widely used in production.
  • Sliding window counter fixes the fixed-window boundary spike using two counters and a weighted estimate. Cloudflare reported it was nearly exact on real traffic.
  • In a distributed setup, the check-and-update must be atomic. A Redis Lua script is the standard answer.
  • Decide out loud whether you fail open or fail closed, and why. Interviewers almost always ask.
  • Return 429 with Retry-After and remaining-quota headers so well-behaved clients can back off.

What Does the Interviewer Want From "Design a Rate Limiter"?

The interviewer wants to see you turn a vague prompt into a concrete, scalable component and reason about trade-offs under concurrency. Open with clarifying questions. Each answer changes the design:

  • What are we limiting? Requests per user, per API key, per IP, per endpoint, or a combination.
  • What scale? Requests per second across the fleet, and number of distinct keys. Ten million active keys changes your memory math.
  • Hard or soft limit? Can a client go 1% over in rare cases, or is the limit contractual or security-related?
  • Single region or global? A global exact limit across regions is expensive. Most systems accept per-region limits.
  • What happens on rejection? Drop with 429, queue for later, or degrade the response.
  • Latency budget? The limiter runs on every request, so it should add very little time, typically a single round trip to a nearby cache.

A reasonable set of assumptions to state: 1 million requests per second at peak, 10 million API keys, per-key limits like 100 requests per second with burst, soft accuracy, and a target of a few milliseconds of added latency. This is the same requirements-first structure covered in our system design interview framework.

Where Should the Rate Limiter Live: Client, Gateway, or Service?

The rate limiter should usually live at the API gateway or as shared middleware, with optional limits inside services for expensive operations. Each placement has a different job.

PlacementStrengthsWeaknessesUse it for
Client sideSaves network round trips, smooths retry stormsCannot be trusted; any client can bypass itSDK courtesy throttling, backoff
API gateway or edgeOne choke point, rejects traffic before it costs anything, central configGateway must share state across instancesPer-user and per-key quotas
Middleware in each serviceKnows business context (cost of a query, tenant tier)Duplicated logic unless packaged as a libraryEndpoint-specific or cost-based limits
Sidecar or external limit serviceLanguage-agnostic, reusableExtra network hopLarge microservice fleets

Real systems combine these. Envoy, for example, supports a local token bucket filter on each proxy plus a global rate limit service backed by Redis, per the Envoy documentation. Stripe has described a similar layered approach on its engineering blog: a request rate limiter, a concurrent requests limiter, and load shedders that drop lower-priority traffic when the fleet is overloaded. If you are interviewing there, the Stripe interview guide explains why reliability topics like this come up so often.

In the interview, say you will put the limiter at the gateway, keep rules in a config store, and cache those rules locally on each gateway node.

Which Rate Limiting Algorithm Should You Choose?

Choose token bucket by default, sliding window counter when boundary accuracy matters, and sliding log only when exact counts justify the memory. This table is the core of a strong answer, so practice drawing it from memory.

AlgorithmHow it worksMemory per keyBurstsAccuracyBest fit
Token bucketTokens refill at rate r up to capacity b; each request spends one2 values (tokens, last refill time)Allows bursts up to bExact long-run ratePublic APIs, default choice
Leaky bucketRequests join a FIFO queue drained at a constant rateQueue of up to b itemsSmoothed; excess waits or dropsExact output rateProtecting a fragile downstream, traffic shaping
Fixed window counterCount requests per window like 12:00:00 to 12:00:591 counterUp to 2x the limit at window boundariesWeakestSimple internal quotas, daily caps
Sliding window logStore each request timestamp; count those in the last window1 entry per request in windowNo boundary spikeExactLow-volume, high-value actions (logins, payouts)
Sliding window counterWeighted sum of current and previous window counts2 countersSmall approximation errorNear exactHigh-volume APIs that need smooth limits

A few points interviewers like to hear:

  • The fixed window boundary problem. With a limit of 100 per minute, a client can send 100 at 12:00:59 and 100 at 12:01:00, which is 200 requests in two seconds. Sliding windows exist to fix this.
  • Token bucket versus leaky bucket. Token bucket controls admission and lets a burst through instantly. Leaky bucket controls processing rate and makes output smooth. NGINX's limit_req module is a well-known leaky bucket implementation.

Rate limiter follow-ups come fast: "Now make it distributed." "What if Redis dies?" "How do you handle a burst?" If you tend to blank when the interviewer pivots, TechScreen runs invisibly during Zoom, Google Meet, or Teams screen shares and gives you real-time structure and talking points for system design rounds. Start with 3 free tokens, no credit card.

Get started free →

Token Bucket: How It Works in Code

A token bucket is a counter that refills continuously at a fixed rate up to a maximum capacity, and each request is allowed only if it can spend a token. You do not need a background thread to add tokens. Compute the refill lazily when a request arrives.

import threading
import time


class TokenBucket:
    def __init__(self, capacity: int, refill_per_second: float):
        self.capacity = capacity
        self.refill_per_second = refill_per_second
        self.tokens = float(capacity)
        self.last_refill = time.monotonic()
        self.lock = threading.Lock()

    def allow(self, cost: int = 1) -> bool:
        with self.lock:
            now = time.monotonic()
            elapsed = now - self.last_refill
            self.tokens = min(self.capacity, self.tokens + elapsed * self.refill_per_second)
            self.last_refill = now
            if self.tokens >= cost:
                self.tokens -= cost
                return True
            return False

Things to point out while writing it:

  • time.monotonic() avoids bugs when the wall clock jumps backward.
  • The cost parameter lets an expensive endpoint spend 10 tokens while a cheap one spends 1. This is how you answer "how do you limit by cost, not count?"
  • Capacity sets the burst size and refill rate sets the sustained rate. A bucket with capacity 20 and refill 5 per second allows a burst of 20, then 5 per second.
  • The lock matters even on one machine. Our concurrency interview guide covers why read-modify-write without a lock is a race.

Sliding Window Counter: How It Works in Code

The sliding window counter estimates the count over the last full window by weighting the previous window's count by how much of it still overlaps. If you are 25% into the current minute, the estimate is 75% of the previous minute's count plus the current minute's count.

import threading
import time


class SlidingWindowCounter:
    def __init__(self, limit: int, window_seconds: int):
        self.limit = limit
        self.window = window_seconds
        self.counts: dict[int, int] = {}
        self.lock = threading.Lock()

    def allow(self) -> bool:
        with self.lock:
            now = time.time()
            current_start = int(now // self.window) * self.window
            previous_start = current_start - self.window
            overlap = 1 - (now - current_start) / self.window

            previous = self.counts.get(previous_start, 0)
            current = self.counts.get(current_start, 0)
            estimated = previous * overlap + current

            if estimated >= self.limit:
                return False

            self.counts[current_start] = current + 1
            for start in [s for s in self.counts if s < previous_start]:
                del self.counts[start]
            return True

The approximation assumes requests in the previous window were spread evenly, which breaks for very spiky traffic. Cloudflare's engineering blog reported that across 400 million requests from 270,000 sources, only 0.003% were allowed or limited incorrectly, but that is a measurement on real traffic, not a guarantee.

If you want exact counts, the sliding log variant stores every timestamp in a Redis sorted set: ZREMRANGEBYSCORE drops old entries, ZCARD counts the rest, and ZADD records the new request. It is precise but costs memory proportional to the limit, so 10 million keys at 1,000 requests per window gets expensive fast.

How Do You Build a Distributed Rate Limiter With Redis?

A distributed rate limiter keeps counters in a shared, low-latency store, usually Redis, so every gateway node enforces the same limit for the same key. In-memory buckets on each node do not work alone: with 20 gateway nodes behind a round-robin load balancer, a client could get up to 20 times its limit.

The high-level design:

  1. Request hits a gateway node.
  2. The node builds the key, for example rl:{apikey_123}:tb, and looks up the rule from a locally cached config.
  3. The node runs an atomic script in Redis that refills, checks, and decrements in one step.
  4. Allowed: forward to the service and attach quota headers. Rejected: return 429 with Retry-After.
  5. Rules live in a config service or database and are pushed or polled into gateway caches.

The race condition interviewers look for

The naive approach is GET the count, compare in application code, then SET the new value. Two gateway nodes can both read 99 against a limit of 100, both allow, and both write 100. The limit leaks under exactly the load you built it for.

A second classic bug: calling INCR and then EXPIRE as two separate commands. If the process dies between them, the key never expires and that client is blocked forever. Fix it by making the whole sequence atomic.

Redis executes a Lua script without running other commands in the middle, so the read, decision, and write cannot interleave, as described in the Redis scripting documentation. Here is a token bucket as a Lua script:

local capacity = tonumber(ARGV[1])
local rate = tonumber(ARGV[2])
local cost = tonumber(ARGV[3])

local t = redis.call('TIME')
local now = tonumber(t[1]) * 1000 + math.floor(tonumber(t[2]) / 1000)

local state = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(state[1]) or capacity
local ts = tonumber(state[2]) or now

local elapsed = math.max(0, now - ts) / 1000
tokens = math.min(capacity, tokens + elapsed * rate)

local allowed = 0
if tokens >= cost then
  tokens = tokens - cost
  allowed = 1
end

redis.call('HSET', KEYS[1], 'tokens', tokens, 'ts', now)
redis.call('PEXPIRE', KEYS[1], math.ceil(capacity / rate * 1000) + 1000)

return {allowed, tostring(tokens)}

Design choices worth saying out loud:

  • Clock source. Using Redis TIME gives every gateway node the same clock, which removes skew between application servers.
  • Expiry. The TTL equals the time to refill a full bucket plus a margin, so idle keys clean themselves up and memory stays bounded.
  • Sharding. Use Redis Cluster and shard by the limit key. A script must only touch keys in one hash slot, so if one script checks several tiers for the same user, wrap the shared part in a hash tag like {apikey_123}.
  • Memory math. Two small hash fields are tiny, but Redis adds per-key overhead, so budget a few hundred bytes per key and show the arithmetic for 10 million keys.

For deeper follow-ups on replication, eviction, and hot keys, our distributed cache design guide covers the same Redis trade-offs from the cache side.

How Do You Handle Bursts, Multi-Tier Limits, and Headers?

Bursts, tiers, and headers are where a working design becomes a production-quality one. Expect at least one of these as a follow-up.

Bursts. Token bucket handles bursts by design: capacity is the burst, refill rate is the average. If the interviewer wants zero bursts, lower the capacity toward 1 or switch to a leaky bucket. If they want bursts but also a hard ceiling per minute, stack two limits.

Multi-tier limits. Real APIs enforce several limits at once, for example:

TierExample rulePurpose
Per second20 requests per second per keyStop tight loops and retry storms
Per minute or hour1,000 per minute per keyEnforce plan quota
Per endpoint5 per minute on /login per IPAbuse and brute-force protection
GlobalFleet-wide concurrency capProtect the backend during incidents

Check all tiers for a request in one Lua call so you do not spend from one bucket and then reject on another. Return the most restrictive tier's values in the headers. Plan tiers (free, pro, enterprise) are just different rule values looked up by the client's plan.

Response headers. Reject with HTTP 429 Too Many Requests, defined in RFC 6585, and include Retry-After. Most APIs also send quota headers on every response. GitHub's REST API, for example, returns x-ratelimit-limit, x-ratelimit-remaining, x-ratelimit-used, and x-ratelimit-reset. The IETF HTTPAPI working group has an Internet-Draft standardizing RateLimit-Policy and RateLimit fields. It is still a draft in 2026, so mention it as the direction the industry is heading, not a finished standard.

HTTP/1.1 429 Too Many Requests
Retry-After: 12
RateLimit-Policy: "per-minute";q=1000;w=60
RateLimit: "per-minute";r=0;t=12
Content-Type: application/json

{"error": "rate_limited", "retry_after_seconds": 12}

Clients should honor Retry-After and back off exponentially with jitter, or every rejected client retries at once.

What Are the Failure Modes and Follow-Up Questions?

The failure mode every interviewer asks about is the counter store going down. Have an answer ready, then work through the others.

Failure or follow-upStrong answer
Redis is unavailableFail open for general API limits so a cache outage does not become an API outage; fall back to a local in-memory bucket per node with the limit divided by node count. Fail closed for security limits like login attempts.
Redis adds too much latencyUse a local token bucket per node and sync with Redis periodically, accepting slight overshoot. Some large systems batch counter updates for this reason.
Hot key (one huge customer)Split the key across N sub-keys with limit/N each, or give that customer a dedicated shard.
Multi-regionEnforce per-region limits locally, and replicate usage asynchronously if a global quota is required. Exact global limits require cross-region coordination on every request, which is usually too slow.
Clock skewUse Redis server time inside the script, or a monotonic clock on a single node.
Rule changesStore rules in a config service, cache on gateways with a short TTL or push updates, and version them.
Limiter itself overloadedShed load: reject lowest-priority traffic first, as Stripe describes with its load shedders.
Abuse across many IPsLimit by account or API key in addition to IP; IP limits alone miss distributed attackers and punish shared NATs.

Two more follow-ups that come up at senior and staff level:

  • "How do you test it?" Unit tests with an injected fake clock, load tests that hit boundaries, and a shadow mode that logs would-be rejections before enforcing a new rule.
  • "How do you observe it?" Metrics for allowed and rejected counts per rule, Redis latency percentiles, and alerts when rejection rates jump, which usually signals a client bug or an attack.

If the interviewer keeps pushing on trade-offs rather than components, they are probing for scope and judgment. Our staff engineer interview guide covers what that bar looks like.

A 45-Minute Answer Plan

Spend most of your time on the distributed design and failure handling, not on reciting algorithms:

  1. Minutes 0 to 5: requirements. Limit key, scale, hard or soft, single or multi-region, rejection behavior.
  2. Minutes 5 to 12: placement and API. Gateway limiter, rules config, 429 plus headers.
  3. Minutes 12 to 22: algorithm choice. Draw the comparison table, pick token bucket or sliding window counter, write the core logic.
  4. Minutes 22 to 34: distributed design. Redis, the race condition, Lua atomicity, sharding, TTLs, memory math.
  5. Minutes 34 to 45: failure modes and extensions. Fail open versus closed, hot keys, multi-tier, multi-region, monitoring.

Narrate each decision as you make it. Interviewers score the reasoning as much as the diagram, which is why thinking out loud matters as much here as in coding rounds. Rehearse this plan at least twice under a timer using the methods in our guide on how to practice system design interviews, then pair it with a related classic like the URL shortener design, which reuses the same key-value and caching ideas. Rate limiting also shows up regularly in backend engineer interviews as a deep-dive inside larger API design questions.

Knowing the token bucket is the easy part. Holding the Redis race condition, fail-open trade-off, and multi-tier headers in your head while an interviewer interrupts is harder. TechScreen is an invisible AI interview assistant that stays hidden during screen shares on Zoom, Google Meet, Teams, HackerRank, and CoderPad, and helps you structure system design answers in real time. Try it with 3 free tokens, no credit card required.

Get started free →

Frequently Asked Questions

Which rate limiting algorithm should I pick in a system design interview?

Default to the token bucket for public APIs because it allows short bursts while enforcing a long-run average, uses only two values per key, and is what many production systems use. Switch to a sliding window counter when the interviewer cares about smooth enforcement at window boundaries with low memory. Choose a sliding log only when exact accuracy matters more than memory, such as low-volume, high-value actions like login attempts or payouts.

What is the difference between token bucket and leaky bucket?

A token bucket controls how many requests are admitted: tokens refill at a fixed rate, each request spends one, and a full bucket lets a burst through at once. A leaky bucket controls how fast requests are processed: requests enter a queue that drains at a constant rate, so output is smooth and bursts wait or get dropped. Token bucket suits APIs where bursts are fine; leaky bucket suits protecting a downstream that needs steady load.

Why does a distributed rate limiter need Redis Lua scripts or atomic operations?

Multiple application servers read and write the same counter at the same time. If each server reads the count, checks it, and writes it back in separate steps, two servers can both read 99 against a limit of 100 and both allow the request. Running the read, decision, and write inside a single Lua script makes the sequence atomic on the Redis server, because Redis executes a script without interleaving other commands.

What HTTP status code and headers should a rate limiter return?

Return HTTP 429 Too Many Requests, defined in RFC 6585, when a client exceeds its limit. Add a Retry-After header telling the client how many seconds to wait. Most APIs also send remaining-quota headers on every response; GitHub, for example, uses x-ratelimit-limit, x-ratelimit-remaining, and x-ratelimit-reset. The IETF HTTPAPI working group is standardizing RateLimit and RateLimit-Policy fields, still at Internet-Draft stage in 2026.

Should a rate limiter fail open or fail closed when Redis is down?

Most public API rate limiters fail open: if the counter store is unreachable, requests are allowed so a cache outage does not become a full API outage. Pair that with a local in-memory limiter as a fallback so the service still has rough protection. Fail closed only for security-sensitive limits, such as login attempts or one-time password checks, where letting unlimited traffic through is worse than briefly rejecting legitimate users.

Where should a rate limiter live in the architecture?

In most designs the rate limiter lives at the API gateway or as middleware in front of services, because that is the single place every request passes through and it can reject traffic before it costs anything. Client-side limiting is useful but cannot be trusted. Service-level limits inside individual services add protection for expensive internal operations. Strong answers usually combine a gateway limiter for per-client quotas with service-level load shedding.

How long should I spend on a rate limiter question in a 45-minute interview?

Spend about five minutes on requirements, five to eight on placement and API, ten on algorithm choice with a comparison, ten to twelve on the distributed design with Redis and race conditions, and the rest on failure modes and follow-ups. Interviewers usually push hardest on the distributed part, so avoid spending twenty minutes explaining algorithms and running out of time before Redis, atomicity, and failure handling.

Ready to use AI assistance in your next interview?

TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.

Ace your next interview →