Design a rate limiter
Throttle a flood of requests four different ways - token bucket, leaking bucket, fixed window and sliding window. Open the tap, watch requests get accepted or 429'd in real time, and feel exactly where each algorithm leaks or bursts.
12 min read
A rate limiter decides, request by request, who gets through and who gets a 429 Too Many Requests. There are four standard algorithms. They all bound the same average rate, and they behave completely differently the moment traffic is bursty, which is the only kind of traffic there is.
One client can consume a service built for all of them
Rate limiting is usually introduced as a fairness or billing feature. Its real job is structural: without it, any single caller, malicious, buggy or merely enthusiastic, sets the load for everybody.
Queues are first in, first out, not fair. If one client sends four-fifths of the traffic, four-fifths of every queue position is theirs, and everyone else waits behind it and times out. Their retries make it worse. A load balancer does not help: it spreads the traffic evenly across servers, which means it spreads the outage evenly too.
A rate limiter is the component that can say “this caller has had enough”, and it has to reject cheaply, before authentication and before the database, ideally at the edge. Otherwise the rejection costs as much as the work.
Same average rate, four different shapes of pain
Token bucket holds up to N tokens and refills at a steady rate; each request spends one. Leaking bucket queues requests and drains them at a fixed rate. Fixed window counts requests per second (or minute) and resets on the boundary. Sliding window blends the previous window into the current one. Same average, very different handling of a burst.
12.0 / 12 tokens
The token bucket accumulates permission while a client is idle, so a client that has been quiet can spend it all at once. That matches how real clients behave, and it is why it is the industry default. The leaking bucket has no memory of good behaviour: it drains at exactly the configured rate no matter what, which is what you want when the thing downstream cannot absorb a spike at all.
A fixed window lets through twice what you configured
The fixed window is the easiest to build: one counter per client, reset on the minute. It has one well-known hole, and it shows up exactly when someone goes looking for it.
Both batches fall inside their own window, so neither breaks the rule, yet a limit meant to bound any rolling minute just allowed 200 requests in about a second. The sliding window counter fixes it by weighting the previous window's count by how much of it still overlaps the rolling window. It costs one extra number per key; Cloudflare measured it misjudging about 0.003% of requests.
Fixed windows still have a place. One integer, one increment, one expiry: for coarse quotas like “10,000 calls a month”, a 2x overshoot at one boundary does not matter. Use a sliding window when the limit is a defence rather than a quota.
The algorithm was the easy part
Everything so far assumed one process holding one counter. In production the limiter runs on every instance behind the load balancer, and the counter becomes a distributed-systems problem.
A shared counter also has to be updated atomically. If two servers each read 99, decide there is room, and both increment, the limit leaks under exactly the load it exists for. Redis handles this with an atomic script or a sorted set per client.
A production limiter also owes the client an honest answer: a 429 with Retry-After and X-RateLimit-Remaining headers, so a well-behaved client can back off instead of guessing. That is the difference between shedding load and triggering the retry storm in the retries lab.
Which one should you reach for?
Token bucket is the default for API rate limiting today: it bounds the average while allowing the short bursts real clients produce, in constant memory that maps onto one Redis counter. Reach past it for a specific reason.
The short version
- Without a limiter, one noisy client sets the latency for everyone.
- Token bucket allows saved-up bursts; leaking bucket enforces a perfectly smooth rate.
- Fixed windows allow 2x at the boundary; sliding windows close the hole for one extra number.
- Per-server counters multiply the limit by the fleet size; share them and update atomically.