Design a rate limiter

Throttle a flood of requests four different ways - token bucket, leaking bucket, fixed window and sliding window. Open the tap, watch requests get accepted or 429'd in real time, and feel exactly where each algorithm leaks or bursts.

12 min read

A rate limiter decides, request by request, who gets through and who gets a 429 Too Many Requests. There are four standard algorithms. They all bound the same average rate, and they behave completely differently the moment traffic is bursty, which is the only kind of traffic there is.

One client can consume a service built for all of them

Rate limiting is usually introduced as a fairness or billing feature. Its real job is structural: without it, any single caller, malicious, buggy or merely enthusiastic, sets the load for everybody.

the noisy clienteveryone else
Everyone else gets 22% of their requests served. The service is 4.6x over capacity, and most of the queue belongs to one caller.
4x
One buggy retry loop at four times capacity, and everyone else is mostly timing out. Turn the limiter on.

Queues are first in, first out, not fair. If one client sends four-fifths of the traffic, four-fifths of every queue position is theirs, and everyone else waits behind it and times out. Their retries make it worse. A load balancer does not help: it spreads the traffic evenly across servers, which means it spreads the outage evenly too.

A rate limiter is the component that can say “this caller has had enough”, and it has to reject cheaply, before authentication and before the database, ideally at the edge. Otherwise the rejection costs as much as the work.

Same average rate, four different shapes of pain

Token bucket holds up to N tokens and refills at a steady rate; each request spends one. Leaking bucket queues requests and drains them at a fixed rate. Fixed window counts requests per second (or minute) and resets on the boundary. Sliding window blends the previous window into the current one. Same average, very different handling of a burst.

12.0 / 12 tokens

+10/s refill-1 per request
A bucket holds up to capacity tokens and refills at a steady rate. Each request spends one token; no token, no entry. Because a full bucket can be spent at once, it allows short bursts - which is exactly why Amazon and Stripe use it. Used by most API gateways.
Accepted
0
Rejected (429)
0
Accept rate
0%
Latest requests, newest on the right
Best for: Public APIs and gateways that must bound the average rate but still let real clients fire short bursts.
Watch out: A client can drain a full bucket instantly - size the burst to what your backend can actually absorb.
8/s
12
10/s
Each opens on a few seconds of traffic with a burst in the middle. Press play, then send a burst into each algorithm.

The token bucket accumulates permission while a client is idle, so a client that has been quiet can spend it all at once. That matches how real clients behave, and it is why it is the industry default. The leaking bucket has no memory of good behaviour: it drains at exactly the configured rate no matter what, which is what you want when the thing downstream cannot absorb a spike at all.

A fixed window lets through twice what you configured

The fixed window is the easiest to build: one counter per client, reset on the minute. It has one well-known hole, and it shows up exactly when someone goes looking for it.

Limit 100 a minute. 100 requests at 0:59, then 100 more 1 second after the reset. 200 got through in 2 seconds, every one of them legal under a fixed window.
1 s
A fixed window has no memory across the boundary. A sliding window weights the previous minute by how much it still overlaps.

Both batches fall inside their own window, so neither breaks the rule, yet a limit meant to bound any rolling minute just allowed 200 requests in about a second. The sliding window counter fixes it by weighting the previous window's count by how much of it still overlaps the rolling window. It costs one extra number per key; Cloudflare measured it misjudging about 0.003% of requests.

Fixed windows still have a place. One integer, one increment, one expiry: for coarse quotas like “10,000 calls a month”, a 2x overshoot at one boundary does not matter. Use a sliding window when the limit is a defence rather than a quota.

The algorithm was the easy part

Everything so far assumed one process holding one counter. In production the limiter runs on every instance behind the load balancer, and the counter becomes a distributed-systems problem.

Configured: 100 a minute per client. Experienced: 1,000 a minute. Each server keeps its own counter, so each grants its own 100. Free, and wrong; the limit changes every time you autoscale.
10
Ten servers with their own counters enforce 1,000 a minute, not 100.

A shared counter also has to be updated atomically. If two servers each read 99, decide there is room, and both increment, the limit leaks under exactly the load it exists for. Redis handles this with an atomic script or a sorted set per client.

A production limiter also owes the client an honest answer: a 429 with Retry-After and X-RateLimit-Remaining headers, so a well-behaved client can back off instead of guessing. That is the difference between shedding load and triggering the retry storm in the retries lab.

Which one should you reach for?

Token bucket is the default for API rate limiting today: it bounds the average while allowing the short bursts real clients produce, in constant memory that maps onto one Redis counter. Reach past it for a specific reason.

Token bucketthe defaultPublic APIs and gateways that must bound the average rate but still let real clients fire short bursts. Used by AWS API Gateway, Stripe, GitHub API, Envoy, NGINX upstream.
Leaking bucketTraffic shaping - feeding a fragile downstream (payments, legacy DB, a metered 3rd-party API) a perfectly constant rate. Used by NGINX limit_req, Shopify, job / message pipelines.
Fixed windowCoarse quotas and billing counters - '10k requests/day', '5 OTPs/hour' - where the exact edge doesn't matter. Used by In-house quota counters, early-stage APIs.
Sliding windowAccurate, fair limiting at high scale - abuse/DDoS protection at the edge - without the fixed-window loophole. Used by Cloudflare, Kong, Redis cell-rate limiters.

The short version

  • Without a limiter, one noisy client sets the latency for everyone.
  • Token bucket allows saved-up bursts; leaking bucket enforces a perfectly smooth rate.
  • Fixed windows allow 2x at the boundary; sliding windows close the hole for one extra number.
  • Per-server counters multiply the limit by the fleet size; share them and update atomically.