Retries

Retrying a failed request seems harmless - until every client retries at once and buries a struggling service. Compare naive retries, fixed delay, exponential backoff and jitter, and watch the retry storm form (or not).

9 min read

Retrying a failed request is the most natural instinct in distributed systems, and one of the most dangerous. The fix is never “do not retry”. It is how you space the retries out, and knowing when to stop.

Every client helpfully retries, all at once

A service wobbles. Every client does the obvious, reasonable thing and tries again straight away. Now the struggling server receives its normal traffic plus a copy of everything that just failed, which makes more things fail, which sends more retries.

Attempts reaching the server each second
The outage ended at 15 s. The service never recovers: retries keep it overloaded, so it keeps failing. 4.0 attempts per request; 29% succeed.
5
A seven-second outage, and the clients keep the service down long after it ends.

That is a feedback loop, not an addition. Each retry arrives while the server is less able to handle it than before. A service that would have shrugged off a brief outage is held down by its own clients, and the load does not subside when the original cause does. The retry is not the bug. Retrying at the same moment as everyone else is the bug.

Wait longer each time, and do not all wait the same amount

Exponential backoff doubles the wait after each failure: 1 second, then 2, then 4. It gives the server geometrically more breathing room early in an outage. It leaves a second problem untouched.

Busiest half second: 1000 retries. Every client runs the same timer from the same start, so they arrive together at 1, 3, 7 and 15 seconds.
Pure backoff lowers the average and keeps the clients perfectly in step. Jitter breaks the step.

Clients that failed together run the same timer from the same start, so pure backoff gives quiet periods punctuated by thundering herds. Jitter is the fix: randomise each wait so a thousand clients pick a thousand different moments. It is one line of code, and it is the difference between a smooth recovery and a drumbeat of self-inflicted spikes. Fixed delays are worse, because they never back off at all.

Notice what spacing cannot do. Switch the storm figure above to backoff with jitter: the spikes smooth out, but the service still does not recover. Spacing changes when the retries arrive, not how many there are, and each request still makes the same number of attempts.

Backoff decides when to retry. Something else decides whether

Perfect jittered backoff still sends traffic to a service that is completely down. And in a chain of services each layer multiplies the one below it: if A retries three times, each of those attempts makes B retry three times, and so on down.

One user action
1
Service A
3
Service B
9
Service C
27
Service D (down)
81
One click becomes 81 requests at a service that is already down.
4
3
Four layers making three attempts each send 81 requests to the broken service. Retry at one layer and it is three.

The fix for the storm, then, is a limit on how much retrying is allowed at all. Here is the same outage with retries capped at 10% of traffic; anything over the budget fails fast instead of piling on.

Attempts reaching the server each second
The outage ended at 15 s and the service was healthy at once. 1.1 attempts per request; 88% succeed.
5
The same outage with retries capped at 10% of traffic: the server never leaves its capacity, whatever the strategy.

Three rules follow. Retry at one layer of a call chain, not at every layer. Express the budget as a share of traffic (“retries may not exceed 10% of requests”) rather than a count per request, so a widespread failure cannot multiply your load. And put a circuit breaker underneath: after enough consecutive failures, stop calling and fail fast. That is also the only way the struggling service gets the quiet it needs to recover. The server's side of this story is in the queueing lab.

The short version

  • Retries can keep a server overloaded long after the outage that triggered them.
  • Exponential backoff spaces retries out and jitter breaks up the herd, but neither caps the total.
  • Retries multiply down a call chain, so retry at one layer only.
  • Cap retries as a share of traffic and add a circuit breaker so a dead service gets quiet.