Queueing
Requests pile up faster than a server can drain them. Send traffic into FIFO, LIFO and priority queues, watch the line grow, and see tail latency explode as load approaches capacity.
9 min read
A queue lets a server absorb a burst instead of dropping it. What a queue cannot do is add capacity. Almost every surprising thing about latency under load follows from that one sentence.
Latency does not degrade gracefully. It has a cliff
Utilisation is the arrival rate divided by the service rate: how busy the server is. Below 100% it keeps up on average and the queue drains. The intuition is that waiting grows in proportion to load. It does not. In the classic model, queueing delay grows as ρ / (1 - ρ), where ρ is utilisation.
At 50% that factor is 1. At 90% it is 9, at 99% it is 99, and the server is “not overloaded” the whole way. That is what makes it dangerous: the capacity dashboard is green while the slowest requests are on fire. Watch it happen on a real queue.
The median barely moves; the 99th percentile climbs long before utilisation reaches 100%. By the time the queue is visibly long, the tail is already gone. Below capacity a queue turns burstiness into delay, which is useful. At capacity it turns everything into delay.
Same throughput, completely different victims
The discipline decides who goes next: first in first out (FIFO), newest first (LIFO), or a priority class first. It cannot change how fast the server works, so throughput is identical under all three. What it changes is the distribution of waiting, and the distribution is what users actually feel.
FIFO is fair: everyone waits about the same, so median and tail stay close. LIFO makes the median look great, because most requests are the newest and get served almost at once, but the cost lands on the requests at the bottom of the stack, which may never be served at all. Same average, far worse tail, and an easy data structure to reach for by accident.
Priority lets you choose who absorbs the delay: checkout stays fast while an analytics job waits. That works as long as the high-priority class cannot saturate the server on its own.
A full queue should say no, not grow
Under sustained overload an unbounded queue grows forever. A request queued ten minutes into an overload will be served long after the user gave up, refreshed and sent another. You still do the work, nobody is there to receive it, and the memory holding the queue climbs toward an out-of-memory crash.
A bound turns that invisible, unbounded failure into an immediate, visible, recoverable one. That is backpressure: when the queue is full, the honest answer is no. Size the queue for the burst you expect to absorb, not the overload you hope to survive, and make the rejection loud. A 429 with a retry hint is far kinder than a request still technically in progress behind forty thousand others. What the caller should do with that no is the subject of the retries lab.
The short version
- Queueing delay grows as ρ / (1 - ρ): nine service times at 90%, ninety-nine at 99%.
- Watch p99, not queue length; the tail goes first.
- Disciplines share throughput but not pain: LIFO hides starvation behind a good median.
- Bound every queue and reject the overflow; a queue absorbs bursts, it does not add capacity.