Inside LLM inference
Drive a live serving engine. Fire prompts at one GPU and watch prefill chunks, decode streams and KV blocks fight over the same step budget - then break it: turn off continuous batching, shrink the KV pool until requests get preempted, and see tokens/sec collapse.
18 min read
A serving engine is the machinery around the model weights: a scheduler that decides what runs in the next forward pass, and a memory manager for the attention state (the KV cache) that every in-flight token keeps on the card. This article builds one up a problem at a time. The figures run a small simulated engine; press play on any of them.
A single request barely touches the GPU
A request happens in two acts. Prefill reads the whole prompt in a few big forward passes. Decode then produces one token per pass, because token n+1 depends on token n.
Each forward pass has a budget of thousands of tokens, and one decoding sequence contributes exactly one. It does not matter how big the GPU is; with one user the pass is almost entirely empty, so you pay for a card that is mostly waiting. The fix is never a faster card. It is more sequences sharing the same pass.
A batch that has to finish together wastes most of it
Put several sequences in one pass and they share the cost of reading the weights. The question is what happens when one finishes. Static batching makes the batch start and end together, so a finished slot sits empty until the slowest member is done. Continuous batching treats the batch as a set that changes every step, so a freed slot is refilled on the next one.
With static batching, seven short requests can finish in 30 tokens and sit idle, billed in full, while one runs to 256. Continuous batching keeps occupancy high, so throughput follows arrivals instead of the slowest request. It is the single biggest win in modern serving engines.
One long prompt can freeze every stream on the box
Prefill wants thousands of tokens in one pass. If a long prompt takes the whole pass, every user already streaming gets nothing that step. Their sequences are fine; they are just frozen.
A few steps like that is a visible stall for everyone on the machine, and the longer the context, the longer the freeze. Chunked prefill caps how much of a pass any one prompt may take, spreading it across steps so decode keeps its rhythm in between. Admitting a long context then costs the new user a little latency instead of costing every existing user their stream.
You run out of memory long before you run out of compute
Every token in flight keeps its keys and values (the KV cache) on the card, and that state grows with every token generated. Modern engines hand out memory in blocks as a sequence grows (paged attention) instead of reserving the worst case up front. When the pool runs dry, something has to give.
KV state is not a cache you can politely drop a line from; it is the sequence. Evict one (preemption) and every token it generated goes with it, to be recomputed from scratch later. That is why serving engines watch KV pressure more nervously than utilisation: running out of memory does not slow you down, it destroys work you already paid for.
Two prompts that share a prefix should not both pay for it
A system prompt, a long document, a chat history: the same tokens arrive over and over. Their KV state can be computed once and reused across requests, but only in whole blocks, because cache lookups happen block by block.
The incomplete final block is prefilled again on every request, forever. It is the kind of alignment detail that never shows up in a diagram and always shows up in the bill: on a busy system, padding a prompt to a block boundary is free performance.
Guess several tokens cheaply, then check them all at once
Decode is memory bound, so a forward pass that checks five candidate tokens costs barely more than one that produces one. Speculative decoding uses a small draft model to propose several tokens, and the big model verifies them in a single pass, keeping the ones it agrees with.
The draft is a sequence, so the tokens are not independent: once one is rejected, everything after it is built on a wrong prefix and thrown away. Surviving to position k costs the acceptance rate to the power k. High-acceptance workloads (code, structured output, repetitive text) get real speed-ups; open-ended chat with a weak draft model can get slower.
Every trick here is the same trick: keep the forward pass full without letting anyone hold the machine hostage. The hardware limits underneath it are in the GPU lab.
The short version
- Decode makes one token per sequence per pass, so a lone request wastes the GPU.
- Continuous batching refills finished slots every step.
- Chunked prefill stops long prompts from freezing everyone else.
- KV memory caps concurrency; prefix caching and speculative decoding skip work you can avoid.