Inside LLM inference

Drive a live serving engine. Fire prompts at one GPU and watch prefill chunks, decode streams and KV blocks fight over the same step budget - then break it: turn off continuous batching, shrink the KV pool until requests get preempted, and see tokens/sec collapse.

18 min read

A serving engine is the machinery around the model weights: a scheduler that decides what runs in the next forward pass, and a memory manager for the attention state (the KV cache) that every in-flight token keeps on the card. This article builds one up a problem at a time. The figures run a small simulated engine; press play on any of them.

A single request barely touches the GPU

A request happens in two acts. Prefill reads the whole prompt in a few big forward passes. Decode then produces one token per pass, because token n+1 depends on token n.

The request
#1
decode13/96
idle slot
What each forward pass spent
prefill tokens decode tokensan empty column is a step that produced nothing
Tokens coming out
batching lets the scheduler swap a finished sequence out of the running set
Step budget used
4%
Budget per step
2048 tokens
Output
31 tok/s
Sequences
1
1.0k tokens
96 tokens
One tall prefill spike, then a sliver per step: one token out of a 2,048-token budget. Send a few more requests and watch the slivers stack.

Each forward pass has a budget of thousands of tokens, and one decoding sequence contributes exactly one. It does not matter how big the GPU is; with one user the pass is almost entirely empty, so you pay for a card that is mostly waiting. The fix is never a faster card. It is more sequences sharing the same pass.

A batch that has to finish together wastes most of it

Put several sequences in one pass and they share the cost of reading the weights. The question is what happens when one finishes. Static batching makes the batch start and end together, so a finished slot sits empty until the slowest member is done. Continuous batching treats the batch as a set that changes every step, so a freed slot is refilled on the next one.

Batch slots
#1
decode59/64
#2
decode59/64
#3
decode58/64
#4
decode57/64
#5
decode57/64
#6
decode56/64
idle slot
idle slot
Slots busy
6 of 8
Queued
5
Output
200 tok/s
Completed
0
2.7 req/s
64 tokens
Static batching: finished slots sit empty while the queue grows. Turn on continuous batching and press play.

With static batching, seven short requests can finish in 30 tokens and sit idle, billed in full, while one runs to 256. Continuous batching keeps occupancy high, so throughput follows arrivals instead of the slowest request. It is the single biggest win in modern serving engines.

One long prompt can freeze every stream on the box

Prefill wants thousands of tokens in one pass. If a long prompt takes the whole pass, every user already streaming gets nothing that step. Their sequences are fine; they are just frozen.

What each forward pass spent, last 48 steps
prefill tokens decode tokensan empty column is a step that produced nothing
Streams in flight
#1
decode30/64
#2
decode5/64
#3
decode5/64
#4
decode5/64
#5
decode5/64
Frozen steps
10
Median first token
1050 ms
Output
55 tok/s
Prefill
whole prompt
A 16k prompt arrived and took whole forward passes: solid teal, and every stream froze. Turn chunking on and send another.

A few steps like that is a visible stall for everyone on the machine, and the longer the context, the longer the freeze. Chunked prefill caps how much of a pass any one prompt may take, spreading it across steps so decode keeps its rhythm in between. Admitting a long context then costs the new user a little latency instead of costing every existing user their stream.

You run out of memory long before you run out of compute

Every token in flight keeps its keys and values (the KV cache) on the card, and that state grows with every token generated. Modern engines hand out memory in blocks as a sequence grows (paged attention) instead of reserving the worst case up front. When the pool runs dry, something has to give.

KV blocks on the card, coloured by the sequence holding them (256 tokens each)
KV used
88%
Preemptions
1
Waiting
3
Completed
2
24k tokens
4k
2.7/s
Long prompts fill the pool fast. Shrink it and press play: admission stalls first, then running sequences are evicted and their work thrown away.

KV state is not a cache you can politely drop a line from; it is the sequence. Evict one (preemption) and every token it generated goes with it, to be recomputed from scratch later. That is why serving engines watch KV pressure more nervously than utilisation: running out of memory does not slow you down, it destroys work you already paid for.

Two prompts that share a prefix should not both pay for it

A system prompt, a long document, a chat history: the same tokens arrive over and over. Their KV state can be computed once and reused across requests, but only in whole blocks, because cache lookups happen block by block.

Each bar is a 16-token block; red is the incomplete last block
Reused from cache
2032 tokens
Recomputed every request
15 tokens
Prefill saved
99%
2047 tokens
2,047 tokens leaves a 15-token tail recomputed on every request. Make it 2,048 and the whole prefix is free.

The incomplete final block is prefilled again on every request, forever. It is the kind of alignment detail that never shows up in a diagram and always shows up in the bill: on a busy system, padding a prompt to a block boundary is free performance.

Guess several tokens cheaply, then check them all at once

Decode is memory bound, so a forward pass that checks five candidate tokens costs barely more than one that produces one. Speculative decoding uses a small draft model to propose several tokens, and the big model verifies them in a single pass, keeping the ones it agrees with.

100%always
80%draft 1
64%draft 2
51%draft 3
41%draft 4
Chance each position survives verification
Tokens per pass
3.36
Extra verify cost
+72%
Net speed-up
1.95x
80%
At 80% acceptance the draft yields about 3.4 tokens a pass, not 4.2. Drag below 40% and speculation costs more than it saves.

The draft is a sequence, so the tokens are not independent: once one is rejected, everything after it is built on a wrong prefix and thrown away. Surviving to position k costs the acceptance rate to the power k. High-acceptance workloads (code, structured output, repetitive text) get real speed-ups; open-ended chat with a weak draft model can get slower.

Every trick here is the same trick: keep the forward pass full without letting anyone hold the machine hostage. The hardware limits underneath it are in the GPU lab.

The short version

  • Decode makes one token per sequence per pass, so a lone request wastes the GPU.
  • Continuous batching refills finished slots every step.
  • Chunked prefill stops long prompts from freezing everyone else.
  • KV memory caps concurrency; prefix caching and speculative decoding skip work you can avoid.