How a GPU runs your model
Warps, occupancy and the memory wall, made playable. Starve an SM of registers until it has nothing to switch to, watch stalls stop being hidden, then move a kernel along the roofline and see why one user generating one token uses a fraction of a percent of the card.
9 min read
Before an inference engine can be clever, it has to live inside a very particular machine: thousands of small cores that only go fast when they all do the same thing, behind memory that is nowhere near fast enough to feed them. Every scheduling trick in a serving engine is a response to the three figures below.
A core waiting on memory needs somewhere else to be
GPU threads run in groups of 32 called warps, in lockstep. A streaming multiprocessor (SM) keeps many warps resident, and when one stalls on a memory read it switches to another in a single cycle. That is the whole trick for hiding memory latency. Occupancy is how many warps it has to switch to, and registers are what you pay for them with.
Each SM has one register file of 65,536 registers, shared by every warp on it. At 32 registers per thread, 64 warps fit; at 128, only 16. Giving each thread more registers feels like it should help, since values held in registers are memory trips avoided. But the moment a warp stalls with nothing queued behind it, the issue slots go empty and the core waits. Branches have a cost too: the 32 lanes of a warp share one program counter, so an if/else runs both sides with half the lanes masked off.
Every level down the hierarchy costs an order of magnitude
The tensor cores are not the scarce resource. Getting bytes to them is. Memory comes in levels, each bigger and slower than the last, and a kernel's working set decides which one it lives in.
Shared memory answers in about 30 cycles; L2, on the same chip but shared by all 132 SMs, in about 200; HBM in about 500. Those are cliffs, not slopes, and the cost shows up as how many other warps you need ready just to paper over each wait. Model weights are gigabytes, so they live in HBM, and they are read again for every single decode step. No amount of occupancy makes that read free.
Why one user gets a fraction of a percent of the card
Arithmetic intensity is the number of floating-point operations done per byte moved from memory. Below a certain intensity, the ridge point, a kernel waits on memory and the tensor cores sit idle. Above it, compute is the limit. The roofline plots both limits at once.
Generating one token for one user is a matrix-vector product: it reads every weight in the model and does about one operation per byte read. The ridge on this card is around 295 operations per byte, so decode sits two and a half orders of magnitude below it, and the headline 989 TFLOP/s is irrelevant.
That is the whole argument for batching. Every extra sequence decoded together reuses the same weight bytes for another token, which walks the marker up the memory roof for free. Packing batches is what a serving engine is really doing, and the inference lab picks up the story from there.
The short version
- GPUs hide memory latency by switching between resident warps; registers limit how many fit.
- Each step down the memory hierarchy costs roughly ten times more.
- Decode for one user does about 1 FLOP per byte, far below the ~295 ridge: memory bound.
- Batching reuses each weight read across many sequences, which is why serving engines batch.