How a GPU runs your model

Warps, occupancy and the memory wall, made playable. Starve an SM of registers until it has nothing to switch to, watch stalls stop being hidden, then move a kernel along the roofline and see why one user generating one token uses a fraction of a percent of the card.

9 min read

Before an inference engine can be clever, it has to live inside a very particular machine: thousands of small cores that only go fast when they all do the same thing, behind memory that is nowhere near fast enough to feed them. Every scheduling trick in a serving engine is a response to the three figures below.

A core waiting on memory needs somewhere else to be

GPU threads run in groups of 32 called warps, in lockstep. A streaming multiprocessor (SM) keeps many warps resident, and when one stalls on a memory read it switches to another in a single cycle. That is the whole trick for hiding memory latency. Occupancy is how many warps it has to switch to, and registers are what you pay for them with.

issuing (black)readystalled on memorygrey: no warp resident
Occupancy
50%
Issue slots used
75%
Warps resident
32 of 64
Divergence cost
1x
64
256
55%
At 64 registers per thread each SM keeps 32 warps to switch between. Push registers to 128 and watch the slots empty and the stalls show through.

Each SM has one register file of 65,536 registers, shared by every warp on it. At 32 registers per thread, 64 warps fit; at 128, only 16. Giving each thread more registers feels like it should help, since values held in registers are memory trips avoided. But the moment a warp stalls with nothing queued behind it, the issue slots go empty and the core waits. Branches have a cost too: the 32 lanes of a warp share one program counter, so an if/else runs both sides with half the lanes masked off.

Every level down the hierarchy costs an order of magnitude

The tensor cores are not the scarce resource. Getting bytes to them is. Memory comes in levels, each bigger and slower than the last, and a kernel's working set decides which one it lives in.

Registers256 KB per SM
~1 cycle
Shared memory / L1256 KB per SM
~30 cycles
L2 cache50 MB, shared
~200 cycles
HBM380 GB
~500 cycles
A 3.1 MB working set lands in l2 cache. To hide a 200-cycle wait you need about 50 other warps ready to run.
3.1 MB
Each level down costs roughly an order of magnitude. Drag past 256 KB, then past 50 MB.

Shared memory answers in about 30 cycles; L2, on the same chip but shared by all 132 SMs, in about 200; HBM in about 500. Those are cliffs, not slopes, and the cost shows up as how many other warps you need ready just to paper over each wait. Model weights are gigabytes, so they live in HBM, and they are read again for every single decode step. No amount of occupancy makes that read free.

Why one user gets a fraction of a percent of the card

Arithmetic intensity is the number of floating-point operations done per byte moved from memory. Below a certain intensity, the ridge point, a kernel waits on memory and the tensor cores sit idle. Above it, compute is the limit. The roofline plots both limits at once.

FLOP per byte
1
Achieved
3 TF
Of peak
0.34%
Limited by
memory
1
One user decoding one token sits at the far left: about 1 FLOP per byte, a third of a percent of the card. Raise the batch and it climbs the roof.

Generating one token for one user is a matrix-vector product: it reads every weight in the model and does about one operation per byte read. The ridge on this card is around 295 operations per byte, so decode sits two and a half orders of magnitude below it, and the headline 989 TFLOP/s is irrelevant.

That is the whole argument for batching. Every extra sequence decoded together reuses the same weight bytes for another token, which walks the marker up the memory roof for free. Packing batches is what a serving engine is really doing, and the inference lab picks up the story from there.

The short version

  • GPUs hide memory latency by switching between resident warps; registers limit how many fit.
  • Each step down the memory hierarchy costs roughly ten times more.
  • Decode for one user does about 1 FLOP per byte, far below the ~295 ridge: memory bound.
  • Batching reuses each weight read across many sequences, which is why serving engines batch.