The agent loop
Think, call a tool, observe, repeat. Watch an agent work a task turn by turn while its context window fills, then choose what happens when it runs out - truncate, compact, or retrieve - and watch the agent forget a result it still needed.
12 min read
An agent is not a model. It is a loop around a model: a harness that decides what goes into the next prompt, runs whatever tool the model asks for, and decides when to stop. This article adds one failure at a time. Each figure runs an agent to the end with your settings; scrub back through it, change a setting, or run it again.
Every turn re-sends everything that happened before it
The loop is think, call a tool, run it, read the result, repeat. The model is stateless, so to take turn ten the harness sends turns one through nine again, plus the system prompt and every tool definition. Nothing is broken yet, and it is already expensive.
Because the prompt grows every turn, the total billed is the sum of a growing series: it scales with the square of the turn count. Doubling the turns roughly quadruples the bill. Prompt caching softens the constant, not the shape. Parallel tool calls finish the same task in fewer turns, which makes them the cheapest optimisation a harness has.
Sooner or later it does not fit, and something has to go
Tool results are the biggest things in the transcript, and they never stop arriving. When the context window fills, the harness has to choose what the agent stops knowing.
The context is the memory. There is nowhere else a conclusion was written down, so dropping a message drops the finding, and the agent goes and gets it again, paying for the tool call, the tokens and the turn. Dropping the oldest is cheapest and most brutal. Summarising keeps a low-fidelity trace of everything, which is how a compacted agent ends up confidently half-remembering a number. Offloading results to a store and retrieving them on demand costs a turn per lookup, but nothing is ever truly gone.
The call will not parse, and the tool times out
Real harnesses spend a surprising share of their turns neither thinking nor working: they reject a malformed tool call, feed the error back, and ask again, or they retry a tool that timed out.
A repair is a full round trip: the whole context goes out again so the model can see the error. At turn fifteen that is fifteen turns of history re-sent to fix a missing argument. The fixes are boring and worth more than they look: validate calls against the schema locally, hand the model the exact error, and keep tool results small.
An agent with no exit condition will try forever
When a tool keeps failing, nothing in the model stops it from calling the same tool again. Each turn it sees a failure and the obvious next action, which is to try again. Something outside the model has to notice that no progress is being made.
A loop guard is a counter and a rule: after N turns with no new progress, stop and hand back what you have. It is unglamorous, and it is the difference between a failed task and a failed task that also cost fifty dollars. How those turns are arranged across several agents is the subject of the next lab.
The short version
- Each turn re-sends the whole history, so cost grows with the square of the turns.
- When the window fills, dropped results must be redone; offloading keeps them retrievable.
- Malformed calls and tool errors each cost a full turn; validate locally and keep results small.
- Only the harness can stop a stuck agent: add a loop guard.