The agent loop

Think, call a tool, observe, repeat. Watch an agent work a task turn by turn while its context window fills, then choose what happens when it runs out - truncate, compact, or retrieve - and watch the agent forget a result it still needed.

12 min read

An agent is not a model. It is a loop around a model: a harness that decides what goes into the next prompt, runs whatever tool the model asks for, and decides when to stop. This article adds one failure at a time. Each figure runs an agent to the end with your settings; scrub back through it, change a setting, or run it again.

Every turn re-sends everything that happened before it

The loop is think, call a tool, run it, read the result, repeat. The model is stateless, so to take turn ten the harness sends turns one through nine again, plus the system prompt and every tool definition. Nothing is broken yet, and it is already expensive.

THINKTOOL CALLEXECUTEOBSERVE6/6 doneturn 7
Prompt size sent on each turn
prompt size billed on each turn · peak 4,723 tok
Turns
7
Last prompt
4.7k tokens
Growth since turn 1
5.2x
Tokens billed
15.9k
6 steps
turn 7, the end
Every turn re-sends everything before it, so each bar is taller than the last. Double the task length and the bill roughly quadruples.

Because the prompt grows every turn, the total billed is the sum of a growing series: it scales with the square of the turn count. Doubling the turns roughly quadruples the bill. Prompt caching softens the constant, not the shape. Parallel tool calls finish the same task in fewer turns, which makes them the cheapest optimisation a harness has.

Sooner or later it does not fit, and something has to go

Tool results are the biggest things in the transcript, and they never stop arriving. When the context window fills, the harness has to choose what the agent stops knowing.

Context window
5,565 / 6,000 tok
system + tools history
Findings still in context: teal held, red done and then lost
12345678910
Transcript (faded messages have been evicted)
thinkevicted
I am missing what step 5 returned. I will have to get it again.
tool callevicted
run_tests()
observationevicted
re-ran step 5 from scratch
thinkevicted
I am missing what step 7 returned. I will have to get it again.
tool callevicted
run_tests()
observationevicted
re-ran step 7 from scratch
thinkevicted
I am missing what step 6 returned. I will have to get it again.
tool callevicted
run_tests()
observationevicted
re-ran step 6 from scratch
thinkevicted
I am missing what step 1 returned. I will have to get it again.
tool callevicted
run_tests()
observationevicted
re-ran step 1 from scratch
think120 tok
I am missing what step 8 returned. I will have to get it again.
tool call70 tok
run_tests()
observation431 tok
re-ran step 2 from scratch
think120 tok
I am missing what step 8 returned. I will have to get it again.
tool call70 tok
run_tests()
observation520 tok
re-ran step 3 from scratch
think120 tok
I am missing what step 8 returned. I will have to get it again.
tool call70 tok
run_tests()
observation542 tok
re-ran step 8 from scratch
think120 tok
I am missing what step 4 returned. I will have to get it again.
tool call70 tok
run_tests()
observation718 tok
re-ran step 4 from scratch
think120 tok
I am missing what step 5 returned. I will have to get it again.
tool call70 tok
run_tests()
observation662 tok
re-ran step 5 from scratch
think120 tok
I am missing what step 7 returned. I will have to get it again.
tool call70 tok
run_tests()
observation652 tok
re-ran step 6 from scratch
Steps redone
92
Turns used
101
Tokens billed
553.5k
Window
6k tokens
Still unfinished after 101 turns: the agent keeps redoing work the window threw away.
6k tokens
turn 101, the end
Dropping the oldest messages drops findings, and the agent pays to redo them. Try summarising, then offloading, with the same window.

The context is the memory. There is nowhere else a conclusion was written down, so dropping a message drops the finding, and the agent goes and gets it again, paying for the tool call, the tokens and the turn. Dropping the oldest is cheapest and most brutal. Summarising keeps a low-fidelity trace of everything, which is how a compacted agent ends up confidently half-remembering a number. Offloading results to a store and retrieving them on demand costs a turn per lookup, but nothing is ever truly gone.

The call will not parse, and the tool times out

Real harnesses spend a surprising share of their turns neither thinking nor working: they reject a malformed tool call, feed the error back, and ask again, or they retry a tool that timed out.

Where each turn went
progressrepairtool errorcompaction
Transcript
think120 tok
Step 1 of the task: I need read_file.
tool call70 tok
read_file()
tool error60 tok
read_file failed: timeout after 30s. Retrying next turn.
think120 tok
Step 1 of the task: I need read_file.
tool call70 tok
read_file()
tool error60 tok
read_file failed: timeout after 30s. Retrying next turn.
think120 tok
Step 1 of the task: I need read_file.
tool call70 tok
read_file()
observation464 tok
read_file returned 21 matches
think120 tok
Step 2 of the task: I need grep.
tool call70 tok
grep()
observation423 tok
grep returned 10 matches
think120 tok
Step 3 of the task: I need run_tests.
tool call70 tok
run_tests()
observation837 tok
run_tests returned 34 matches
think120 tok
Step 4 of the task: I need web_search.
tool call70 tok
web_search()
observation730 tok
web_search returned 13 matches
think120 tok
Step 5 of the task: I need edit_file.
tool call70 tok
edit_file()
observation386 tok
edit_file returned 24 matches
think120 tok
Step 6 of the task: I need list_dir.
tool call70 tok
list_dir()
observation780 tok
list_dir returned 39 matches
final200 tok
Task complete. Answer assembled from all 6 findings.
Repairs
0
Tool errors
2
Turns wasted
25%
Tokens billed
21.7k
25%
20%
turn 9, the end
A quarter of tool calls failing and a fifth malformed: a large share of turns make no progress, and each one re-sends the whole context.

A repair is a full round trip: the whole context goes out again so the model can see the error. At turn fifteen that is fifteen turns of history re-sent to fix a missing argument. The fixes are boring and worth more than they look: validate calls against the schema locally, hand the model the exact error, and keep tool results small.

An agent with no exit condition will try forever

When a tool keeps failing, nothing in the model stops it from calling the same tool again. Each turn it sees a failure and the obvious next action, which is to try again. Something outside the model has to notice that no progress is being made.

THINKTOOL CALLEXECUTEOBSERVE0/6 done41 turns, no progress
Turn outcomes
Turns
42
Turns without progress
41
Tokens billed
247.9k
Status
aborted
No guard, no progress, out of turns. This is the failure mode a loop guard exists to catch.
95%
turn 42, the end
A tool that fails 95% of the time, and no guard: the agent keeps calling it until it runs out of turns. Turn the guard on and it stops after four turns of no progress.

A loop guard is a counter and a rule: after N turns with no new progress, stop and hand back what you have. It is unglamorous, and it is the difference between a failed task and a failed task that also cost fifty dollars. How those turns are arranged across several agents is the subject of the next lab.

The short version

  • Each turn re-sends the whole history, so cost grows with the square of the turns.
  • When the window fills, dropped results must be redone; offloading keeps them retrievable.
  • Malformed calls and tool errors each cost a full turn; validate locally and keep results small.
  • Only the harness can stop a stuck agent: add a loop guard.