Agent graphs

Wire a multi-agent graph and run it. Fan out workers, add a critic loop, flip parallel execution on and off, and watch the Gantt chart redraw as the critical path - not the total work - decides how long the run takes.

12 min read

Once you have more than one agent, the interesting question stops being “how good is the model?” and becomes “what shape is the graph?”. Total work decides your bill. The critical path, the longest chain of steps that must happen one after another, decides how long the user waits. This article pulls those two numbers apart one change at a time.

A single agent pays for every subtask in series

One agent, one task list, one thing happening at a time. The wall clock and the total work are the same number, and the user waits for all of it.

single agent4 subtasks, one after another
Timeline
single agent
12.1s
Wall clock (user waits)
12.3 s
Agent time (you pay)
12.1 s
Estimated cost
$0.044
Tokens
15k
4
12.5 s, the end
One agent: the wall clock and the bill are the same number. Add a subtask and both grow together.

With one agent there is no difference between work done and time elapsed; they are one quantity measured in two units. The only lever is doing less work. Every idea below is about breaking that link: moving the wall clock without moving the bill, or the other way round.

Split the work and the wall clock stops tracking the bill

A planner splits the task, workers each take a piece, and a reducer merges what comes back. If the workers run at the same time, the user waits for the longest chain, not the sum.

plannersplit the taskworker 1subtask 1worker 2subtask 2worker 3subtask 3worker 4subtask 4reducermerge findings
Timeline: the critical path is orange
planner
2.4s
worker 1
2.6s
worker 2
3.7s
worker 3
4.9s
worker 4
3.2s
reducer
2.1s
Wall clock (user waits)
9.8 s
Agent time (you pay)
18.9 s
Speed-up
1.92x
Estimated cost
$0.068
4
10.0 s, the end
Fanned out, the user waits for the planner, the slowest worker and the reducer. You still pay for every worker, plus the planner and reducer.

Parallelism buys latency, never cost. You still pay for every subtask, and now also for a planner and a reducer, so total agent time goes up a little while the wall clock falls a lot. Turn parallel execution off and the graph looks identical but runs as long as before: a chain wearing a graph costume, paying the coordination overhead for none of the benefit.

Your rate limit decides how wide the graph really is

You can declare seven workers. Whether seven run at once is up to your provider's rate limit and your connection pool. Past the limit, extra workers queue and start in later waves.

Timeline: workers past the limit queue into a staircase
planner
2.4s
worker 1
2.6s
worker 2
3.7s
worker 3
4.9s
worker 4
3.2s
worker 5
4.3s
worker 6
2.6s
worker 7
3.7s
reducer
2.1s
Wall clock (user waits)
19.1 s
Agent time (you pay)
29.6 s
Speed-up
1.55x
Estimated cost
$0.106
5 workers cannot start until a slot frees, so they wait in line on the critical path.
7
2
19.3 s, the end
Seven workers, two allowed at once: a staircase. Raise the limit one notch at a time and watch the wall clock fall, then stop improving.

Every wave lands directly on the critical path, so the timeline turns from a block into a staircase. A worker added beyond the limit is billed in full and shortens nothing, which makes over-applied fan-out the easiest way to build an agent graph that is worse than the loop it replaced.

A quality loop with no cap is an open-ended bill

Add a critic that scores the output and sends it back for revision, and the graph stops being a one-way flow: it can loop. Each revision adds a full pass to the clock and closes only part of the gap to the quality bar.

plannersplit the taskworker 1subtask 1worker 2subtask 2worker 3subtask 3reducermerge findingscriticscore the outputrevision 1apply critic notescriticscore the outputrevision 2apply critic notescriticscore the outputrevision 3apply critic notescriticscore the output
Quality against the bar of 90
93 out of 100 after 3 revisions
Wall clock
26.1 s
Revisions
3 of 3
Quality
93
Estimated cost
$0.107
90
3
26.3 s, the end
Each revision costs a full pass and closes about half the remaining gap. A bar of 90 takes three rounds; the cap is what keeps the bill finite.

Returns halve while the price per pass stays flat: a first revision might take output from 60 to 80, the second to about 90. A bar of 95 can cost several passes and still not be reached. That is why the maximum number of rounds is not a tuning knob; it is the thing standing between you and an unbounded bill.

Wall clock follows the critical path. Cost follows total agent time. Quality follows the critic loop, the one part of the graph that can run away from you.

The short version

  • One agent: wall clock and cost are the same number.
  • Fan-out cuts the wall clock to the critical path and slightly raises the cost.
  • Past your concurrency limit, extra workers only add cost.
  • Critic loops give diminishing returns at full price; always cap the rounds.