A cheaper delegated build that failed verification

A June 2026 orchestration experiment cut repeated-context cost, but its first crew result did not compile. Here is what the run supports and what it does not.

By Faruk Alpay · Published · Updated · 7 min read

  • agent orchestration
  • benchmarks
  • verification

In June 2026 we compared a long single-agent build with an operator-and-worker design. The recorded crew run cost about $0.032. The single-agent run cost about $0.86, a ratio close to 27 to 1.

The crew run did not compile on its first measured pass.

That fact changes the headline. The experiment shows a large difference in recorded token cost and exposes where orchestration time went. It does not show a 27-times cheaper completed build, because the outcomes were not equivalent.

The architecture under test

A long agent loop repeatedly sends a growing history back to the model. Tool calls, file contents, command output, and prior reasoning accumulate. In the measured solo run, 29 model steps consumed 4,371,865 input tokens and 78,029 output tokens.

The orchestrated path bounds that growth by separating planning from execution:

task
  -> operator plan and file ownership
  -> dependency waves
  -> isolated worker contexts
  -> build, test, and run
  -> review of actual command output
  -> bounded repair waves

Each worker receives the plan, a file manifest, and the files owned by its batch. The context bundle is capped. It does not inherit the operator's entire conversation.

The current scheduler in server/src/orchestrator/spawn.js groups independent batches into dependency waves. Workers inside a wave run concurrently under a bounded pool. The default concurrency is four and the configured maximum is eight. A later wave starts only after the current wave settles.

Within one multi-file batch, micro-batching may generate files sequentially so each file can see the files already written by that batch. Parallelism happens across independent batches, not across every file indiscriminately.

This corrects the original description of the system, which said workers were spawned one at a time.

What the first experiment recorded

The task was a standard-library Rust benchmark package with several algorithms, a timing harness, an SVG figure, CSV output, integration tests, and real build and run commands. The three approaches were executed on the same task and machine, according to the run notes.

ApproachWall timeRecorded costObserved outcome
Fast single-agent loopabout 228 s$0.86Compiled; four tests passed; every timing printed 0 ms
Deep single agentmore than 9 minnot completedNo file was produced before the run was stopped
Operator plus workersabout 207 s$0.032Review found a Rust string-delimiter error; first measured pass did not compile

The crew trace attributed about 174 seconds to planning and review, and roughly 22 seconds to seven worker spawns. Those numbers support a latency diagnosis: worker execution was inexpensive, while the operator remained the serial bottleneck.

They do not support an outcome-equivalent cost comparison. The solo result passed its tests but produced meaningless timing data. The crew result reached review quickly but had a compile error. Both outputs failed an important acceptance criterion, in different ways.

Why the token count fell

The strongest architectural explanation is context isolation. In the solo loop, every additional step replays a larger prefix. Total billed input grows rapidly even if each new action is small.

Workers start from compact contexts. If five independent batches each need a manifest and their owned files, they do not each need the full transcript of the other four batches. This bounds repeated input and makes concurrency possible.

That is an explanation from the implementation and token trace, not a causal estimate from the benchmark. The experiment changed several variables at once: model role, context shape, scheduling, and review policy. One run cannot assign a percentage of the savings to each mechanism.

Verification is part of the result

The original presentation treated the cost ratio as the main result and the compile failure as a caveat. The order should be reversed. A build benchmark first needs a common acceptance test, then cost and latency can be compared among runs that pass it.

For this task, a defensible acceptance contract would include:

  1. cargo build exits successfully.
  2. cargo test passes the required tests.
  3. cargo run produces the declared CSV and SVG artifacts.
  4. Timing values have sufficient resolution and are not all zero.
  5. Every expected file exists and contains non-empty content.

Only runs that pass all five should enter a completed-build cost comparison. Failed runs remain valuable for diagnosing architecture, but they belong in separate outcome classes.

The production orchestrator follows this order. Workers write their bounded file sets, structural lint can reject malformed artifacts, the runtime executes real verification commands, and review reads those results. Failed verification can create bounded fix batches that use the deeper escalation tier.

What we can claim from this run

The evidence supports four limited conclusions:

  • A long solo loop can become input-heavy because it repeatedly sends an expanding context.
  • Isolated worker contexts recorded far lower input cost in this experiment.
  • Independent dependency batches can run concurrently without sharing the full transcript.
  • Planning and review dominated the measured crew latency.

The evidence does not establish these broader claims:

  • that orchestration is 27 times cheaper for completed builds;
  • that the result generalizes across repositories or task sizes;
  • that the observed wall-time difference is statistically stable;
  • that worker speed is more important than operator quality.

The run notes do not include enough information for an external reproduction. They lack a pinned repository commit, container image, hardware description, repeated trials, and a variance estimate. That is a measurement gap.

The next benchmark protocol

A replacement benchmark should pin the task fixture, source commit, container digest, provider settings, and model configuration. It should run each approach several times in randomized order and publish every terminal outcome, including timeouts and invalid artifacts.

The report should separate:

  • operator input and output tokens;
  • worker input and output tokens;
  • tool and build time;
  • time to first file;
  • time to the first verified green result;
  • repair rounds;
  • pass rate for every acceptance check.

Median and tail latency matter, but pass rate comes first. A cheap failed build is not a cheaper success.

The architectural direction remains useful. Bounded contexts, explicit file ownership, dependency-aware concurrency, and real verification address concrete failure modes in long agent loops. The first experiment helped identify those mechanisms. Its cost ratio should remain in the record with the failed outcome attached to it.

All dev blogs · RSS feed

Lightcap