07 August 2026
This week we built and ran our first controlled comparison of a solo coding agent and a coordinated swarm.
Initial methodology
We designed the experiment to compare two complete operating strategies:
- Solo: one frontier model plans, implements, and verifies the task.
- Swarm: the same frontier model plans while up to four cheaper worker models implement, review, and resolve conflicts.
Both conditions receive the same task, repository, public tests, hidden-test isolation, environment, tools, permissions, and starting state. We measure hidden accuracy, time to working code, final active AI time, model cost, and human intervention. Runs are unattended, paired by a predeclared trial seed, and evaluated without exposing hidden feedback to the agents.
Our setup
The fixture was MiniFlow, a Python 3.13 workflow engine using only the standard library. Agents had to implement validation, deterministic concurrent scheduling, retries, timeouts, cancellation, checkpoint and resume behavior, events, and a command-line interface.
The fixture contained 25 public tests. A host-only evaluator scored 198 hidden cases across 11 equally weighted categories and 45 specification requirements. The fixture, specification, evaluator, harness, prompts, and environment were content-addressed and frozen before launch.
The runs used:
- Solo:
gpt-5.6-sol, high reasoning effort. - Swarm planner:
gpt-5.6-sol, high reasoning effort. - Swarm workers:
gpt-5.6-terra, high reasoning effort, with four worker slots. - Environment: isolated Linux ARM64 container, two CPUs, 3 GB memory, no network, and gateway-only repository access.
- Timing: a 20-minute operating target and a 60-minute active-AI emergency ceiling for each condition.
Results
Both conditions completed unattended with valid telemetry and reconciled usage.
| Measure | Solo | Swarm |
|---|---|---|
| Final hidden score | 192/198 (97.0%) | 198/198 (100.0%) |
| First successful build | 6m 39s | 4m 12s |
| Full public-test pass | 8m 38s | 8m 37s |
| Active-AI final result | 8m 48s | 8m 55s |
| Model calls | 1 | 9 |
| Imputed list-price cost | $0.055 | $0.374 |
| Human interventions | 0 | 0 |
The swarm passed every hidden case. Solo missed six checkpoint and resume cases and passed every other hidden category. The accuracy difference was 3.0 percentage points; swarm cost about 6.8 times as much.
The implementations were structurally different. Solo produced 1,016 lines of Python across its API, CLI, and engine, with 727 lines concentrated in the engine. Swarm produced 1,867 lines across focused API, CLI, engine, runtime, and validation modules.
All four swarm workers started concurrently with non-overlapping ownership. Three partial integrations built successfully but could not pass the public suite until the engine and API arrived. The fourth integration completed the system and passed the pinned checks. The planner then stopped without launching a reviewer or resolver.
Surprises
- The entire 3.0-point accuracy difference came from one behavior. When resuming a cancelled task with a scheduled retry, solo cleared the recorded delay and retried immediately. All six misses were variants of that case; every other hidden case passed.
- Swarm produced 84% more Python than solo, but split it into focused modules. Its largest module was 570 lines, compared with solo’s 727-line engine.
- Parallel work did not create an integration tail. Four workers committed four independently owned slices, and the final tree needed no review or resolution task.
- Three workers correctly handed back buildable but publicly failing partial systems. The planner accepted those intermediate states and waited for the missing engine and API instead of treating each failure as a new defect.
- The swarm produced its first successful build 2m 27s earlier, but the two conditions reached a full public pass only one second apart. Parallelism accelerated the first buildable structure, not final completion.
- At five minutes, hidden snapshots were 0/198 for solo and 6/198 for swarm. Both then reached their final scores before nine minutes, showing a sharp integration step rather than gradual progress.
- Solo left its implementation as working-tree changes. Swarm’s integration protocol produced four focused commits, giving it a clearer implementation history as a side effect of coordination.
Update — 8 August 2026: three predeclared pairs
The result above came from a single paired run, and we said at the time that one pair was evidence the pipeline worked, not a sufficient sample. We then revised the protocol (3.0.0: a hardened harness, a symmetric prompt set, and model-per-role manifests that let us separate topology from model economics), re-ran calibration, and executed three fresh predeclared pairs — six unattended runs, all with valid telemetry and reconciled cost.
The single-pair headline did not replicate. The calibration pilot reversed it (solo 198/198, swarm 192/198), and the matrix scattered:
| Pair | Solo | Swarm | Swarm − solo |
|---|---|---|---|
| 1 | 88/198 (44.4%) | 192/198 (97.0%) | +52.5 pp |
| 2 | 88/198 (44.4%) | 94/198 (47.5%) | +3.0 pp |
| 3 | 198/198 (100.0%) | 198/198 (100.0%) | 0.0 pp |
Medians: solo 44.4% vs swarm 97.0% accuracy; solo $0.045 vs swarm $0.297 imputed cost (6.6x); solo 10m 30s vs swarm 13m 30s active AI time; both conditions 3 of 3 unattended completions.
What we learned:
- Outcomes are bimodal. Every run finished at either 97–100% or 44–48%, never in between, and the three collapsed runs (two solo, one swarm) share an almost identical failure signature: a coherent slice of the event and observability contract skipped wholesale while the public tests still pass. Single-run comparisons on this fixture are draws from that distribution — including the one we published above.
- The swarm’s advantage here is robustness, not quality at the top. When both conditions succeeded they tied at 100%, and the swarm cost 5.4x the tokens and ran slower. The swarm avoided the collapse mode in two of the three seeds where it appeared; with three pairs, that observation is worth exactly one-versus-two runs of evidence.
- Our predeclared decision gate passes, with an honest asterisk. Accuracy and autonomy pass; the cost criterion fails decisively (6.6x); the speed criterion passes only because solo never reached 80% of hidden checks in two of three runs. We are publishing the gate outcome either way, as the protocol now requires.
Full per-run tables, the gate evaluation, and limitations are in the repository report for this matrix, with every run’s raw and evaluated bundle content-addressed.