Each chart on this page shows one dot per run, grouped by model and reasoning effort. Outlined dots are wave-1 runs. Filled dots are the replicate runs from waves 2 and 3. A short vertical tick marks the average of a cell with more than one run.
This page covers 34 finished runs. Seven wave-3 runs are still running: two of Opus 5.5 high, two of Fable 5.1 high and three of Fable 5.1 max. They are not shown yet.
17.1One idea decides most runs
A run's first big jump almost always comes from the same rule. When the run is ahead, it closes the opponent's live scoring target before opening a new one. It is worth 6 to 9 points of pool score in one step. We mark the first iteration at which a run benched a clean candidate at 60% or more, which only that idea reaches.
When each run found the key idea
Iteration of the first clean candidate at 60% or more. Dots at "never" did not find it within the budget.
- Fable
- Opus
- Sonnet 5
- GPT-6 Astra
- GPT-5.6 Luna
Two thirds of the finished runs reached 60% within 25 iterations. Five never reached it, and they are among the weakest runs in the benchmark: GPT-6 Astra medium in wave 1, Sonnet 5 high-a, both Luna high-a and high-b, and the 13-iteration older Opus run. Two more runs, Sonnet 5 high-b and Luna high-c, crossed 60% and stalled just above it. Reaching 60% is necessary for a strong run, but not enough.
A single run can mislead. Astra medium's first run never found the idea, and its two replicates found it at iterations 20 and 69 and scored normally. Sonnet 5's first high run never found it, and its third high run reached 73.3%. That is why every cell has replicate runs.
17.2Strength, run by run
Head-to-head Elo of each run's champion
One dot per run. Elo from the ladder, S1 = 1000.
- Fable
- Opus
- Sonnet 5
- GPT-6 Astra
- GPT-5.6 Luna
Runs of the same model at the same effort differ by 4 to 182 Elo. Opus 5.5 medium is the most repeatable cell, with two runs 4 Elo apart. Most other cells spread 10 to 50 Elo. That spread is as large as most gaps between models, so a single run can't rank two models.
The best champions of Fable 5.1 and Opus 5.5 sit within a few Elo of each other at the top. The best GPT-6 Astra runs sit about 10 Elo lower. Sonnet 5 and GPT-5.6 Luna sit further back.
17.3Reasoning effort
Effort is pinned per run, so each model's rows compare efforts directly. With the runs finished so far:
| Model | Effort | Runs | Elo of each run | Average |
|---|---|---|---|---|
| Sonnet 5 | medium | 2 | 1095, 1124 | 1110 |
| Sonnet 5 | high | 3 | 921, 1060, 1101 | 1028 |
| GPT-5.6 Luna | medium | 2 | 1049, 1104 | 1076 |
| GPT-5.6 Luna | high | 3 | 928, 928, 1102 | 986 |
| GPT-6 Astra | low / medium / high / xhigh | 3 each | 960–1143 | 1124 / 1074 / 1126 / 1130 |
| Opus 5.5 | medium / high | 2 / 1 | 1150, 1154 / 1132 | 1152 / 1132 |
| Fable 5.1 | medium / high / max | 2 / 1 / 1 | 1118, 1144 / 1157 / 1156 | 1131 / 1157 / 1156 |
- Sonnet 5 and Luna average lower at high effort. The difference comes from runs that missed the key idea: one of three Sonnet high runs and two of three Luna high runs. Each model's best high-effort run matches its medium runs. So high effort made these models miss the idea more often. It did not lower the ceiling.
- GPT-6 Astra is flat from low to xhigh. Its medium average is lower only because of the one wave-1 run that never found the idea.
- Fable 5.1 gains from medium to high and levels off at max. High and max each have one finished run so far.
- Opus 5.5 high reached the best pool score of any run, but rates about 20 Elo below Opus 5.5 medium head to head. Its champion remembers each opponent's habits within a game, which fits the fixed pool.
Effort does change how a run works. At higher effort, the Claude runs wrote stricter keep rules for themselves and discarded more of their own positive results. Chapter 18.2 explains why the loop's own rules stop filtering after the first jump.
17.4What effort costs
Output tokens per iteration
Total output tokens, including thinking and reasoning, divided by benched candidates. Log scale.
- Fable
- Opus
- Sonnet 5
- GPT-6 Astra
- GPT-5.6 Luna
Within every model, tokens rise with effort. Astra's cost climbs step by step from low to xhigh. Sonnet 5 high used about 1.5 times the tokens of medium. Across engines, Codex runs spend 2 to 5 times fewer output tokens than Claude runs on the same loop, although the two engines count tokens differently.
17.5The bull bug in practice
Some runs noticed that a triple aimed at the bull skipped its difficulty penalty (chapter 19.4). Some avoided it, and some used it. Every champion that used it plays almost as well with the bug patched: the patch costs 1.5 to 3 points of pool score. In wave 2, only Opus 5.5 medium-b used it. Its X222 fell from 77.9% raw to 76.3% patched, level with its clean sibling X230 at 76.5%.
17.6Caveats
- Few runs per cell. One to three runs per model and effort. Differences under about 50 Elo between cells are within run-to-run spread.
- Two harnesses in wave 1. Wave-1 Astra runs used a proposal-only driver. From wave 2 on, every run uses the same full agent loop.
- One seed, one pool. A run selects its champion on the fixed pool at one seed. The ladder tests the champion against other strong players.
- Shared machines. Runs shared a Mac or a VM. The prompt's 180-second hang rule keeps slow benches from being discarded (chapter 18.3).