Cricket Bench

Strategy writing: what the runs found

Nearly every strong run found the same idea: when ahead, shut the opponent's scoring lane. Whether and when a run found it explains more of its result than its model or its reasoning effort. Replicate runs show how much of any single score is luck.

Runs strategy writing, waves 1–3Budget 115 benched candidates per runScores bull-patched pool score and ladder Elo

Each chart on this page shows one dot per run, grouped by model and reasoning effort. Outlined dots are wave-1 runs. Filled dots are the replicate runs from waves 2 and 3. A short vertical tick marks the average of a cell with more than one run.

This page covers 34 finished runs. Seven wave-3 runs are still running: two of Opus 5.5 high, two of Fable 5.1 high and three of Fable 5.1 max. They are not shown yet.

17.1One idea decides most runs

A run's first big jump almost always comes from the same rule. When the run is ahead, it closes the opponent's live scoring target before opening a new one. It is worth 6 to 9 points of pool score in one step. We mark the first iteration at which a run benched a clean candidate at 60% or more, which only that idea reaches.

When each run found the key idea

Iteration of the first clean candidate at 60% or more. Dots at "never" did not find it within the budget.

  • Fable
  • Opus
  • Sonnet 5
  • GPT-6 Astra
  • GPT-5.6 Luna
Reading it. The toggle switches the same rows to ladder Elo, pool score or tokens per iteration. Source: data/csv/runs_all.csv, curves_all.csv.

Two thirds of the finished runs reached 60% within 25 iterations. Five never reached it, and they are among the weakest runs in the benchmark: GPT-6 Astra medium in wave 1, Sonnet 5 high-a, both Luna high-a and high-b, and the 13-iteration older Opus run. Two more runs, Sonnet 5 high-b and Luna high-c, crossed 60% and stalled just above it. Reaching 60% is necessary for a strong run, but not enough.

A single run can mislead. Astra medium's first run never found the idea, and its two replicates found it at iterations 20 and 69 and scored normally. Sonnet 5's first high run never found it, and its third high run reached 73.3%. That is why every cell has replicate runs.

17.2Strength, run by run

Head-to-head Elo of each run's champion

One dot per run. Elo from the ladder, S1 = 1000.

  • Fable
  • Opus
  • Sonnet 5
  • GPT-6 Astra
  • GPT-5.6 Luna
Reading it. A 10-Elo gap is about a 51.4% head-to-head edge. Runs still in progress have no Elo yet. Source: ladder/results/ratings.csv.

Runs of the same model at the same effort differ by 4 to 182 Elo. Opus 5.5 medium is the most repeatable cell, with two runs 4 Elo apart. Most other cells spread 10 to 50 Elo. That spread is as large as most gaps between models, so a single run can't rank two models.

The best champions of Fable 5.1 and Opus 5.5 sit within a few Elo of each other at the top. The best GPT-6 Astra runs sit about 10 Elo lower. Sonnet 5 and GPT-5.6 Luna sit further back.

17.3Reasoning effort

Effort is pinned per run, so each model's rows compare efforts directly. With the runs finished so far:

Average Elo per model and effort · finished runs only · data/csv/runs_all.csv
ModelEffortRunsElo of each runAverage
Sonnet 5medium21095, 11241110
Sonnet 5high3921, 1060, 11011028
GPT-5.6 Lunamedium21049, 11041076
GPT-5.6 Lunahigh3928, 928, 1102986
GPT-6 Astralow / medium / high / xhigh3 each960–11431124 / 1074 / 1126 / 1130
Opus 5.5medium / high2 / 11150, 1154 / 11321152 / 1132
Fable 5.1medium / high / max2 / 1 / 11118, 1144 / 1157 / 11561131 / 1157 / 1156

Effort does change how a run works. At higher effort, the Claude runs wrote stricter keep rules for themselves and discarded more of their own positive results. Chapter 18.2 explains why the loop's own rules stop filtering after the first jump.

17.4What effort costs

Output tokens per iteration

Total output tokens, including thinking and reasoning, divided by benched candidates. Log scale.

  • Fable
  • Opus
  • Sonnet 5
  • GPT-6 Astra
  • GPT-5.6 Luna
Not shown. Wave-1 Astra runs, because a driver did the tooling and the model only wrote proposals. Early runs with no surviving transcript. Source: data/csv/tokens.csv.

Within every model, tokens rise with effort. Astra's cost climbs step by step from low to xhigh. Sonnet 5 high used about 1.5 times the tokens of medium. Across engines, Codex runs spend 2 to 5 times fewer output tokens than Claude runs on the same loop, although the two engines count tokens differently.

17.5The bull bug in practice

Some runs noticed that a triple aimed at the bull skipped its difficulty penalty (chapter 19.4). Some avoided it, and some used it. Every champion that used it plays almost as well with the bug patched: the patch costs 1.5 to 3 points of pool score. In wave 2, only Opus 5.5 medium-b used it. Its X222 fell from 77.9% raw to 76.3% patched, level with its clean sibling X230 at 76.5%.

17.6Caveats