Cricket Bench

Methods and data

Every run, every deviation from protocol, how the numbers were checked, and the raw files.

Sources read-only run repos, journals, research logs, session transcriptsRebench every champion re-run raw and bull-patchedLadder every player gated before rating

A.1Every run

One row per run, in both benchmarks and all three waves. Pool scores are bull-patched. Pool scores from the two benchmarks use different pools, so compare them within a benchmark only.

All runs

A.2Download the data

A.3How the numbers were checked

A.4Deviations from protocol

Every deviation below is also noted where it matters. Together they are the main reason to read cross-model comparisons as descriptive.

Known deviations · runs/README.md, research/02-heuristic-arms.md, research/03-rl-arms.md
RunDeviationEffect
GPT-6 Astra, wave 1, strategy writingRun by a driver script. The model proposed each strategy, and the script built, benched and kept it.Measures proposal quality, not agentic work. Token counts are not comparable with full agent sessions. Waves 2 and 3 run Astra as a full agent.
GPT-6 Astra, wave 1, self-trainingA coordinator framed the deliverable as "a reproducible environment first".A different objective. All three efforts stopped far inside the time budget.
Claude Fable 5, wave 1After 25 iterations the run expanded its own opponent pool, seven times, to 19 bots.Only iterations 1 to 25 are comparable.
Claude Opus 5, wave 1Counted multi-configuration sweeps as single iterations.Only 106 of the 115 claimed iterations can be identified.
Claude Opus 5.5 high, wave 1The runner counted only keep and discard commits. The run also gave itself within-game opponent memory.125 strategies for a reported 115. Its champion at 115 is the same, X189.
Claude Opus 5.5 medium, wave 1A session hit a usage limit, and the runner started another full session.129 iterations. Its champion at 115 is X212 (74.6%). Its final champion X227 scores 76.6%.
Claude Fable 5.1 max, wave 1Stopped at 56 iterations after the run declared its rule family exhausted.The fewest iterations of the modern runs. Wave 3 adds three full-length max runs.
All wave-2 strategy runsThe loop's 10-second hang rule discarded good candidates on the shared machines. The prompt now sets 180 seconds, and all 16 runs restarted from scratch at 16:15 ET on 24 September.The aborted attempts are kept outside the runs as git bundles, tagged hang10s.
Claude Opus 5.5 medium-b, wave 216 candidates were committed as "noted", which the runner did not count.131 benched candidates. Its champion at 115 is X209 (76.4% patched).
GPT-5.6 Luna medium-b, wave 2One session scripted a batch of 14 candidates. All failed to compile at the same line, and the script did not stop.14 of 115 iterations produced nothing. Kept as a real result: loop.md counts a compile error as a discard.
Codex strategy runs, wave 2The Codex account hit its weekly usage limit at 17:37 ET on 24 September.The runners waited. After a reset, work resumed at 17:48 ET. No iteration was lost.
GPT-5.6 Luna medium, wave 2, self-trainingThe first session stopped to ask for design approval, because Codex loaded a global skill that requires sign-off.The continue prompt says no human is available. Only runs that needed a second session saw that line.
Wave-2 self-training runsFour runs trained on the Mac at once, alongside five strategy runs and the ladder. Threads were capped at 4 per run.The training budget is wall-clock time, so shared compute may have reduced effective training.

A.5Known gaps

A.6Glossary

Run
One complete attempt by one model, at one effort, on one benchmark, from the pinned starting code. Older notes call it an arm.
Wave
A batch of runs launched together. Wave 1 ran from July to September. Wave 2 started on 24 September and wave 3 on 27 September.
Pool
The fixed bots a run is scored against. Each benchmark has its own pool (chapter 19.3).
Pool score
Mean win rate against the pool: per-opponent win rate, averaged over skill profiles, then over opponents.
Iteration
One strategy written and benched. Kept, discarded and noted candidates all count.
Discovery
The first iteration at which a run benched a clean candidate at 60% or more. It marks when the run found the lane-shutdown idea.
Clean champion
The best strategy a run kept that never aims a triple at the bull.
Exploit champion
A champion that aims a triple at the bull. "Patched" means re-scored with the bull bug fixed.
Elo
A Bradley–Terry rating from the head-to-head ladder, anchored at S1 = 1000. A 10-point gap is about a 51.4% edge.
IT3
Fable 5's league-trained self-training agent from July, and the ladder leader.

A.7History

The benchmark grew out of a seven-month project that started as an attempt to teach a computer to play bar darts. The history page tells that story.