This benchmark asks a model to do open-ended research with a fast, honest scoreboard. Each iteration is one hypothesis: write a strategy, test it on 825,000 games, decide whether it earned its place. The deliverable is the best strategy the run finds. What we watch is how the model gets there.
18.1The task
The scaffold is commit 7b098ec of the Darts-Cricket repo. It holds the engine, the bots, the benchmark script and two documents: loop.md, the rules of the loop, and strategies_journal.md, the run's lab notebook. The loop is:
- Read the journal end to end: the baseline, what has been tried, and the open territory.
- Pick one structural idea. Do not repeat a discarded idea without a real change of mechanism.
- Write it as a new
if (strat_id == N)branch inchoose_throw()infast_sim.c, and register its name. - Bench it with
./autoresearch_strategies/bench.sh NAME. The script recompiles the engine and plays the full pool. - Decide keep or discard under the keep rules (18.2).
- Journal first. Write the entry before reverting anything, because the journal is the only record of a discard.
- Commit as
autoresearch: keep NAMEorautoresearch: discard NAME, then loop. Never ask to continue.
The goal is to beat E12, the best classic bot, at 54.6%. A run targets 115 iterations.
Source: autoresearch_strategies/loop.md at 7b098ec.
18.2The score and the keep rules
bench.sh prints five metrics. The headline is the mean win rate: for each opponent, average the win rate over the three skill profiles, then average over the eleven opponents. It also reports the worst opponent's average, the best, and how many opponents the candidate beats. At 25,000 games per matchup the standard error is about 0.3 points per matchup. The scaffold journal says it plainly: "1 pp is real, 0.3 pp is noise."
A candidate is kept if any one of these holds:
- its mean is at least 55.1%, which beats E12 by half a point; or
- it beats at least three top-tier bots and its mean is at least 53%; or
- its worst opponent is at least 48% and it beats one top-tier bot at 52% or more.
These rules were written for the first few iterations. After a run's first big jump, every candidate clears rule 1, so the rules stop filtering. What a run does about that gap is part of what the benchmark shows. Some runs kept nearly everything. Others wrote stricter rules for themselves, such as multi-seed checks.
18.3How we run it
The prompt
Every session gets the same prompt. Only the iteration count, the effort and the model name change.
You are continuing a clean-slate strategy autoresearch arm on this branch. Read autoresearch_strategies/loop.md and autoresearch_strategies/strategies_journal.md end to end, including all prior session summaries, then run the loop exactly as loop.md describes for N more iterations and stop.
Rules for this arm: (1) State the pinned model and reasoning effort in the session header. (2) A Python venv with numpy is on PATH, and
bench.shworks as-is. Write scratch files under$TMPDIR, not/tmp. Treat a bench as hung only when no METRIC lines appear after 180 seconds. (3) Continue strategy IDs from the journal and commit after every iteration. (4) Do not read other git branches or worktrees. Do not modifyloop.md,fast_sim_wrapper.pyor the opponent pool. (5) After N iterations, append a session summary naming the champion and its metrics, commit it, and stop.
Source: runs/prompts/heuristic.md (rules condensed; the full text is in the repo).
The runner
A small shell script, runs/run_heuristic.sh, drives each run. It counts the autoresearch: keep and discard commits since the scaffold. While the count is below 115, it starts a new session for the next 15 iterations, or fewer at the end. Each session is a fresh agent process that knows the run only through the repo and its journal.
An iteration is one benched candidate. loop.md has a third outcome, "noted": the journal keeps the entry and the code is discarded. The runner did not count noted candidates, so a few runs benched more than 115. Opus 5.5 medium-b benched 131. The results in chapter 17 use each run's champion as of its 115th candidate, so every run has the same budget.
If a session ends without a new commit, the runner assumes a usage limit, waits 10 minutes, and tries again. It gives up after 36 empty sessions in a row in wave 2, which is 6 hours, and after 144 in wave 3, which is 24 hours. On 24 September the Codex account hit its weekly limit at 17:37 ET. The runners waited, and after a reset the first new iteration landed at 17:48 ET.
The two harnesses
| Engine | Command | Host and sandbox |
|---|---|---|
| Claude Code | claude -p PROMPT --model M --effort E --strict-mcp-config --dangerously-skip-permissions, with the effort also committed in .claude/settings.json | Mac. No sandbox, the same as the wave-1 Claude runs. |
| Codex | codex exec -m M -c model_reasoning_effort=E --ignore-user-config -c approval_policy="never" PROMPT | code-01, a GCP VM. Codex's workspace-write sandbox, with the run's .git writable. Writes outside the run fail. |
Claude runs stay on the Mac for a practical reason. code-01 runs Ubuntu 24.04, and its AppArmor profile for bwrap breaks Claude Code's sandbox helper, so every sandboxed shell command fails there. Codex's sandbox works on code-01, and code-01 holds other work, so Codex runs there only inside the sandbox.
Shared-machine rules
- A private temp directory.
loop.mdsuggests writing the bench table to/tmp/table.txt. With 11 runs on one machine, runs would overwrite each other's tables. Each run gets its ownTMPDIRinside its repo, and the prompt says to use it. - A 180-second hang rule.
loop.mdsays a bench with no result after about 10 seconds has hung. That assumed a bench takes about 2 seconds. Under shared load a bench took about 8 seconds on code-01 and 13 seconds on the Mac. In the first 15 minutes of wave 2, the Codex runs discarded many good candidates as "benchmark timeout". At 16:15 ET on 24 September we raised the threshold to 180 seconds in the prompt and restarted all 16 wave-2 strategy runs from scratch. The aborted attempts are kept as git bundles outside the runs, taggedhang10s.
Sources: runs/run_heuristic.sh, runs/README.md (protocol, deviations and caveats).
18.4Three waves of runs
The benchmark grew in three waves. Wave 1 produced one run per model and effort. That turned out to be too few, so waves 2 and 3 add replicate runs under one uniform harness.
| Wave | Dates | Runs | Models and efforts | Harness |
|---|---|---|---|---|
| 1 | 1 Jul – 22 Sep 2026 | 14 | Claude Opus (older), Fable 5, Opus 5, Fable 5.1 medium/high/max, Opus 5.5 medium/high; GPT-6 Astra low to ultra | Claude: full agent sessions. Astra: a proposal-only driver, where the model proposed each strategy and a script built, benched and kept it. One run per cell; 13 to 129 iterations. |
| 2 | 24 Sep 2026 | 16 | Sonnet 5 medium ×2, high ×1; Opus 5.5 medium; Fable 5.1 medium; GPT-5.6 Luna medium ×2, high ×1; GPT-6 Astra low, medium, high, xhigh ×2 each | The uniform runner above, 115 iterations each. Astra now runs as a full agent. |
| 3 | 27 Sep 2026 | 11 | Sonnet 5 high ×2; Opus 5.5 high ×2; Fable 5.1 high ×2, max ×3; GPT-5.6 Luna high ×2 | Same as wave 2. Brings every high and max cell to three runs. |
Wave 1's deviations, such as the Astra driver, runs that stopped early or overshot, and a run that expanded its own opponent pool, are listed in full under Methods & Data.
Sources: runs/manifest.tsv, runs/manifest_wave3.tsv, research/data/heuristic_arms.csv.
18.5From run to result
A finished run leaves a repo, a journal and a commit per iteration. Four steps turn that into the numbers in chapter 17.
- Cap at 115. Only the first 115 benched candidates count, including noted ones. This changes the champion in two runs: Opus 5.5 medium in wave 1, where X212 (74.6%) replaces X227 (76.6%), and Opus 5.5 medium-b in wave 2, where X209 (76.4% patched) replaces X222.
- Pick the champions. Each run's best clean strategy, which never aims triple at the bull, and its best exploit strategy if it has one. The ladder agent benches every strategy in the run to confirm the choice and the journal's scores.
- Rescore with the bull patch. Each champion is re-benched raw and with the patch (19.4) at 25,000 games per matchup. In wave 2 only one run used the bull bug: Opus 5.5 medium-b's X222 fell from 77.87% raw to 76.32% patched, and its clean X230 scored 76.55%.
- Mark the discovery. Almost every strong run found the same idea: when ahead, shut the opponent's scoring lane. It is worth 6 to 9 points in one step. We record the first iteration at which a run benched a clean candidate at 60% or more. That is a simple, uniform proxy for the moment the run found the idea.
- Enter the ladder. Each champion must reproduce its recorded score in the ladder's own harness. It then plays every other champion head to head (chapter B).
Sources: ladder/wave2/rebench/champions.md; scripts/build_runs.py (the discovery threshold).