None of this was planned as a benchmark. It started with a question from a real bar league: what is the best way to play cricket darts? Every attempt to answer that question produced a better simulator, a harder opponent pool, or a faster scoring loop. Around July the question turned around. The game was well understood, so the useful thing to measure was the researcher.
There is a second story running underneath. The code was written with AI agents from day one. The commit trailers name Claude Opus 4.6 and Sonnet 4.6 in February and March, Opus 4.7 in April, Fable 5 and Opus 4.8 in July, and Fable 5.1, Opus 5.5 and GPT-6 Astra in September. The project's tools improved as the models did. Eventually the models themselves became the thing under test.
19.1The game, in one paragraph
Cricket uses seven targets: 20 down to 15, plus the bull. Three marks close a target (a single is one mark, a double two, a triple three; the bull has no triple). Once you have closed a target that your opponent has not, extra hits score its face value. You win by closing all seven while at least level on points. The catch is that aim is not outcome. Every dart samples a result from a skill profile. A pro aiming at a triple hits it 41% of the time and misses the number entirely 14% of the time. So every decision trades speed against risk, and scoring against defence.
19.2The eras
-
Feb 5–18
Learning from scratch fails
Three parallel agents built the engine in a day. Tabular Q-learning and a Double DQN both stalled near 45%. A branching actor-critic (A2C) went through nineteen versions and twelve catalogued bugs. The worst: misses were accidentally disabled, the game became deterministic, and a "91% win rate going first" turned out to be physics.
Best learned agent: 49.3% against the hard pool, after 19 versions.
That same 11-bot pool later became the RL benchmark pool.
-
Feb 7 – Mar 26
Reproducing the paper, then out-simulating it
The reference was Frongello's 2018 study of cricket strategy. All 17 of its strategies were rebuilt as one parameterised class and played across 11 skill levels. A harder bull (a 0.75 difficulty multiplier) changed the rankings. One tournament ran 86 hours on the homelab. The whole dataset reached about 18.7 billion games. On 26 March a new rule, E12, became the best classic bot. It covers a lane the opponent has closed when it is one mark from finishing.
E12: #1 of 30 classic strategies. It becomes the baseline for everything after.
-
Feb 18 – Mar 11
Search, and the first autoresearch loop
AlphaZero v1 scored 0% against 28 bots (under-trained). MuZero was built and never produced a result. Then came the idea that shaped everything later. An agent follows a written program: edit one file (an MCTS evaluation function), run a tournament, keep or revert, log, never stop. Forty-six experiments took the MCTS agent from 1.3% to 48%.
At 99% accuracy the agent still scored 47%. The wall was strategy, not aim.
-
Apr 17–21
The false summit, then the loop that stuck
A strategy called Phase Switch looked like #1 at 9 of 11 skill levels. Two simulator bugs, a 200-turn cap and truncated games credited to the wrong side, had inflated it. After the fix it ranked 8th to 10th. The lesson became a rule: no result is kept on one measurement.
On 20 April, commit
7b098ecadded a strategy loop in the style of Karpathy's autoresearch. A model writes a new strategy branch in C, benches it against 11 bots in two seconds, and keeps or discards it. In two days the main lineage went from E12's 54.6% to X188 at 60.7%, "The Shape Reader". Then its own descendants were added to the pool, and "progress" X165 lost to its grandparent X109 at 36.9%.The scaffold commit 7b098ec is where every benchmark arm still starts.
-
May – Jun
Ten weeks of silence
Zero sessions and zero commits in any repository.
-
Jul 1–2
Clean slates: the first model comparison
The loop was restarted from
7b098econ isolated branches with fresh agents. Claude Opus ran 13 iterations and reached 55.2%. Claude Fable 5 found "faucet denial" at iteration 11, +9.4 points in one step, and after 25 iterations it beat the 89-entry main lineage by 2.7 points on that lineage's own pool. Fresh eyes beat accumulated history.The same day, a clean-room RL task was written: train a learned agent that beats the 11-bot pool, with about four hours of compute and no access to prior work. Fable 5 cloned E12 by behaviour cloning, then ran PPO. Run 1 learned never to finish a game. Run 2 reached 75.8% in 26 minutes of training. A league loop then produced IT3, "The Closer".
From 49.3% (19 versions, February) to 75.8% (one session, July).
-
Jul 2–7
A ladder and a ground truth
A vectorised C environment (10.6× faster) made a unified Bradley–Terry Elo ladder possible: 27 artifacts, 253 C pairings at 4,000 games each. IT3 topped it at 1222. It had zero intransitive cycles. An exact endgame tablebase of 4.16 billion states then graded every champion's endgame. Even IT3 gives away 0.027 win probability per game, and a tablebase endgame beats it by 2.8 points head-to-head.
The champion is strong, hard to exploit (a trained adversary wins only 53.2%), and still not optimal.
-
Jul 24 – Sep 24
The benchmark
The clean-slate procedure was written down as a protocol. The model under test forks from
7b098ecwith no memory tools and no other branches, and runs a target of 115 iterations. A second model independently rebuilds the harness and re-runs it. Opus 5 reached 64.4%. Fable 5.1 reached 75.9%. On 6 September reasoning effort became a controlled variable, and GPT-6 Astra ran six effort levels.The Fable 5.1 max-effort arm then found that the engine was wrong. A triple aimed at the bull skipped the difficulty penalty, and that was worth about four points. The independent review had missed it. From then on, arms were compared on their clean champions. Opus 5.5 finished the set: 77.8% clean on the heuristic task, and 77.9% on the RL task.
Fourteen heuristic arms, ten clean-room RL arms, three model families, one scaffold.
19.3What carried forward
Four things built in the early months are what make the benchmark work:
- A fast, validated engine. The C simulator matches the Python rules to within 1.8 points at 5,000 games, and plays a million games in about two seconds.
- A fixed opponent pool. Eleven bots that span the strategy space, from S1 (the pure closer) to E12 (the old champion).
- The loop. One file, one number, keep or discard, journal everything. It was written for an agent in March, and every heuristic arm still runs it.
- Scepticism about single numbers. The false summit, the deterministic-game bug, the grandparent that beat its descendant, and the bull artifact each taught the same lesson. That lesson is why every result on this site carries its pool, seed and game count.
The strategy side of the story, including the Frongello reproduction, The Shape Reader and The Closer, is told in full at darts.mattalldian.com. It is frozen as of 2 July 2026.