Cricket Bench

Seven months, from a Q-table to a benchmark

The project began as an attempt to teach a computer to play bar darts. Each failure left behind a sharper instrument. By September the instrument was the interesting part: a fixed game, fixed opponents, and a two-second score that any research agent could be pointed at.

Span 5 Feb – 24 Sep 2026Commits 215 on main, 1,025 across all branchesSources git history, journals, session logs

None of this was planned as a benchmark. It started with a question from a real bar league: what is the best way to play cricket darts? Every attempt to answer that question produced a better simulator, a harder opponent pool, or a faster scoring loop. Around July the question turned around. The game was well understood, so the useful thing to measure was the researcher.

There is a second story running underneath. The code was written with AI agents from day one. The commit trailers name Claude Opus 4.6 and Sonnet 4.6 in February and March, Opus 4.7 in April, Fable 5 and Opus 4.8 in July, and Fable 5.1, Opus 5.5 and GPT-6 Astra in September. The project's tools improved as the models did. Eventually the models themselves became the thing under test.

19.1The game, in one paragraph

Cricket uses seven targets: 20 down to 15, plus the bull. Three marks close a target (a single is one mark, a double two, a triple three; the bull has no triple). Once you have closed a target that your opponent has not, extra hits score its face value. You win by closing all seven while at least level on points. The catch is that aim is not outcome. Every dart samples a result from a skill profile. A pro aiming at a triple hits it 41% of the time and misses the number entirely 14% of the time. So every decision trades speed against risk, and scoring against defence.

19.2The eras

  1. Feb 5–18

    Learning from scratch fails

    Three parallel agents built the engine in a day. Tabular Q-learning and a Double DQN both stalled near 45%. A branching actor-critic (A2C) went through nineteen versions and twelve catalogued bugs. The worst: misses were accidentally disabled, the game became deterministic, and a "91% win rate going first" turned out to be physics.

    Best learned agent: 49.3% against the hard pool, after 19 versions.

    That same 11-bot pool later became the RL benchmark pool.

  2. Feb 7 – Mar 26

    Reproducing the paper, then out-simulating it

    The reference was Frongello's 2018 study of cricket strategy. All 17 of its strategies were rebuilt as one parameterised class and played across 11 skill levels. A harder bull (a 0.75 difficulty multiplier) changed the rankings. One tournament ran 86 hours on the homelab. The whole dataset reached about 18.7 billion games. On 26 March a new rule, E12, became the best classic bot. It covers a lane the opponent has closed when it is one mark from finishing.

    E12: #1 of 30 classic strategies. It becomes the baseline for everything after.

  3. Feb 18 – Mar 11

    Search, and the first autoresearch loop

    AlphaZero v1 scored 0% against 28 bots (under-trained). MuZero was built and never produced a result. Then came the idea that shaped everything later. An agent follows a written program: edit one file (an MCTS evaluation function), run a tournament, keep or revert, log, never stop. Forty-six experiments took the MCTS agent from 1.3% to 48%.

    At 99% accuracy the agent still scored 47%. The wall was strategy, not aim.

  4. Apr 17–21

    The false summit, then the loop that stuck

    A strategy called Phase Switch looked like #1 at 9 of 11 skill levels. Two simulator bugs, a 200-turn cap and truncated games credited to the wrong side, had inflated it. After the fix it ranked 8th to 10th. The lesson became a rule: no result is kept on one measurement.

    On 20 April, commit 7b098ec added a strategy loop in the style of Karpathy's autoresearch. A model writes a new strategy branch in C, benches it against 11 bots in two seconds, and keeps or discards it. In two days the main lineage went from E12's 54.6% to X188 at 60.7%, "The Shape Reader". Then its own descendants were added to the pool, and "progress" X165 lost to its grandparent X109 at 36.9%.

    The scaffold commit 7b098ec is where every benchmark arm still starts.

  5. May – Jun

    Ten weeks of silence

    Zero sessions and zero commits in any repository.

  6. Jul 1–2

    Clean slates: the first model comparison

    The loop was restarted from 7b098ec on isolated branches with fresh agents. Claude Opus ran 13 iterations and reached 55.2%. Claude Fable 5 found "faucet denial" at iteration 11, +9.4 points in one step, and after 25 iterations it beat the 89-entry main lineage by 2.7 points on that lineage's own pool. Fresh eyes beat accumulated history.

    The same day, a clean-room RL task was written: train a learned agent that beats the 11-bot pool, with about four hours of compute and no access to prior work. Fable 5 cloned E12 by behaviour cloning, then ran PPO. Run 1 learned never to finish a game. Run 2 reached 75.8% in 26 minutes of training. A league loop then produced IT3, "The Closer".

    From 49.3% (19 versions, February) to 75.8% (one session, July).

  7. Jul 2–7

    A ladder and a ground truth

    A vectorised C environment (10.6× faster) made a unified Bradley–Terry Elo ladder possible: 27 artifacts, 253 C pairings at 4,000 games each. IT3 topped it at 1222. It had zero intransitive cycles. An exact endgame tablebase of 4.16 billion states then graded every champion's endgame. Even IT3 gives away 0.027 win probability per game, and a tablebase endgame beats it by 2.8 points head-to-head.

    The champion is strong, hard to exploit (a trained adversary wins only 53.2%), and still not optimal.

  8. Jul 24 – Sep 24

    The benchmark

    The clean-slate procedure was written down as a protocol. The model under test forks from 7b098ec with no memory tools and no other branches, and runs a target of 115 iterations. A second model independently rebuilds the harness and re-runs it. Opus 5 reached 64.4%. Fable 5.1 reached 75.9%. On 6 September reasoning effort became a controlled variable, and GPT-6 Astra ran six effort levels.

    The Fable 5.1 max-effort arm then found that the engine was wrong. A triple aimed at the bull skipped the difficulty penalty, and that was worth about four points. The independent review had missed it. From then on, arms were compared on their clean champions. Opus 5.5 finished the set: 77.8% clean on the heuristic task, and 77.9% on the RL task.

    Fourteen heuristic arms, ten clean-room RL arms, three model families, one scaffold.

19.3What carried forward

Four things built in the early months are what make the benchmark work:

The strategy side of the story, including the Frongello reproduction, The Shape Reader and The Closer, is told in full at darts.mattalldian.com. It is frozen as of 2 July 2026.