Cricket Bench

Two benchmarks for research agents, played on one dartboard

Cricket Bench gives a model an open research problem with a fast, honest score. In one benchmark the model writes a better darts strategy by hand. In the other it builds and trains an agent that learns to play. Every run starts from pinned code and plays fixed bots. Every champion then meets every other on one Elo ladder.

Models Claude Fable 5.1, Opus 5.5, Sonnet 5, earlier Claude · GPT-6 Astra, GPT-5.6 LunaRuns strategy writing and self-training, three wavesLadder 2,701 pairs · 9.06M games

Most benchmarks ask a model for an answer. These two ask it to run an investigation: propose an idea, build it, measure it against opponents that never change, decide whether the evidence is good enough, and repeat. The game is cricket darts. The simulator is fast, the bots are fixed, and the score is checked by a second, independent tournament.

20.1The two benchmarks

Strategy writing

Hand-coded strategies in C · 115 iterations · 11 fixed bots at three skill levels

The model gets a working C engine and one baseline strategy. Each iteration it writes one new strategy, benches it against the pool, keeps or discards it by fixed rules, and journals why. The deliverable is the best strategy it found. This tests idea generation, experimental discipline and reading an unfamiliar engine.

How a run works (18) · Results (17)

Self-training

A learned agent · about 4 hours of training compute · 11 fixed bots at pro level

The model gets the game in Python, the bots, and PyTorch. It must build something that learns to beat the bots. Any learning method is allowed, but the policy has to be learned, not hand-written. Pass is a mean win rate above 50%. This tests method choice, debugging a learning system and honest evaluation.

How a run works (16) · Results (15)

20.2How every run works

Both benchmarks follow the same six steps. Chapter 19 covers the parts they share: the game, the engines, the bots and the isolation rules.

  1. Pin. Each run gets a fresh git repository that holds only the pinned starting code. The model and its reasoning effort are fixed for the whole run.
  2. Launch. A script starts an unattended agent session, Claude Code or Codex, with a fixed prompt. It restarts the session after usage limits and stops when the run is complete.
  3. Work. The model writes code, benches it and records every decision in its journal. Each step is a git commit, so the whole run can be replayed.
  4. Freeze. The run's own final benchmark names its champion.
  5. Rescore. We re-run each champion with a known engine bug patched out (chapter 19.4).
  6. Rate. Every champion from both benchmarks joins one head-to-head ladder (chapter B).

20.3What we found

  1. Strategy writing turns on one idea, and finding it is partly luck.

    Almost every strong run found the same rule: when ahead, shut the opponent's scoring lane. It is worth 6 to 9 points in one step. The runs that never found it are the weakest runs in the benchmark, whatever their model or effort. 17.1

  2. One run per model and effort is not enough.

    Two or three runs of the same model at the same effort differ by 4 to 182 Elo. That is as large as most differences between models. So every cell now has replicate runs, and the site shows each run as its own dot. 17.2

  3. More reasoning effort did not reliably help.

    With replicates, no model gets consistently stronger as its effort rises. More effort always costs more tokens. 17.3

  4. In self-training, the choice of method decides the score.

    Runs that learned a value network and searched over the known dice scored 71 to 78% against the pool. Linear scorers scored 60 to 63%. Two newer models fell back to imitating a bot and scored close to 50%. 15.3

  5. A high pool score is not a strong player.

    The best pool scores in both benchmarks belong to players that rank outside the ladder's top 15. The ladder leader is still IT3, a self-trained agent from July. B.4

20.4How to read this site