Cricket Bench

The ladder: every champion plays every other

Each benchmark scores a run against its own fixed pool of bots. That says how well a champion beats those bots. It does not say how strong the champion is. The ladder answers that. Every champion from both benchmarks plays every other one, on one engine with the bull bug fixed, and one Elo scale ranks them all.

Players 74 (strategy champions, trained agents, 9 hand-coded anchors)Games 2,701 pairs · 9.06M gamesFit Bradley–Terry, S1 = 1000

The two benchmarks use different engines, different bot pools and different game counts, so their pool scores can't be compared. The ladder puts every champion on one scale. Each one plays every other one, thousands of games per pair. A single rating model turns the results into Elo.

B.1How the ladder is played

The ladder protocol · research/05-ladder.md §1
ItemRule
PlayersThe best clean champion of every strategy-writing run, plus its best exploit champion where one exists. The final model of every self-training run. The July artifacts IT3, IT2_det, PPO_run2 and DIST. Nine hand-coded anchors: S1, S2, S6, S10, S14, E1, E10, E12 and PS.
SkillPro against pro. The first thrower alternates every game.
EngineThe C engine for pairs of C strategies and numpy networks. The Python rules of record, game.py, for any pair with a Python agent. The engine is recorded per match.
Bull bugPatched for everyone. A triple aimed at the bull uses the double-aim outcome distribution. Clean champions play bit-identically with or without the patch.
Games4,000 per pair on the C engine and 2,000 per pair on game.py.
SeedsA fixed base per pair, from a hash of the two names. Game g uses base + g. Any pair can be replayed exactly.
StallsA game stops at 2,000 darts. The side ahead on points loses, because it could have ended the game by closing and did not.
RatingA Bradley–Terry fit over every pair, anchored at S1 = 1000, with 400/ln 10 per natural unit. Standard errors come from the Fisher information and a 200-sample bootstrap.
New playersA new player plays only its own pairs. Existing pairs are never replayed. Refitting the wave-1 players from the same match data reproduces their ratings exactly.

B.2Gates: every player must reproduce itself first

A port that changes behaviour would rate the wrong player. So before any player is rated, the ladder's copy has to reproduce that player's own recorded result.

No player failed a gate, so none was excluded. Source: ladder/gates/gate_summary.json.

B.3The ratings

Head-to-head Elo, every player

Circle = strategy-writing champion. Diamond = self-training agent. Small grey dot = hand-coded anchor. Bars show ±2 standard errors.

  • Fable
  • Opus
  • Sonnet 5
  • GPT-6 Astra
  • GPT-5.6 Luna
Reading it. A 10-Elo gap is about a 51.4% head-to-head edge. Standard errors are 1.0 to 1.2 Elo, so gaps of a few Elo are real but small. The chart shows each run's champions within its first 115 benched candidates (chapter 18.5). Five later champions are rated too and kept in the raw files. Source: ladder/results/ratings.csv.

B.4What the ladder shows

  1. A self-trained agent from July is still first.

    IT3, a Fable 5 league-trained network, leads every champion from both benchmarks. It beats each of the top strategy champions 51 to 54% head to head.

  2. The best champions of several models are nearly tied.

    The top strategy champions from Fable 5.1 and Opus 5.5 sit within a few Elo of each other. The best GPT-6 Astra champions sit about 10 Elo behind them. The best Sonnet 5 and GPT-5.6 Luna champions sit further back.

  3. Runs of the same model spread widely.

    Two or three runs of one model at one effort can differ by 4 to 182 Elo. That spread is as large as most gaps between models. Chapter 17 shows why: most of it comes from whether a run found one key idea.

  4. A top pool score does not make a top player.

    Pool score and ladder rating agree only loosely: the rank correlation is 0.85 for strategy champions and 0.75 for trained agents. The best pool score in the strategy benchmark, Opus 5.5 high's X189, rates 1132, about 25 Elo below the top strategy champions. Its champion remembers each opponent's habits within a game, which works against the fixed pool and less well against strong, varied opponents.

B.5Caveats

Sources: research/05-ladder.md §1–10; ladder/results/ratings.csv, pool_vs_ladder.md, sensitivity.md, fit.json.