Cricket Bench

Self-training: what the models built

The method a run chose decided its score more than the model, the effort or the training time. Runs that learned a value function and searched over the known dice did best. Two new models never got reinforcement learning to work and shipped copies of the bots instead.

Runs 15 · 16 rated agentsPool B · pro vs pro · 2,000–4,000 games per matchupBest IT3 · 1169.5 Elo

Every run faced the same eleven bots with the same budget. The pool scores run from 50.2% to 77.9%. On the head-to-head ladder the same agents run from 917 to 1170 Elo, and the two orders do not agree. July's league champion, IT3, is still the strongest agent of either benchmark.

15.1Strength, head to head

Ladder rating of every self-training agent

Elo from the head-to-head round robin, S1 = 1000. Bars are ±2 standard errors. Hover a row for the pool score.

  • Claude Fable
  • Claude Opus
  • GPT-6 Astra
  • Claude Sonnet 5
  • GPT-5.6 Luna
Sources: ladder/results/ratings.csv (Bradley–Terry fit over 2,701 pairs, 2,000 games per pair that involves a Python agent). The hand-coded anchors are left out here. S1 sits at 1000 and the best hand-coded bot, S14, at 1032.0. Table 15.2 lists every value.

15.2What each run built

Method, pool score and ladder rating · sorted by rating
AgentWaveMethodPool %Elo
Fable 5 · IT31PPO run 2, fine-tuned for determinism, then league self-play against 14 members74.71169.5
Fable 5 · IT2_det1An earlier league iteration, greedy74.951156.1
Fable 5.1 medium (b)1Value net + 1-dart expectimax, Monte-Carlo targets with a per-dart discount of 0.9870.911143.7
Fable 5 · PPO run 21Behaviour clone of E12, then PPO against the pool; sampled policy75.791137.8
Opus 5.5 medium (b)2Value net over positions, model-based TD(λ), 1-dart expectimax77.91134.0
Opus 5.5 medium1Afterstate value net, TD(λ) with expected backups, 1-dart expectimax77.881131.1
GPT-6 Astra high (b)2Behaviour clone of S2, then PPO; sampled policy72.621128.7
Fable 5.1 medium1Afterstate value net + expectimax, Monte-Carlo win/loss targets, ply 272.251122.2
Fable 5 · AlphaZero cell1Constrained to AlphaZero-style self-play with chance nodes, depth 365.631111.3
GPT-6 Astra low1Cross-entropy method over 20 linear weights62.501101.5
Opus 5.5 high1Afterstate value net, λ-returns with model-based max backup, 1-ply76.911099.4
GPT-6 Astra high1Cross-entropy method over 55 linear weights61.351097.5
GPT-6 Astra ultra1Cross-entropy method over 32 linear weights, 8 s of training60.161037.5
Fable 5 · DIST1Hand-executable rules distilled from PPO run 2 (an analysis, not training)60.711024.8
Sonnet 5 medium2Behaviour clone of 13 bots. Both RL fine-tunes made it worse, so it shipped with no RL50.19924.5
GPT-5.6 Luna medium2Neural imitation of S2, after a tabular Q-learner learned nothing55.17916.8
Opus (version unrecorded)1Planned a policy-gradient actor-critic, trained no policy———

Pool scores are each run's own final benchmark. Fable 5.1 medium is rated at ply 2, the setting its run reported (71.58% at ply 1). IT3's pool score is derived from the 11 pool cells of its official 2,000-game table. Sources: research/data/runs_all.csv, rl_arms.csv, rl_wave2.csv, ladder/results/ratings.csv.

15.3The method family decides the score

The engine exposes the exact probability of every dart outcome. A run that reads the engine first sees that it does not have to learn the dice. It only has to learn how good a position is, and then it can search one dart ahead over the known outcomes.

Five runs built exactly that: two on Fable 5.1 and three on Opus 5.5. Each learned a value network over positions and chose throws by expectimax over the true outcome distribution. Two runs took the other obvious route, cloning a strong bot and then improving it with PPO: Fable 5 in July, and GPT-6 Astra high in wave 2. The three wave-1 Astra runs, under their coordinator, tuned small linear scorers with the cross-entropy method. Two wave-2 runs never got past imitation.

⊗

Ranked by pool score: value net + expectimax with TD targets (76.9–77.9%) > clone then PPO (72.6–75.8%) > value net + expectimax with Monte-Carlo targets (70.9–72.3%) > AlphaZero-style self-play (65.6%) > linear cross-entropy method (60.2–62.5%) > imitation only (50.2–55.2%).

Sources: research/03-rl-arms.md, cross-arm finding 2; research/data/rl_wave2.csv.

  1. Opus 5.5 medium replicated itself

    The wave-2 run chose the same family again: a value network, model-based TD(λ) and 1-dart expectimax. It reached 77.9% on the pool, the same as wave 1 (77.88%). The ladder ratings agree to 2.9 Elo: 1134.0 and 1131.1. It trained for about 2 hours and used 54,602 output tokens. Against 18 bots it never trained on, it scored 82.6%.

  2. Astra's second high-effort run chose a stronger method

    Without the coordinator framing, GPT-6 Astra high cloned S2 from 80,073 demonstration states and then ran PPO. It reached 72.62% over 44,000 games after 37.2 minutes of training. That rates 1128.7, against 1097.5 for the wave-1 Astra high run and its linear scorer.

  3. Sonnet 5 and Luna shipped copies of the bots

    Sonnet 5's DQN diverged. It then cloned 13 bots and scored 51.4% in a quick check. Two RL fine-tunes on top of the clone both made it worse, one down to 13.9%, so it shipped the clone at 50.19%. Luna's tabular Q-learner won almost no games in its smoke test, so it trained a network to imitate S2 and shipped that at 55.17%. Both rate below S1. The Sonnet clone loses to S2 43.5% head to head, and the Luna clone beats S2 51.6%.

Within the value-net family, the pool gap between Opus 5.5 (77–78%) and Fable 5.1 (71–72%) follows the training target. The Opus runs bootstrapped with TD(λ) and model-based backups. The Fable runs regressed on final win or loss. Search depth barely mattered: one dart deeper added 0 to 1.3 points in every run that tried it. No run tested the target head to head, so this is an inference across runs, not a controlled result.

15.4Every run met the same trap

A game only ends when someone closes all seven targets while level or ahead on points. A learner paid only for winning can find that not losing is easier. It keeps scoring and never closes. Every run that trained a policy met some version of this.

The non-termination trap
RunFormFix
Fable 5 PPORun 1 learned to hoard points and refuse to close. Training win rate rose to 0.77 by survivorship. 13–34% of greedy games deadlocked.A dart cap with a loss for stalling. The shipped policy is sampled, because its greedy version still deadlocks.
Fable 5.1The value saturated at 1.0 in won positions, so every action tied and argmax threw single-15 forever.A win bonus plus a tie-break, with a regression test.
Fable 5.1 (b)Points races into the thousands, because Monte-Carlo win/loss targets carry no notion of time.Score clipping and a per-dart discount of γ = 0.98, confirmed by a controlled A/B test.
Opus 5.5 highA bootstrapped self-loop became a fixed point. 64 of 100 games against E1 hit the cap.Diagnosed with a traced game. γ < 1 fixed it but cost 0.5–1.1 points. Under a pre-registered rule the run shipped its γ = 1 model and documented the stall risk.
GPT-6 Astra high (b)Greedy decoding caused scoring loops.Evaluate and ship the sampled policy at temperature 1.0. Later checkpoints showed no stalls, and all 44,000 final games finished.
Sonnet 5RL fine-tuning brought back very slow games. One evaluation game did not finish in 120 s.Dropped the fine-tune. The shipped clone still stalled 13 times in its gate, and the ladder charged it 32 stall losses.

The runs that handled it best traced a single stalled game and explained the mechanism before touching the reward. That is the move a human RL researcher makes, and it is visible in the journals.

15.5Opponent modelling bought about one point

Each bot has a fixed habit, so it seems obvious that an agent should learn who it is playing. Five runs tried this: sampling weights, opponent-tendency features, evidence inputs and an auxiliary opponent-ID loss. The best result was +0.5 to +1.1 points, all of it against the "racing" bots. The wave-2 Opus run found +0.9 points, within two to three standard errors.

Two measurements bound the headroom. A trained best response adds at most about 1.8 points (Fable 5's league loop). Specialists trained on one opponent do no better than the generalist on that opponent. The three chase bots, S10, S14 and S16, hold every method to 57–63%. The limit is the game's tempo, not knowledge of the opponent.

15.6Pool score is not strength

A high pool score says an agent punishes these eleven bots. It does not say the agent plays cricket well. The ladder shows the gap for self-training more sharply than for strategy writing.

The best new agent from wave 2, Opus 5.5 medium (b), rates 34.5 Elo below IT3. No new model has matched July's league champion.

15.7Caveats