Every run faced the same eleven bots with the same budget. The pool scores run from 50.2% to 77.9%. On the head-to-head ladder the same agents run from 917 to 1170 Elo, and the two orders do not agree. July's league champion, IT3, is still the strongest agent of either benchmark.
15.1Strength, head to head
Ladder rating of every self-training agent
Elo from the head-to-head round robin, S1 = 1000. Bars are ±2 standard errors. Hover a row for the pool score.
- Claude Fable
- Claude Opus
- GPT-6 Astra
- Claude Sonnet 5
- GPT-5.6 Luna
15.2What each run built
| Agent | Wave | Method | Pool % | Elo | |
|---|---|---|---|---|---|
| Fable 5 · IT3 | 1 | PPO run 2, fine-tuned for determinism, then league self-play against 14 members | 74.7 | 1169.5 | |
| Fable 5 · IT2_det | 1 | An earlier league iteration, greedy | 74.95 | 1156.1 | |
| Fable 5.1 medium (b) | 1 | Value net + 1-dart expectimax, Monte-Carlo targets with a per-dart discount of 0.98 | 70.91 | 1143.7 | |
| Fable 5 · PPO run 2 | 1 | Behaviour clone of E12, then PPO against the pool; sampled policy | 75.79 | 1137.8 | |
| Opus 5.5 medium (b) | 2 | Value net over positions, model-based TD(λ), 1-dart expectimax | 77.9 | 1134.0 | |
| Opus 5.5 medium | 1 | Afterstate value net, TD(λ) with expected backups, 1-dart expectimax | 77.88 | 1131.1 | |
| GPT-6 Astra high (b) | 2 | Behaviour clone of S2, then PPO; sampled policy | 72.62 | 1128.7 | |
| Fable 5.1 medium | 1 | Afterstate value net + expectimax, Monte-Carlo win/loss targets, ply 2 | 72.25 | 1122.2 | |
| Fable 5 · AlphaZero cell | 1 | Constrained to AlphaZero-style self-play with chance nodes, depth 3 | 65.63 | 1111.3 | |
| GPT-6 Astra low | 1 | Cross-entropy method over 20 linear weights | 62.50 | 1101.5 | |
| Opus 5.5 high | 1 | Afterstate value net, λ-returns with model-based max backup, 1-ply | 76.91 | 1099.4 | |
| GPT-6 Astra high | 1 | Cross-entropy method over 55 linear weights | 61.35 | 1097.5 | |
| GPT-6 Astra ultra | 1 | Cross-entropy method over 32 linear weights, 8 s of training | 60.16 | 1037.5 | |
| Fable 5 · DIST | 1 | Hand-executable rules distilled from PPO run 2 (an analysis, not training) | 60.71 | 1024.8 | |
| Sonnet 5 medium | 2 | Behaviour clone of 13 bots. Both RL fine-tunes made it worse, so it shipped with no RL | 50.19 | 924.5 | |
| GPT-5.6 Luna medium | 2 | Neural imitation of S2, after a tabular Q-learner learned nothing | 55.17 | 916.8 | |
| Opus (version unrecorded) | 1 | Planned a policy-gradient actor-critic, trained no policy | — | — | — |
Pool scores are each run's own final benchmark. Fable 5.1 medium is rated at ply 2, the setting its run reported (71.58% at ply 1). IT3's pool score is derived from the 11 pool cells of its official 2,000-game table. Sources: research/data/runs_all.csv, rl_arms.csv, rl_wave2.csv, ladder/results/ratings.csv.
15.3The method family decides the score
The engine exposes the exact probability of every dart outcome. A run that reads the engine first sees that it does not have to learn the dice. It only has to learn how good a position is, and then it can search one dart ahead over the known outcomes.
Five runs built exactly that: two on Fable 5.1 and three on Opus 5.5. Each learned a value network over positions and chose throws by expectimax over the true outcome distribution. Two runs took the other obvious route, cloning a strong bot and then improving it with PPO: Fable 5 in July, and GPT-6 Astra high in wave 2. The three wave-1 Astra runs, under their coordinator, tuned small linear scorers with the cross-entropy method. Two wave-2 runs never got past imitation.
Ranked by pool score: value net + expectimax with TD targets (76.9–77.9%) > clone then PPO (72.6–75.8%) > value net + expectimax with Monte-Carlo targets (70.9–72.3%) > AlphaZero-style self-play (65.6%) > linear cross-entropy method (60.2–62.5%) > imitation only (50.2–55.2%).
Sources: research/03-rl-arms.md, cross-arm finding 2; research/data/rl_wave2.csv.
Opus 5.5 medium replicated itself
The wave-2 run chose the same family again: a value network, model-based TD(λ) and 1-dart expectimax. It reached 77.9% on the pool, the same as wave 1 (77.88%). The ladder ratings agree to 2.9 Elo: 1134.0 and 1131.1. It trained for about 2 hours and used 54,602 output tokens. Against 18 bots it never trained on, it scored 82.6%.
Astra's second high-effort run chose a stronger method
Without the coordinator framing, GPT-6 Astra high cloned S2 from 80,073 demonstration states and then ran PPO. It reached 72.62% over 44,000 games after 37.2 minutes of training. That rates 1128.7, against 1097.5 for the wave-1 Astra high run and its linear scorer.
Sonnet 5 and Luna shipped copies of the bots
Sonnet 5's DQN diverged. It then cloned 13 bots and scored 51.4% in a quick check. Two RL fine-tunes on top of the clone both made it worse, one down to 13.9%, so it shipped the clone at 50.19%. Luna's tabular Q-learner won almost no games in its smoke test, so it trained a network to imitate S2 and shipped that at 55.17%. Both rate below S1. The Sonnet clone loses to S2 43.5% head to head, and the Luna clone beats S2 51.6%.
Within the value-net family, the pool gap between Opus 5.5 (77–78%) and Fable 5.1 (71–72%) follows the training target. The Opus runs bootstrapped with TD(λ) and model-based backups. The Fable runs regressed on final win or loss. Search depth barely mattered: one dart deeper added 0 to 1.3 points in every run that tried it. No run tested the target head to head, so this is an inference across runs, not a controlled result.
15.4Every run met the same trap
A game only ends when someone closes all seven targets while level or ahead on points. A learner paid only for winning can find that not losing is easier. It keeps scoring and never closes. Every run that trained a policy met some version of this.
| Run | Form | Fix |
|---|---|---|
| Fable 5 PPO | Run 1 learned to hoard points and refuse to close. Training win rate rose to 0.77 by survivorship. 13–34% of greedy games deadlocked. | A dart cap with a loss for stalling. The shipped policy is sampled, because its greedy version still deadlocks. |
| Fable 5.1 | The value saturated at 1.0 in won positions, so every action tied and argmax threw single-15 forever. | A win bonus plus a tie-break, with a regression test. |
| Fable 5.1 (b) | Points races into the thousands, because Monte-Carlo win/loss targets carry no notion of time. | Score clipping and a per-dart discount of γ = 0.98, confirmed by a controlled A/B test. |
| Opus 5.5 high | A bootstrapped self-loop became a fixed point. 64 of 100 games against E1 hit the cap. | Diagnosed with a traced game. γ < 1 fixed it but cost 0.5–1.1 points. Under a pre-registered rule the run shipped its γ = 1 model and documented the stall risk. |
| GPT-6 Astra high (b) | Greedy decoding caused scoring loops. | Evaluate and ship the sampled policy at temperature 1.0. Later checkpoints showed no stalls, and all 44,000 final games finished. |
| Sonnet 5 | RL fine-tuning brought back very slow games. One evaluation game did not finish in 120 s. | Dropped the fine-tune. The shipped clone still stalled 13 times in its gate, and the ladder charged it 32 stall losses. |
The runs that handled it best traced a single stalled game and explained the mechanism before touching the reward. That is the move a human RL researcher makes, and it is visible in the journals.
15.5Opponent modelling bought about one point
Each bot has a fixed habit, so it seems obvious that an agent should learn who it is playing. Five runs tried this: sampling weights, opponent-tendency features, evidence inputs and an auxiliary opponent-ID loss. The best result was +0.5 to +1.1 points, all of it against the "racing" bots. The wave-2 Opus run found +0.9 points, within two to three standard errors.
Two measurements bound the headroom. A trained best response adds at most about 1.8 points (Fable 5's league loop). Specialists trained on one opponent do no better than the generalist on that opponent. The three chase bots, S10, S14 and S16, hold every method to 57–63%. The limit is the game's tempo, not knowledge of the opponent.
15.6Pool score is not strength
A high pool score says an agent punishes these eleven bots. It does not say the agent plays cricket well. The ladder shows the gap for self-training more sharply than for strategy writing.
- The two highest pool scores, both Opus 5.5 medium at 77.9%, rate 1134.0 and 1131.1, well below the top agents.
- Fable 5.1 medium (b) has the lowest value-net pool score, 70.91%, but it is the best-rated agent of any run since July, at 1143.7. It beats the wave-1 Opus 5.5 medium agent 59.7% head to head.
- Opus 5.5 high scores 76.91% on the pool and rates 1099.4. The Opus 5.5 medium agent beats it 60.2%.
- IT3, rated first at 1169.5, scores 74.7% on the pool. It trained in a league against 14 members, including strong hand-coded champions. The pool-only runs never saw a strong opponent.
The best new agent from wave 2, Opus 5.5 medium (b), rates 34.5 Elo below IT3. No new model has matched July's league champion.
15.7Caveats
- Few runs per cell. Only Fable 5.1 medium, Opus 5.5 medium and Astra high have two runs. Every other cell has one. The strategy benchmark shows that runs of one cell can differ by 50 Elo or more.
- Different harnesses. The wave-1 Astra runs worked under a coordinator whose objective was a reproducible environment. The wave-2 Astra run did not. Their gap mixes model behaviour with framing.
- Shared compute. In wave 2, four runs trained on one Mac with other jobs, capped at 4 threads each. The budget is wall-clock time, so these runs got less compute than wave 1.
- Imitation is allowed. TASK.md allows heuristics as "training curricula" and names "imitation+improvement" as a method. A pure clone meets the letter of the task. The Sonnet run called its own result a gray area.
- Tokens exclude training. Token counts measure the model's work in the session. They do not include the compute the trainer used.
- Effort unrecorded before September. The Fable 5 and older Opus runs ran at the default effort, which was never written down.