The two benchmarks use different engines, different bot pools and different game counts, so their pool scores can't be compared. The ladder puts every champion on one scale. Each one plays every other one, thousands of games per pair. A single rating model turns the results into Elo.
B.1How the ladder is played
| Item | Rule |
|---|---|
| Players | The best clean champion of every strategy-writing run, plus its best exploit champion where one exists. The final model of every self-training run. The July artifacts IT3, IT2_det, PPO_run2 and DIST. Nine hand-coded anchors: S1, S2, S6, S10, S14, E1, E10, E12 and PS. |
| Skill | Pro against pro. The first thrower alternates every game. |
| Engine | The C engine for pairs of C strategies and numpy networks. The Python rules of record, game.py, for any pair with a Python agent. The engine is recorded per match. |
| Bull bug | Patched for everyone. A triple aimed at the bull uses the double-aim outcome distribution. Clean champions play bit-identically with or without the patch. |
| Games | 4,000 per pair on the C engine and 2,000 per pair on game.py. |
| Seeds | A fixed base per pair, from a hash of the two names. Game g uses base + g. Any pair can be replayed exactly. |
| Stalls | A game stops at 2,000 darts. The side ahead on points loses, because it could have ended the game by closing and did not. |
| Rating | A Bradley–Terry fit over every pair, anchored at S1 = 1000, with 400/ln 10 per natural unit. Standard errors come from the Fisher information and a 200-sample bootstrap. |
| New players | A new player plays only its own pairs. Existing pairs are never replayed. Refitting the wave-1 players from the same match data reproduces their ratings exactly. |
B.2Gates: every player must reproduce itself first
A port that changes behaviour would rate the wrong player. So before any player is rated, the ladder's copy has to reproduce that player's own recorded result.
- Trained agents. Each adapter replays its run's own pool at 2,000 games per opponent, and must land within 1 point of the recorded pool score. Every one landed within 0.43 points. IT3 and IT2_det reproduced all 11 official cells exactly.
- C strategies. A sample of 45 strategies was checked against its own branch's engine on 20,000 games per cell. All passed.
- Stateful strategies. Three strategies remember the opponent within a game. Each game must match the same game played alone in a fresh process. They matched, 43 of 43 games each.
- Co-loading. Every Python agent must choose the same action with all the others loaded in one process. All 13 did, on 300 probe positions each.
No player failed a gate, so none was excluded. Source: ladder/gates/gate_summary.json.
B.3The ratings
Head-to-head Elo, every player
Circle = strategy-writing champion. Diamond = self-training agent. Small grey dot = hand-coded anchor. Bars show ±2 standard errors.
- Fable
- Opus
- Sonnet 5
- GPT-6 Astra
- GPT-5.6 Luna
B.4What the ladder shows
A self-trained agent from July is still first.
IT3, a Fable 5 league-trained network, leads every champion from both benchmarks. It beats each of the top strategy champions 51 to 54% head to head.
The best champions of several models are nearly tied.
The top strategy champions from Fable 5.1 and Opus 5.5 sit within a few Elo of each other. The best GPT-6 Astra champions sit about 10 Elo behind them. The best Sonnet 5 and GPT-5.6 Luna champions sit further back.
Runs of the same model spread widely.
Two or three runs of one model at one effort can differ by 4 to 182 Elo. That spread is as large as most gaps between models. Chapter 17 shows why: most of it comes from whether a run found one key idea.
A top pool score does not make a top player.
Pool score and ladder rating agree only loosely: the rank correlation is 0.85 for strategy champions and 0.75 for trained agents. The best pool score in the strategy benchmark, Opus 5.5 high's X189, rates 1132, about 25 Elo below the top strategy champions. Its champion remembers each opponent's habits within a game, which works against the fixed pool and less well against strong, varied opponents.
B.5Caveats
- The scale moves when players join. Adding waves 2 and 3 lowered the top ratings by about 11 Elo. Compare ratings within one fit only.
- The anchor is awkward. S1 beats strong players relatively well but loses to E12, S2 and S6 head to head. So most hand-coded bots rate 89 to 109 Elo lower than on the July ladder. Differences among the top players do not depend on the anchor.
- Two engines. Pairs with a Python agent ran on
game.py, and the rest on the C engine. On six test pairs they agreed within |z| < 1.2, but they use different random streams. - Stalls. 451 games stalled, almost all involving self-trained agents that keep scoring instead of closing. The alternative stall rules move no rating by more than 2.1 Elo.
- Cycles. There are 53 rock-paper-scissors triangles with every edge above 52%. None involves the top 18 players.
Sources: research/05-ladder.md §1–10; ladder/results/ratings.csv, pool_vs_ladder.md, sensitivity.md, fit.json.