CHSM8: a chess model that plays in the style you describe
The final CHSM8 model (formerly chsm8-dpo; that name redirects here). Lineage: chsm8-pt
(pretraining) β chsm8-sft (style SFT) β this model (DPO).
210.5M parameters; the prompt is encoded by a frozen google/embeddinggemma-300m. No search; every move is legal by construction.
chsm8-sft further trained with DPO so that it follows a natural-language
style prompt more strongly ("trade queens as early as possible", "push your pawns aggressively up the board"...).
Same architecture and usage as chsm8-sft: decoder over factorized move tokens, cross-attention to a frozen
google/embeddinggemma-300m encoding of the prompt (all token states as keys) in the top 6 of 12 layers.
Two models in this repo
| file | training | when to use |
|---|---|---|
model.pt (main) |
20-minute DPO run on 1 GPU: 4,366 steps Γ 128 games, about 0.56M games (about 2% of the corpus) | default; strongest prompt adherence |
full_pass/model.pt |
one full pass over 27.08M games on 8 GPUs: 13,428 steps Γ about 2,000 games; checkpoints at 25/50/75% in full_pass/checkpoints/ |
the "DPO 100%" model: same recipe, about 50Γ more exposure |
Both use the identical loss and hyperparameters (below); they differ in run length, batch size and LR schedule length.
Which to pick. model.pt is the main model because it follows style prompts most strongly (66.5% early queen trades vs
58.7%). Treat its lead with caution: it was the best of about 70 short runs, so part of it may be a lucky seed or our own metric
hacking (see below); a retrain with a new seed reached 64.8%. Both models are a large, clear step over SFT, and the differences
between DPO settings are mostly within evaluation noise:
| model | weighted z-score over the five metrics |
|---|---|
| chsm8-sft | β21 |
model.pt (20-min run), 2,000 games |
+2.8 |
full_pass/model.pt, 2,000 games |
β1.4 |
| 20-min recipe retrained (new seed), 1,000 games | β0.1 |
z = distance from the mean of 20 short DPO runs in units of their spread, per metric (queen trades, pawn moves, queenside
castling Γ0.2, no-prompt Elo, attack β no prompt gap), weights adjusted for the correlation between metrics. One 1,000-game eval
has about Β±2 z of noise: the same model.pt weights scored +4.8 on one 1,000-game batch and +0.8 on another. Any DPO setting
landing within about Β±2 of 0 is equally good as far as we can measure; SFT at β21 is far below all of them.
Results (2,000 games per prompt vs Stockfish 17.1 UCI_Elo 2000, 0.05 s/move, greedy legal moves)
| model | queens traded by move 20 "trade queens as early as possible" |
pawn moves per own move "push your pawns aggressively..." |
queenside share of castling "castle queenside and storm..." |
Elo, no prompt | Elo, "attack the king at all costs and sacrifice material..." β no prompt |
|---|---|---|---|---|---|
| chsm8-sft (start) | 53.0% | 0.290 | 97.3% | 1792 | β40 |
model.pt (20-min run) |
66.5% | 0.319 | 99.7% | 1797 | β150 |
full_pass/model.pt |
58.7% | 0.317 | 99.6% | 1812 | β155 |
Precision at 2,000 games: shares Β±1β2 points, Elo Β±21 (95%). Each row pools two independent 1,000-game evals. The metrics are proxies for prompt adherence and strength.
- Prompt adherence rises over SFT on all three style prompts; strength without a prompt is unchanged within noise (the no-text path is pinned).
- Under risky prompts the model follows the prompt and loses more: "attack at all costs, sacrifice material" costs about 150 Elo vs no prompt (SFT: about 40). Playing sacrifices against an accurate engine without search costs strength; every method that raised adherence in our tests (DPO variants, classifier-free guidance) showed this.
Full-pass checkpoints (1,000 games per prompt):
| checkpoint | queens | pawns | queenside | Elo, no prompt | attack β no prompt |
|---|---|---|---|---|---|
25% (full_pass/checkpoints/p25.pt) |
58.1% | 0.317 | 99.7% | 1811 | β134 |
| 50% | 61.4% | 0.320 | 99.7% | 1774 | β91 |
| 75% | 56.9% | 0.317 | 99.8% | 1790 | β118 |
| 100% | 57.1% | 0.319 | 99.6% | 1807 | β159 |
Validation (2,000 val games, cross-entropy on the described side's moves):
| right label | wrong label | no label | pair accuracy | |
|---|---|---|---|---|
20-min run (model.pt, 1,000 val games) |
2.846 | 3.600 | 2.932 | 0.893 |
| SFT | 2.806 | 3.094 | 2.928 | 0.803 |
| full pass 25% | 2.855 | 3.960 | 2.932 | 0.912 |
| full pass 50% | 2.853 | 4.027 | 2.933 | 0.921 |
| full pass 75% | 2.850 | 4.061 | 2.933 | 0.925 |
| full pass 100% | 2.851 | 4.077 | 2.933 | 0.925 |
Pair accuracy = share of games whose moves are more likely under their own description than under another game's.
How far above baseline, and what it costs (2,000 games per row)
Style metrics with no prompt, with a random description from the test split (a different one per game), and with the prompt targeting that metric ("trade queens", "push pawns", "castle queenside"):
| model | prompt | queens traded | pawn moves | queenside castling |
|---|---|---|---|---|
| chsm8-sft | no prompt | 21.3% | 0.251 | 14.1% |
| random description | 21.1% | 0.249 | 10.2% | |
| targeted prompt | 52.9% | 0.290 | 97.3% | |
model.pt (20-min run) |
no prompt | 23.4% | 0.251 | 13.5% |
| random description | 16.4% | 0.251 | 8.3% | |
| targeted prompt | 66.5% | 0.319 | 99.7% | |
full_pass/model.pt |
no prompt | 22.8% | 0.253 | 13.3% |
| random description | 19.1% | 0.254 | 9.7% | |
| targeted prompt | 58.7% | 0.317 | 99.6% |
Elo under each prompt:
| model | none | random description | "trade queens" | "push pawns" | "castle queenside" | "attack at all costs" |
|---|---|---|---|---|---|---|
| SFT | 1792 | 1791 | 1827 | 1819 | 1761 | 1752 |
| DPO 20-min (main) | 1797 | 1722 | 1784 | 1742 | 1717 | 1647 |
| DPO full pass | 1812 | 1726 | 1767 | 1760 | 1750 | 1657 |
- The targeted prompt moves each metric far above both baselines (queen trades about 3Γ, queenside castling from 10β14% to almost 100%); no prompt and a random description give nearly the same style, so the effect comes from the prompt's content.
- SFT keeps its strength under ordinary prompts; DPO follows every prompt harder and pays for it: 15β80 Elo under style
prompts, about 75β85 under a random description, about 150 under "attack at all costs". The model that follows prompts
most (
model.pt) pays the most. Without a prompt both DPO models play as strongly as SFT.
What went wrong, and what we learned
Short runs did not predict the full run. We chose the recipe from about 70 short DPO runs (10β20 min), re-evaluated the best at 1,000 games per prompt, and ran the winner for a full pass. The full pass came out less style-adherent than the 20-minute run it was meant to scale up (queen trades 58.7% vs 66.5% at 2,000 games), so the 20-minute run is now the main model.
Evaluation noise was as large as the effects we optimised. Re-running the same 1,000-game eval of the same checkpoint moved no-prompt Elo by up to 41 and the attack gap by up to 41 (e.g. 20-min run: 1817 β 1776 Elo; long-band 100%: gap β92 β β133), enough to reorder the candidates. Between-recipe spread vs 1,000-game noise: queens 2.1 vs 1.5 points, Elo 10 vs about 15, gap 24 vs about 21, queenside 0.38 vs 0.33. Picking the best of many noisy runs selects partly for luck (winner's curse), i.e. we partly hacked our own metric. Only queen trades and pawn moves separated recipes reliably.
One band vs a random band of five. Pairs from the long description only (1β2 sentences) instead of a random one of the five bands: 25% of the pass 64.2% / 0.318 / 99.1% / 1790 / β107; full pass 60.9% / 0.308 / 99.6% / 1798 / β113 (2,000 games). More queen trades than the random-band full pass but weaker pawn adherence by the end; not published.
Possible causes of the short-vs-full gap (under test):
- Batch size at a fixed learning rate. The 20-min run used 128 games per step; the full pass about 2,000 (256 per GPU Γ 8 GPUs) with the same lr. With Adam each step moves the weights by about the lr regardless of batch, so per game seen the full run updates about 16Γ less; at a similar step count (full pass 25% β 3,400 steps vs 4,366) it had less style (58% vs 66% queens). Small-batch gradient noise may also help the conditioning path.
- LR schedule. The 20-min run's cosine decays within 20 minutes; the full pass spends most of its steps at high lr, which keeps sharpening label discrimination on validation (pair accuracy 0.89 β 0.93) without adding style in play.
- Multi-GPU code path. Hand-written gradient all-reduce, per-rank data streams; a defect there would look the same.
- Test (20-minute runs, 1,000 games per prompt), queen trades: 1 GPU Γ 128 games (new seed) 64.8%; 4 GPUs Γ 32 (same total batch) 63.1%; 4 Γ 256 (8Γ batch, same lr) 61.6%; full pass (about 16Γ batch) 57β58%. Adherence falls as the batch grows, while the multi-GPU path at the same batch matches the single GPU, so no sign of a multi-GPU bug. Raising the lr 2.8Γ at 4 Γ 256 did not bring the style back (61.0%, pair accuracy 0.906 vs 0.889), so a too-small step is not the explanation. Larger-batch runs also saw 4Γ more games (2.5M vs 0.59M) in the same 20 minutes, and the steps between runs are 1β2 standard errors each; we cannot yet separate batch size from exposure or noise.
Robust findings: label-pair DPO raises adherence on all three style prompts; no-prompt strength stays within noise of SFT; stronger adherence comes with a larger Elo drop under "attack at all costs". Finer choices (Ξ², KL weight, pair type, band) are within noise at 1,000 games per prompt. Ranking recipes properly needs repeated runs and β₯2β4k games per prompt.
Training (both models)
- Pairs from the labelled corpus, no sampling: chosen = a game's moves under its own style description, rejected = the same moves under another game's description (one random length band of 5 per game). Moves are always the real human moves.
- Loss: DPO (Ξ² 0.5) on the summed log-prob of the described side's moves, reference = frozen SFT; + NLL on the game's moves (weight 1.0) + KL(reference β policy) per move (weight 1.0) as a strength guard; the no-text cross-attention term stays pinned.
- Decoder lr 2e-7, conditioning lr 1e-4, 5% warmup then cosine, AdamW (0.9, 0.95), grad clip 1.0.
model.pt: 1 GH200, 128 games / β€98k tokens per step, cosine over 20 min of training time, 4,366 steps.full_pass/model.pt: 8 GH200 (2 nodes, data parallel), 256 games / β€197k tokens per GPU, cosine over one pass (27.08M games), 13,428 steps, about 3,300 games/s, 2.0 h.
Files
model.pt, full_pass/model.pt, full_pass/checkpoints/p25.pt, p50.pt, p75.pt; same format as chsm8-sft (chess, cond,
frozen_dummy, args, dpo_args). The model classes and the loading and play code will be released later with the project code.
Limitations
Research checkpoint, no search. Style prompts are followed more strongly at the cost of strength under prompts that ask for risky play. Trained on games of β₯ 2200 Lichess players and top players; prompts describing much weaker play are out of distribution.
