CHSM8: a chess model that plays in the style you describe

The final CHSM8 model (formerly chsm8-dpo; that name redirects here). Lineage: chsm8-pt (pretraining) β†’ chsm8-sft (style SFT) β†’ this model (DPO). 210.5M parameters; the prompt is encoded by a frozen google/embeddinggemma-300m. No search; every move is legal by construction.

chsm8-sft further trained with DPO so that it follows a natural-language style prompt more strongly ("trade queens as early as possible", "push your pawns aggressively up the board"...). Same architecture and usage as chsm8-sft: decoder over factorized move tokens, cross-attention to a frozen google/embeddinggemma-300m encoding of the prompt (all token states as keys) in the top 6 of 12 layers.

Two models in this repo

file training when to use
model.pt (main) 20-minute DPO run on 1 GPU: 4,366 steps Γ— 128 games, about 0.56M games (about 2% of the corpus) default; strongest prompt adherence
full_pass/model.pt one full pass over 27.08M games on 8 GPUs: 13,428 steps Γ— about 2,000 games; checkpoints at 25/50/75% in full_pass/checkpoints/ the "DPO 100%" model: same recipe, about 50Γ— more exposure

Both use the identical loss and hyperparameters (below); they differ in run length, batch size and LR schedule length.

Which to pick. model.pt is the main model because it follows style prompts most strongly (66.5% early queen trades vs 58.7%). Treat its lead with caution: it was the best of about 70 short runs, so part of it may be a lucky seed or our own metric hacking (see below); a retrain with a new seed reached 64.8%. Both models are a large, clear step over SFT, and the differences between DPO settings are mostly within evaluation noise:

model weighted z-score over the five metrics
chsm8-sft βˆ’21
model.pt (20-min run), 2,000 games +2.8
full_pass/model.pt, 2,000 games βˆ’1.4
20-min recipe retrained (new seed), 1,000 games βˆ’0.1

z = distance from the mean of 20 short DPO runs in units of their spread, per metric (queen trades, pawn moves, queenside castling Γ—0.2, no-prompt Elo, attack βˆ’ no prompt gap), weights adjusted for the correlation between metrics. One 1,000-game eval has about Β±2 z of noise: the same model.pt weights scored +4.8 on one 1,000-game batch and +0.8 on another. Any DPO setting landing within about Β±2 of 0 is equally good as far as we can measure; SFT at βˆ’21 is far below all of them.

Results (2,000 games per prompt vs Stockfish 17.1 UCI_Elo 2000, 0.05 s/move, greedy legal moves)

model queens traded by move 20
"trade queens as early as possible"
pawn moves per own move
"push your pawns aggressively..."
queenside share of castling
"castle queenside and storm..."
Elo, no prompt Elo, "attack the king at all costs and sacrifice material..." βˆ’ no prompt
chsm8-sft (start) 53.0% 0.290 97.3% 1792 βˆ’40
model.pt (20-min run) 66.5% 0.319 99.7% 1797 βˆ’150
full_pass/model.pt 58.7% 0.317 99.6% 1812 βˆ’155

Precision at 2,000 games: shares Β±1–2 points, Elo Β±21 (95%). Each row pools two independent 1,000-game evals. The metrics are proxies for prompt adherence and strength.

  • Prompt adherence rises over SFT on all three style prompts; strength without a prompt is unchanged within noise (the no-text path is pinned).
  • Under risky prompts the model follows the prompt and loses more: "attack at all costs, sacrifice material" costs about 150 Elo vs no prompt (SFT: about 40). Playing sacrifices against an accurate engine without search costs strength; every method that raised adherence in our tests (DPO variants, classifier-free guidance) showed this.

Full-pass checkpoints (1,000 games per prompt):

checkpoint queens pawns queenside Elo, no prompt attack βˆ’ no prompt
25% (full_pass/checkpoints/p25.pt) 58.1% 0.317 99.7% 1811 βˆ’134
50% 61.4% 0.320 99.7% 1774 βˆ’91
75% 56.9% 0.317 99.8% 1790 βˆ’118
100% 57.1% 0.319 99.6% 1807 βˆ’159

Validation (2,000 val games, cross-entropy on the described side's moves):

right label wrong label no label pair accuracy
20-min run (model.pt, 1,000 val games) 2.846 3.600 2.932 0.893
SFT 2.806 3.094 2.928 0.803
full pass 25% 2.855 3.960 2.932 0.912
full pass 50% 2.853 4.027 2.933 0.921
full pass 75% 2.850 4.061 2.933 0.925
full pass 100% 2.851 4.077 2.933 0.925

Pair accuracy = share of games whose moves are more likely under their own description than under another game's.

How far above baseline, and what it costs (2,000 games per row)

Style metrics with no prompt, with a random description from the test split (a different one per game), and with the prompt targeting that metric ("trade queens", "push pawns", "castle queenside"):

model prompt queens traded pawn moves queenside castling
chsm8-sft no prompt 21.3% 0.251 14.1%
random description 21.1% 0.249 10.2%
targeted prompt 52.9% 0.290 97.3%
model.pt (20-min run) no prompt 23.4% 0.251 13.5%
random description 16.4% 0.251 8.3%
targeted prompt 66.5% 0.319 99.7%
full_pass/model.pt no prompt 22.8% 0.253 13.3%
random description 19.1% 0.254 9.7%
targeted prompt 58.7% 0.317 99.6%

Elo under each prompt:

model none random description "trade queens" "push pawns" "castle queenside" "attack at all costs"
SFT 1792 1791 1827 1819 1761 1752
DPO 20-min (main) 1797 1722 1784 1742 1717 1647
DPO full pass 1812 1726 1767 1760 1750 1657
  • The targeted prompt moves each metric far above both baselines (queen trades about 3Γ—, queenside castling from 10–14% to almost 100%); no prompt and a random description give nearly the same style, so the effect comes from the prompt's content.
  • SFT keeps its strength under ordinary prompts; DPO follows every prompt harder and pays for it: 15–80 Elo under style prompts, about 75–85 under a random description, about 150 under "attack at all costs". The model that follows prompts most (model.pt) pays the most. Without a prompt both DPO models play as strongly as SFT.

What went wrong, and what we learned

Short runs did not predict the full run. We chose the recipe from about 70 short DPO runs (10–20 min), re-evaluated the best at 1,000 games per prompt, and ran the winner for a full pass. The full pass came out less style-adherent than the 20-minute run it was meant to scale up (queen trades 58.7% vs 66.5% at 2,000 games), so the 20-minute run is now the main model.

Evaluation noise was as large as the effects we optimised. Re-running the same 1,000-game eval of the same checkpoint moved no-prompt Elo by up to 41 and the attack gap by up to 41 (e.g. 20-min run: 1817 β†’ 1776 Elo; long-band 100%: gap βˆ’92 β†’ βˆ’133), enough to reorder the candidates. Between-recipe spread vs 1,000-game noise: queens 2.1 vs 1.5 points, Elo 10 vs about 15, gap 24 vs about 21, queenside 0.38 vs 0.33. Picking the best of many noisy runs selects partly for luck (winner's curse), i.e. we partly hacked our own metric. Only queen trades and pawn moves separated recipes reliably.

One band vs a random band of five. Pairs from the long description only (1–2 sentences) instead of a random one of the five bands: 25% of the pass 64.2% / 0.318 / 99.1% / 1790 / βˆ’107; full pass 60.9% / 0.308 / 99.6% / 1798 / βˆ’113 (2,000 games). More queen trades than the random-band full pass but weaker pawn adherence by the end; not published.

DPO checkpoints, long band (solid) vs random band of 5 (dashed)

Possible causes of the short-vs-full gap (under test):

  • Batch size at a fixed learning rate. The 20-min run used 128 games per step; the full pass about 2,000 (256 per GPU Γ— 8 GPUs) with the same lr. With Adam each step moves the weights by about the lr regardless of batch, so per game seen the full run updates about 16Γ— less; at a similar step count (full pass 25% β‰ˆ 3,400 steps vs 4,366) it had less style (58% vs 66% queens). Small-batch gradient noise may also help the conditioning path.
  • LR schedule. The 20-min run's cosine decays within 20 minutes; the full pass spends most of its steps at high lr, which keeps sharpening label discrimination on validation (pair accuracy 0.89 β†’ 0.93) without adding style in play.
  • Multi-GPU code path. Hand-written gradient all-reduce, per-rank data streams; a defect there would look the same.
  • Test (20-minute runs, 1,000 games per prompt), queen trades: 1 GPU Γ— 128 games (new seed) 64.8%; 4 GPUs Γ— 32 (same total batch) 63.1%; 4 Γ— 256 (8Γ— batch, same lr) 61.6%; full pass (about 16Γ— batch) 57–58%. Adherence falls as the batch grows, while the multi-GPU path at the same batch matches the single GPU, so no sign of a multi-GPU bug. Raising the lr 2.8Γ— at 4 Γ— 256 did not bring the style back (61.0%, pair accuracy 0.906 vs 0.889), so a too-small step is not the explanation. Larger-batch runs also saw 4Γ— more games (2.5M vs 0.59M) in the same 20 minutes, and the steps between runs are 1–2 standard errors each; we cannot yet separate batch size from exposure or noise.

Robust findings: label-pair DPO raises adherence on all three style prompts; no-prompt strength stays within noise of SFT; stronger adherence comes with a larger Elo drop under "attack at all costs". Finer choices (Ξ², KL weight, pair type, band) are within noise at 1,000 games per prompt. Ranking recipes properly needs repeated runs and β‰₯2–4k games per prompt.

Training (both models)

  • Pairs from the labelled corpus, no sampling: chosen = a game's moves under its own style description, rejected = the same moves under another game's description (one random length band of 5 per game). Moves are always the real human moves.
  • Loss: DPO (Ξ² 0.5) on the summed log-prob of the described side's moves, reference = frozen SFT; + NLL on the game's moves (weight 1.0) + KL(reference β€– policy) per move (weight 1.0) as a strength guard; the no-text cross-attention term stays pinned.
  • Decoder lr 2e-7, conditioning lr 1e-4, 5% warmup then cosine, AdamW (0.9, 0.95), grad clip 1.0.
  • model.pt: 1 GH200, 128 games / ≀98k tokens per step, cosine over 20 min of training time, 4,366 steps.
  • full_pass/model.pt: 8 GH200 (2 nodes, data parallel), 256 games / ≀197k tokens per GPU, cosine over one pass (27.08M games), 13,428 steps, about 3,300 games/s, 2.0 h.

Files

model.pt, full_pass/model.pt, full_pass/checkpoints/p25.pt, p50.pt, p75.pt; same format as chsm8-sft (chess, cond, frozen_dummy, args, dpo_args). The model classes and the loading and play code will be released later with the project code.

Limitations

Research checkpoint, no search. Style prompts are followed more strongly at the cost of strength under prompts that ask for risky play. Trained on games of β‰₯ 2200 Lichess players and top players; prompts describing much weaker play are out of distribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for LegumMagister/chsm8

Finetuned
(1)
this model

Dataset used to train LegumMagister/chsm8