CHSM8-SFT: chess moves conditioned on a natural-language style description

chsm8-pt fine-tuned to play one side of a game in the style a text describes (e.g. "trade everything and grind the endgame") or like a named top player ("play like Carlsen").

Highlights

Conditioning works: the right description makes the described side's moves more predictable. Test set (8,000 games), cross-entropy on the described side's moves:

right description wrong description (another game's) no description
all 9 labels 2.880 2.929 2.905
5 LLM style descriptions only 2.855
player prompts ("play like {player}", val) 2.977 2.980 2.987

Style descriptions carry most of the signal (about 0.05 nats below no text); player prompts help only slightly so far.

Prompts change how it plays (model's own moves vs Stockfish 2000, no prompt in brackets):

  • "leave your king in the center and go for an open fight": castles in 45% of games [94%]
  • "castle queenside and storm the enemy king with your pawns": 39% of castling is queenside [10%]
  • "trade queens as early as possible": queens off by move 20 in 40% of games [25%]
  • "push your pawns aggressively up the board and grab space": pawn moves 0.289 per move [0.248]
  • "bring your queen out early and push random pawns for no reason": early queen moves 16% [7%], fewer Sicilians
  • "attack the king at all costs...": most checks per move (0.082 [0.072]), shorter games (89 plies [98])
  • every prompt plays more varied openings than no prompt (0.77–0.96 distinct lines per game [0.75])

Prompts don't change how well it plays. No prompt: 1828 Elo (1808–1847, 1,000 games). Strong prompts ("a 2800-rated super grandmaster...") 1765–1836 and weak prompts ("you just learned how the pieces move") 1778–1839 overlap within the confidence intervals: every training game came from a β‰₯ 2200 player, so no description was ever paired with weak play. Prompts that push risky behaviour cost strength (king in the center 1735, "1 minute" 1756).

Elo by prompt

Behaviour by prompt

Data

  • chsm8-post-training: 27.08M train / 1.50M val / 1.50M test game sides (described player rated β‰₯ 2200), each with 5 LLM style descriptions (weight 1.0) and 4 metadata descriptions (rating, time control, opening, trading; weight 0.05).
  • 542,109 train games of 99 FIDE top-100 players (chsm8-top-players) with one "play like {player}" prompt each (49 templates Γ— aliases); about 1% of games for the first 80% of training, ramping to about 9.5% at the end.
  • One pass, every training game exactly once, order random across years.

Conditioning

  • The description is encoded by a frozen google/embeddinggemma-300m (up to 128 tokens). All its token states, projected to the decoder width, are the keys and values of a cross-attention in the top 6 of 12 decoder layers.
  • Each cross-attention output is added as x + tanh(gate) Β· attention, gates initialised at 0 (training starts exactly at the pretrained model). With no description the model is the pretrained policy.
  • Loss: next-move cross-entropy on the described side's moves only; the opponent's moves are context.

Training

  • Length-bucketed batches: up to 128 games per GPU, games Γ— 5 descriptions Γ— longest game ≀ 98,304 tokens. 2 nodes Γ— 4 GH200 (8 GPUs, DDP), about 2,400 games/s, peak 53 GB per GPU, 3.1 h of training.

Final metrics (cross-entropy on the described side's moves)

matched = the game's own description, mismatched = another game's description, zero = no text; p* = player prompts.

metric value
pval_games 8000.0000
pval_matched 2.9782
pval_matched_ce 2.9772
pval_matched_ce_llm 0.0000
pval_matched_ce_meta 2.9772
pval_matched_minus_mismatched -0.0022
pval_matched_minus_zero -0.0092
pval_mismatched_ce 2.9795
pval_zero_ce 2.9865
test_games 8000.0000
test_matched 2.8802
test_matched_ce 2.8797
test_matched_ce_llm 2.8546
test_matched_ce_meta 2.9110
test_matched_minus_mismatched -0.0488
test_matched_minus_zero -0.0251
test_mismatched_ce 2.9285
test_zero_ce 2.9048
val_every 1000.0000
val_games 8000.0000
val_matched 2.8776
val_matched_ce 2.8770
val_matched_ce_llm 2.8532
val_matched_ce_meta 2.9069
val_matched_minus_mismatched -0.0477
val_matched_minus_zero -0.0239
val_mismatched_ce 2.9248
val_zero_ce 2.9010

Post-SFT eval

vs Stockfish 17.1 UCI_Elo 2000, 0.05 s/move, greedy legal moves. Behaviour is measured on the model's own moves.

elo

behaviour

group prompt games Elo (95% CI) game length (plies) captures / own move checks / own move pawn moves / own move queens traded by ply 40 castled castled queenside (of castled) queen moved in first 5 own moves distinct first-8-own-move lines / game Sicilian as Black (1.e4 c5) mean centipawn loss
strong You are Magnus Carlsen in his prime: squeeze every tiny edge until it becomes a win 300 1836 (1797–1871) 101 0.219 0.068 0.246 23% 93% 16% 9% 0.83 100% 200
strong a 2800-rated super grandmaster with flawless technique 300 1829 (1792–1862) 98 0.216 0.070 0.242 17% 95% 8% 5% 0.88 100% 199
strong think like an engine: precise, accurate and ruthless 300 1827 (1789–1862) 95 0.218 0.072 0.227 21% 96% 6% 4% 0.78 100% 186
strong play like a 2750 player against a similar-rated player 299 1818 (1779–1852) 98 0.213 0.068 0.248 18% 94% 12% 7% 0.82 100% 206
strong mature, patient, grandmaster-level positional chess 299 1813 (1776–1847) 102 0.208 0.054 0.245 16% 96% 8% 7% 0.80 100% 201
strong calculate deeply, never blunder, and convert every advantage cleanly 300 1809 (1770–1845) 97 0.231 0.078 0.246 29% 95% 12% 10% 0.90 100% 194
strong play like a strong GM against an opponent of the same strength 299 1791 (1749–1828) 98 0.215 0.070 0.244 24% 95% 8% 7% 0.82 100% 231
strong play the way a reigning world champion would in a title match 300 1765 (1721–1803) 96 0.211 0.060 0.247 15% 95% 11% 5% 0.77 100% 210
weak play like an 800 rated player against a much stronger opponent 300 1839 (1802–1873) 97 0.220 0.064 0.245 18% 96% 9% 5% 0.82 100% 172
weak play badly on purpose 300 1827 (1789–1862) 98 0.221 0.063 0.252 20% 93% 8% 5% 0.84 100% 194
weak a careless beginner who leaves pieces hanging every few moves 300 1817 (1779–1851) 96 0.217 0.060 0.248 27% 93% 8% 5% 0.88 100% 221
weak play like a 1900 player against someone better than you 300 1817 (1777–1853) 98 0.215 0.064 0.242 27% 95% 15% 8% 0.89 100% 197
weak you just learned how the pieces move 300 1814 (1775–1849) 95 0.221 0.069 0.244 21% 95% 8% 6% 0.88 100% 216
weak move instantly without thinking and ignore every threat 300 1792 (1753–1826) 98 0.219 0.061 0.239 25% 95% 8% 5% 0.81 100% 204
weak a distracted casual player blundering in a noisy bar 300 1790 (1750–1825) 94 0.220 0.057 0.242 23% 96% 9% 5% 0.79 100% 207
weak bring your queen out early and push random pawns for no reason 300 1778 (1736–1816) 90 0.234 0.070 0.237 23% 85% 22% 16% 0.95 0% 247
style keep the position full and avoid every exchange 200 1847 (1800–1888) 97 0.225 0.075 0.243 20% 96% 7% 6% 0.86 100% 202
style play like it's a long classical game with plenty of time 200 1834 (1786–1876) 100 0.219 0.073 0.254 21% 96% 11% 4% 0.81 100% 213
style grind out a long endgame, play on for as long as it takes 199 1813 (1764–1855) 102 0.218 0.059 0.256 24% 97% 8% 7% 0.90 100% 205
style keep the queens on the board no matter what 200 1811 (1764–1853) 99 0.220 0.064 0.244 28% 94% 13% 5% 0.90 100% 204
style give lots of checks and keep harassing the enemy king 200 1809 (1762–1850) 97 0.221 0.076 0.242 28% 96% 9% 9% 0.91 100% 209
style trade everything you can and simplify into an endgame 200 1804 (1759–1844) 97 0.219 0.060 0.245 32% 95% 15% 8% 0.91 100% 203
style castle queenside and storm the enemy king with your pawns 200 1800 (1748–1844) 93 0.222 0.067 0.276 20% 92% 39% 6% 0.84 100% 225
style go for a short, sharp game and finish it fast 200 1785 (1736–1828) 99 0.224 0.071 0.238 22% 94% 6% 2% 0.92 0% 204
style go for the Sicilian Defense 199 1781 (1728–1827) 96 0.219 0.071 0.244 28% 92% 12% 6% 0.87 100% 230
style attack the king at all costs and sacrifice material for the initiative 200 1780 (1724–1828) 89 0.232 0.082 0.251 17% 95% 11% 8% 0.80 100% 211
style solid, safe, risk-free chess: wait for your opponent to make mistakes 200 1780 (1729–1824) 100 0.216 0.060 0.243 24% 96% 14% 10% 0.90 100% 222
style push your pawns aggressively up the board and grab space 200 1775 (1725–1818) 93 0.223 0.055 0.289 20% 90% 9% 14% 0.92 100% 249
style trade queens as early as possible 200 1772 (1722–1816) 102 0.214 0.061 0.246 40% 92% 22% 8% 0.92 100% 196
style you have 1 minute, play like it 200 1756 (1701–1803) 94 0.215 0.063 0.249 20% 94% 11% 4% 0.82 100% 219
style leave your king in the center and go for an open fight 199 1735 (1676–1783) 86 0.215 0.050 0.268 18% 45% 34% 5% 0.96 100% 254
none (no prompt) 1000 1828 (1808–1847) 98 0.221 0.072 0.248 25% 94% 10% 7% 0.75 100% 196

Limitations

Research checkpoint, no search. The model imitates how β‰₯ 2200 Lichess players and top players played; prompts describing much weaker play are outside the training data.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for LegumMagister/chsm8-sft

Finetuned
(2)
this model
Finetunes
1 model

Dataset used to train LegumMagister/chsm8-sft