CHSM8
CHSM8 is a research chess move-prediction model. Given a game prefix, it predicts the next chess move while applying a legal-move mask at inference. It is a project-specific PyTorch model, not a standard Transformers model.
Architecture
- Llama-style causal chess decoder: 12 layers, width 960, 15 attention heads, SwiGLU MLP (width 3,520), RoPE, and Q/K normalization.
- Factorized chess output heads predict move kind, source square, destination square, piece, and promotion rather than a flat move vocabulary.
- 210.5M parameters (210,528,719), of which 166.2M were trained in pretraining. The remaining 44.3M form a cross-attention scaffold that was frozen and fed a zero context during pretraining; SFT trains it to read text.
- Trained in bf16 with FlashAttention, fused AdamW, gradient clipping 1.0, and weight decay 0.1.
Training data and schedule
Training data: LegumMagister/chsm8-pre-training
(1.07B filtered Lichess games; the dataset's scripts/ rebuild the exact packed format). Training uses packed Lichess standard-chess game records, represented as move
sequences. The final branch reads the 3p-s packed corpus without replacement
in 512-position sequences; shards 77–127 were globally reshuffled during
training to remove a temporal data-order artefact.
The planned one-pass budget is 78.32B chess positions. Optimisation uses a
position-indexed cosine learning-rate schedule: 5% warmup, peak LR 5e-4, and
minimum LR equal to 10% of the peak. A 1% held-out split is reserved for
validation. The model's legal-move penalty weight is 0.5.
Checkpoint
| File | Step | Description |
|---|---|---|
chsm8_latest.pt |
1,194,960 | Final checkpoint of the training run. |
A weights-only PyTorch dictionary containing model_state_dict, step,
total_training_time, total_tokens_processed, and cfg. The model classes and the loading and play code will be released later with the project code.
Evaluation
Estimated strength: 1837 Elo (95% CI 1817–1856) on Stockfish's UCI_Elo
scale.
| Games | W / D / L | Score | Elo (95% CI) |
|---|---|---|---|
| 999* | 139 / 283 / 577 | 0.281 | 1837 (1817–1856) |
* 1,000 games played; one reached the 240-ply cap without a result and is excluded.
Protocol: colour-balanced games, greedy legal-move decoding (the model always
plays its highest-scoring legal move; no search, no sampling), against
Stockfish 17.1 with UCI_LimitStrength / UCI_Elo=2000, 0.05 s per move,
1 thread, 240-ply cap. Elo = 2000 + 400·log10(S / (1 − S)), with S the score
(win = 1, draw = ½); the interval comes from the per-game standard error.
UCI_Elo is calibrated against engine rating lists, not Lichess or FIDE
ratings, so treat the number as a reproducible project metric.
Checkpoint choice. We originally expected the 85%-of-training checkpoint to be the better model, because it precedes a spike in training loss late in the run, and it did lead after a 200-game evaluation. Over 1,000 games, however, this final checkpoint came out ahead, so it is the one published.
Comparison with other chess models
Head-to-head, 200 games per opponent: 100 book openings, each played with both colours; both sides play their single most likely legal move (no search, no sampling); games past 178 plies count as draws (3 of 1,000).
| Opponent | chsm8-pt W / D / L | chsm8-pt score | Elo advantage (95% CI) |
|---|---|---|---|
| Maia-2 blitz (rating input 2000+) | 102 / 52 / 46 | 0.640 | +100 (+59, +144) |
Karvonen ChessGPT lichess_16layers |
123 / 43 / 34 | 0.723 | +166 (+122, +216) |
| chess-gpt2-uci-12x12x768 | 151 / 37 / 12 | 0.848 | +298 (+249, +359) |
| Waterhorse ChessGPT-base-v1 (2.8B) | 158 / 32 / 10 | 0.870 | +330 (+278, +397) |
| Kempner ChessFormer 50M, ≤1500 data | 163 / 30 / 7 | 0.890 | +363 (+309, +434) |
All compared models, like chsm8-pt, are trained only by supervised next-move / next-token prediction, with no reinforcement learning, reward models or preference optimization.
Limitations
CHSM8 is a research checkpoint with no search, opening book, endgame tablebase, or standard Transformers interface. It has been evaluated only under the project protocol above; use it for research and reproduce results with the project code once it is released.