File size: 2,037 Bytes
477cf8b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fca717c
 
477cf8b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
---
license: mit
tags:
  - chess
  - transformer
  - recurrent
  - policy
  - puzzles
  - pytorch
library_name: pytorch
---

# 99M puzzle expert (squares64)

Tactics specialist. Same **99M** squares64 architecture as
[`avewright/chess-transformer-100m-squares64`](https://huggingface.co/avewright/chess-transformer-100m-squares64),
finetuned on official Lichess puzzles (solver plies after the opponent setup move).

This file is **`latest.pt` at puzzle-FT step 19156** (2026-09-13 18:29 UTC).
Recent train loss ~0.2239. Last puzzle holdout hard CE ~0.2317.

This repo is **not** the generalist incumbent. Do not overwrite that model
with these weights.

## Architecture

| | |
|---|---|
| Params | 98.97M |
| Hidden / heads | 736d / 8 |
| Encoder dim | 256 |
| Trunk | prefix 4 + bank 7×3 unrolls + suffix 4 |
| Effective depth | 29 (15 unique layer modules) |
| Attention | 64×64 squares only |
| Vocab | **1968** compact |

Config is in `model_config.json`.

## Training

- Warm start: public 99M `latest.pt` (weights only; optimizer reset).
- Data: [`Lichess/chess-puzzles`](https://huggingface.co/datasets/Lichess/chess-puzzles)
  `479ea9bc9f681385f5adb23fa27a96c2dc8ae599`
- Split: PuzzleId-level 80/20 (seed 273). Train rows 11195282; eval 2824449.
- Labels: one-hot solver moves. `value_valid=0`.
- Optimizer: Polar-NorMuon + AdamW aux. `muon_lr=0.002`, `adam_lr=3e-5`.

## Inference

Requires `avewright/transform` (`chess_inference.py`, `chess_squares64.py`) and compact vocab 1968.

```python
import os
os.environ["MOVE_VOCAB_VERSION"] = "compact"

import chess
from huggingface_hub import hf_hub_download
from chess_inference import load_checkpoint, get_model_move

path = hf_hub_download("avewright/puzzle-model", "latest.pt")
model = load_checkpoint(path, device="cpu")
board = chess.Board()
move, info = get_model_move(model, board, device="cpu")
print(move, info["top_moves"], info["wdl"])
```

## Files

- `latest.pt` — full train checkpoint
- `model_config.json`
- `train.log`
- `pack.json` — puzzle split counts