Reducer Pong 753M (WebGPU demo)
A fine-tuned copy of Reduction of States' "Reduced by loss tests" 753M model, exported so it can play Pong live in a web browser via WebGPU (demo on reductionofstates.com/research/).
What it is
- Base:
CONSEQUENCE_PRUNE_750M, a 752,977,920-parameter Llama-style decoder (24 layers, hidden 2048, 32 query / 8 KV heads, RoPE, SwiGLU with width 2944, RMSNorm, tied embeddings). It was pruned from a 998M parent that was trained from scratch, then recovery-trained, in Reduction of States' local MLX pilot. The pilot's frozen weights were only read, never modified. - Pong skill: a full fine-tune of a copy of those weights (700 AdamW steps, lr 5e-5).
- Task: imitate a simple "follow the ball" paddle.
- Input: a text game state, e.g.
pong ball x 12 y 7 dx right dy up paddle 9 move. - Output: a single move token (
up/stay/down). - Halfway through, states from the model's own play were added and labeled by the teacher (DAgger).
- Export:
- ONNX with the 24 transformer blocks unchanged, and the vocabulary sliced to the 48 tokens any Pong prompt can contain. The output is the three move logits, which with tied embeddings equal the full model's logits for those tokens.
- Weights quantized to 4-bit (
MatMulNBits, block 32). - File:
pong_q4.onnx, 429,718,753 bytes, sha2564703f09a1e6ef0fae3a377aa9ca09bd3a3802d43a8b969e6acc4cc536f7f1dc8.
Measured (40 games, 400 ticks each, same seeds for every paddle)
| Paddle | Returns | Misses | Agrees with teacher |
|---|---|---|---|
| This model (4-bit, as shipped) | 259 | 3 | 92.9% |
| Random | 5 | 40 | 34.9% |
| Teacher (follow the ball) | 280 | 0 | 100% |
- 4-bit and fp32 choose the same move on 99.2% of 500 random states.
- The ONNX graph matches the MLX model to 1.8e-4 in logits.
- Held-out move accuracy was 94.2%, against 37.4% before fine-tuning.
Limits
This is a demonstration: a small language model driven through a text interface to play a toy game. It is not a general game-playing agent, and it is not evidence about Reducer's pruning quality. The teacher is a trivial rule; the model imitates it imperfectly.
Files
pong_q4.onnx: the model. Inputids(int64, [1, 16]); outputmove_logits[1, 3].pong_tokens.json: word → token-id pieces (in the sliced vocabulary) used to buildids.EVAL.json,FINETUNE.json: raw evaluation and training logs.