Reducer Pong 753M (WebGPU demo)

A fine-tuned copy of Reduction of States' "Reduced by loss tests" 753M model, exported so it can play Pong live in a web browser via WebGPU (demo on reductionofstates.com/research/).

What it is

  • Base: CONSEQUENCE_PRUNE_750M, a 752,977,920-parameter Llama-style decoder (24 layers, hidden 2048, 32 query / 8 KV heads, RoPE, SwiGLU with width 2944, RMSNorm, tied embeddings). It was pruned from a 998M parent that was trained from scratch, then recovery-trained, in Reduction of States' local MLX pilot. The pilot's frozen weights were only read, never modified.
  • Pong skill: a full fine-tune of a copy of those weights (700 AdamW steps, lr 5e-5).
    • Task: imitate a simple "follow the ball" paddle.
    • Input: a text game state, e.g. pong ball x 12 y 7 dx right dy up paddle 9 move.
    • Output: a single move token (up / stay / down).
    • Halfway through, states from the model's own play were added and labeled by the teacher (DAgger).
  • Export:
    • ONNX with the 24 transformer blocks unchanged, and the vocabulary sliced to the 48 tokens any Pong prompt can contain. The output is the three move logits, which with tied embeddings equal the full model's logits for those tokens.
    • Weights quantized to 4-bit (MatMulNBits, block 32).
    • File: pong_q4.onnx, 429,718,753 bytes, sha256 4703f09a1e6ef0fae3a377aa9ca09bd3a3802d43a8b969e6acc4cc536f7f1dc8.

Measured (40 games, 400 ticks each, same seeds for every paddle)

Paddle Returns Misses Agrees with teacher
This model (4-bit, as shipped) 259 3 92.9%
Random 5 40 34.9%
Teacher (follow the ball) 280 0 100%
  • 4-bit and fp32 choose the same move on 99.2% of 500 random states.
  • The ONNX graph matches the MLX model to 1.8e-4 in logits.
  • Held-out move accuracy was 94.2%, against 37.4% before fine-tuning.

Limits

This is a demonstration: a small language model driven through a text interface to play a toy game. It is not a general game-playing agent, and it is not evidence about Reducer's pruning quality. The teacher is a trivial rule; the model imitates it imperfectly.

Files

  • pong_q4.onnx: the model. Input ids (int64, [1, 16]); output move_logits [1, 3].
  • pong_tokens.json: word → token-id pieces (in the sliced vocabulary) used to build ids.
  • EVAL.json, FINETUNE.json: raw evaluation and training logs.
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading