magic-run100x (private)

GPT-2 Small core (85.7M params, vocab 276, 512 positions) trained from scratch on ~1B tokens: 900M FineWeb-Edu (byte-level arithmetic-safe tokenizer) + 25M x 4 synthetic arithmetic evidence families (widened-v2 spec: starts 0-999, mult 2-5, addsub 1-99; 300k distinct programs/family). prod_v1 config: batch 64, lr 3e-4 linear decay, warmup 20, wd 0.01, eps_root 1e-2, seed 0, 30,518 steps, single H100. Final checkpoint step_30517 exported via tools/export_checkpoint.py. Held-out (ADD->MULTIPLY) answer loss: 2.33 (chance ~5.62). Project: Stage-Resolved Data Attribution (MathaMAGICal / algoverse-jpdv).

Downloads last month
7
Safetensors
Model size
85.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support