AegisStep-28M

A ~28M-parameter Process Reward Model trained from scratch. It checks each step of a math solution and says whether it is correct or incorrect.

Built with: Sparse Mixture-of-Experts (4 experts, top-2), Grouped-Query Attention, RoPE, SwiGLU.

Results (10,000 held-out test chains)

Metric Value
AUROC 0.773
Balanced accuracy 69.99%
Always-guess baseline 50.54%
Full-chain accuracy 39.97%

Training

150k Math-Shepherd chains, split by question (no leakage), 3 epochs, AdamW, cosine schedule, dropout 0.1. The best checkpoint was chosen by validation loss.

Limitations

This is a small research baseline, not a reliable verifier. Labels come from Math-Shepherd's automatic labeling, and chains longer than 512 tokens are cut off.

Citation

Math-Shepherd: Wang et al., 2023, arXiv:2312.08935

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train saqiibb/AegisStep-28M

Paper for saqiibb/AegisStep-28M