--- language: - en license: apache-2.0 tags: - llm - pytorch - causal-lm - rune-r1 - reasoning - grpo - rlvr datasets: - HuggingFaceFW/fineweb-edu - rasbt/math_distill metrics: - accuracy pipeline_tag: text-generation base_model: - samueljayasingh/rune-0.3b-base - samueljayasingh/rune-0.3b-sft --- # Rune-R1 (351M) — GRPO Reasoning Model **Rune-R1** is a ~351M parameter decoder-only transformer trained from scratch and aligned for math reasoning via a three-stage pipeline: ``` Pretrain (FineWeb-Edu) → SFT (distilled CoT format) → GRPO (RLVR on math correctness) ``` This repository holds the final checkpoint: the GRPO-tuned policy, starting from [Rune-R1-SFT](https://huggingface.co/samueljayasingh/Rune-R1-sft) and optimized with Group Relative Policy Optimization against a verifiable, rule-based reward for math answer correctness. See [Rune-R1-Base](https://huggingface.co/samueljayasingh/Rune-R1-base) and [Rune-R1-SFT](https://huggingface.co/samueljayasingh/Rune-R1-sft) for the earlier pipeline stages. ## Model Description | | | |---|---| | **Developed by** | samueljayasingh | | **Model type** | Causal language model (text-only) | | **Base model** | [Rune-R1-SFT](https://huggingface.co/samueljayasingh/Rune-R1-sft) (351M, chain-of-thought SFT on top of Rune-R1-Base) | | **Fine-tuning method** | GRPO (Group Relative Policy Optimization) with PPO-style clipping and a KL penalty to a frozen reference (RLVR — reinforcement learning from verifiable rewards) | | **Dataset** | `data/math_train.json` (math word problems with verifiable final answers), evaluated on a 50-example MATH-500 held-out subset | | **Language** | English | | **Tokenizer** | GPT-2 (`tiktoken`) | | **License** | Apache 2.0 | ### Architecture Details | Parameter | Value | |---|---| | Layers | 22 | | Embedding dimension | 1024 | | Attention heads / KV groups | 16 / 4 (GQA) | | Feed-forward hidden dim | 2816 (SwiGLU) | | Context length | 1024 tokens | | Position embeddings | RoPE (base 10,000) | | Normalization | RMSNorm, with QK normalization | ## Intended Uses & Limitations ### Intended Use - Research into RLVR / GRPO-style reasoning fine-tuning at small model scale. - Reference implementation for reward-verified RL post-training pipelines (pretrain → SFT → RL). - Studying reward hacking, KL-regularization tradeoffs, and reasoning-accuracy dynamics under a small RL step budget. ### Limitations - Small model (351M) with a limited RL budget (2,000 steps) — MATH-500 accuracy remains low (0–4% across evaluation checkpoints; see table below) and should not be compared to production-scale reasoning models. - Reward signal is a rule-based correctness check (`\boxed{}` extraction + symbolic grading), so the model may still learn to produce well-formatted but incorrect reasoning that occasionally reward-hacks the verifier. - Inherits base/SFT limitations: 1024-token context, ~5B pretraining tokens, no broad safety/RLHF alignment beyond the math-correctness reward. - **Not suitable for production or user-facing deployment** — this is a research artifact demonstrating the training pipeline, not a competitive reasoning model. ## How to Use ```python import torch import tiktoken from rune.model import CONFIG_350M, RuneModel ckpt = torch.load("pytorch_model.bin", map_location="cpu") model = RuneModel(CONFIG_350M) model.load_state_dict(ckpt) model.eval() enc = tiktoken.get_encoding("gpt2") prompt = "What is 12 * 15?" tokens = torch.tensor([enc.encode(prompt)], dtype=torch.long) # Model responds in "...reasoning...\n\n\\boxed{final_answer}" format. # See rune/generate.py in the source repo for full sampling / KV-cache generation code. ``` The `rune` package (model definition + generation utilities) is available at the Rune-R1 GitHub repository. ## Hardware Trained end to end — pretraining, SFT, and GRPO — on a single rented GPU instance: | Component | Spec | |---|---| | GPU | 1x AMD MI300X | | VRAM | 192 GB | | vCPU | 20 | | RAM | 240 GB | | Boot disk | 720 GB NVMe SSD | | Scratch disk | 5 TB NVMe SSD | | Rate | $1.99/hr | ## Training & Evaluation ### Training Procedure | Parameter | Value | |---|---| | Starting checkpoint | Rune-R1-SFT | | Reference model | Frozen copy of the SFT checkpoint (KL penalty target) | | Training steps | 2,000 | | Rollouts per prompt (group size) | 8 | | Inner epochs per rollout batch | 2 | | Max new tokens (rollout) | 512 | | Sampling temperature / top-p | 0.8 / 0.9 | | PPO clip epsilon | 10.0 | | KL coefficient | 0.001 | | Learning rate | 1e-6 | | Reward function | Rule-based: extract `\boxed{}` answer, symbolically grade vs. ground truth (1.0 / 0.0) | | Eval cadence | MATH-500 (50-example subset), every 100 steps | ### Evaluation Results | Metric | Value | |---|---| | Final MATH-500 accuracy (step 2000) | 0% (50 examples) | | Peak MATH-500 accuracy | 4% (steps 1600, 1900) | | Mean reward per step (over training) | ~0.016 | | Max single-step average reward | 0.75 | | Steps with nonzero reward | 152 / 2001 | MATH-500 accuracy fluctuated in the 0–4% range throughout training rather than improving monotonically, reflecting the small model size and limited RL budget rather than a fully converged reasoning model. ## Citation ```bibtex @misc{RuneR12026, author = {Samuel Jayasingh}, title = {Rune-R1: A 351M Transformer Reasoning Model Trained via Pretrain-SFT-GRPO from Scratch}, year = {2026}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/samueljayasingh/Rune-R1}} } ``` ## Acknowledgements - [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) — pretraining corpus. - [rasbt/math_distill](https://huggingface.co/datasets/rasbt/math_distill) — distilled chain-of-thought SFT data. - [rasbt/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch) — architecture and the pretrain → SFT → GRPO reasoning-from-scratch recipe this pipeline is adapted from.