Qwen2.5-Math-1.5B-Base-GRPO

Built with Qwen. This is the final base → GRPO checkpoint from the OpenR1 token-level study. GRPO started directly from the original Qwen/Qwen2.5-Math-1.5B weights at revision 4a83ca6e4526a4f2da3aa259ec36c259f66b2ab2. No supervised fine-tuning stage preceded this run. The initial snapshot was checked against the public upstream file checksums. This repository contains the exact saved final weights, tokenizer and generation configuration.

Completion and learning signal

The run completed 1,000 optimizer updates, final fixed-probe evaluation, and the independent selected-token probability check. Of those updates, 863 had a nonzero gradient norm and 137 had a zero gradient norm. Completion does not establish improved reasoning performance. Zero-gradient updates can still move weights through Adam momentum or weight decay. The full recorded metrics and learning-signal counts are retained for analysis.

Training recipe

Setting Value
Framework Prime RL 0.9.0
Prime revision ab5de8fff44b2c4a5c85e24b6e6e3f7d57eee7b1
Dataset zbeeb/Staleness-GRPO-DAPO-Math-17k
Dataset revision 53064564abf94eac096877a61d63e92ac4217433
Dataset rows 17,005
Updates 1,000
Batch 32 responses: four prompts × eight responses
Learning rate 1e-06
Scheduler 30-update warmup, then constant
Optimizer AdamW, weight decay 0.01, gradient clipping at norm 1
Loss PPO ratio clipping at 0.2
Advantage Reward minus group mean, without standard-deviation normalization
Reference-policy KL penalty None
Zero-advantage groups Retained; they contribute zero policy-loss gradient
Maximum off-policy age 8 updates
Sampling Temperature 1, top-p 1
Response limit 3,072 tokens
Total context 4,096 tokens
Seed 42
Saved tensor dtypes F32

training-config.json records the resolved settings. Environment variables, headers and credential values are excluded, and cluster-specific absolute paths are replaced by their basenames.

Use

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "zbeeb/Qwen2.5-Math-1.5B-Base-GRPO"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id, dtype=torch.bfloat16, device_map="auto")
messages = [
    {"role": "system", "content": "Please reason step by step, and put your final answer within \\boxed{}."},
    {"role": "user", "content": "Solve 2x + 3 = 11."},
]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=3072, do_sample=True, temperature=1.0, top_p=1.0, eos_token_id=[151643, 151645], pad_token_id=151643)
print(tokenizer.decode(outputs[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Included measurements

metrics.jsonl.gz preserves the unfiltered native trainer, orchestrator, inference and benchmark metrics. study-metrics-rank-*.jsonl contains the trained-token counts, independently checked probabilities, gradient norms and memory measurements. evaluation-summary.jsonl contains the unchanged OpenR1 train/held-out reference-probe summaries at RL step 0 and every 100 updates through 1,000, using teacher forcing and free generation. Probe labels describe membership in the original SFT population; they are not a new split of the RL dataset.

These repositories publish checkpoints and compact run logs. The larger per-token Parquet exports, raw rollout traces and system logs remain preserved in the experiment's persistent storage. No full-vocabulary logit table or optimizer recovery state is included. provenance.json, learning-signal.json and export-manifest.json record the input checkpoint checksums, training arm, completion evidence, source revisions, artifact sizes and SHA-256 hashes.

License and attribution

The original model license is retained verbatim in LICENSE. Notice records upstream attribution and the modifications from SFT and GRPO. This model is distributed under the upstream model license.

Downloads last month
382
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zbeeb/Qwen2.5-Math-1.5B-Base-GRPO

Finetuned
(333)
this model

Dataset used to train zbeeb/Qwen2.5-Math-1.5B-Base-GRPO

Collection including zbeeb/Qwen2.5-Math-1.5B-Base-GRPO