Qwen2.5-Math-1.5B-OpenR1-SFT

Built with Qwen. This is the final supervised fine-tuning checkpoint from a token-level study of the transition from SFT to GRPO. It was trained from Qwen/Qwen2.5-Math-1.5B at revision 4a83ca6e4526a4f2da3aa259ec36c259f66b2ab2. The model weights were modified through supervised fine-tuning; tokenizer and generation settings are the exact saved experiment artifacts.

Training

The SFT stage completed 278 optimizer updates and 20,016 examples, corresponding to one epoch over the selected subset of open-r1/OpenR1-Math-220k at revision e4e141ec9dea9f8326f4d347be56105859b2bd68.

The shared data selection uses complete, locally verified mathematical reasoning traces that fit every model's 4,096-token budget. Overlength examples were excluded rather than truncated. Prompt tokens are masked from the supervised loss. The repository includes the exact selected example IDs and probe IDs in data-selection.json.

Setting Value
Framework Prime RL 0.9.0
Prime RL revision ab5de8fff44b2c4a5c85e24b6e6e3f7d57eee7b1
Seed 42
Sequence length 4096 tokens
Effective batch 72 examples
Learning rate 2e-05
Scheduler cosine
Warmup ratio 0.03
Weight decay 0.01
Gradient clipping 1.0
Training precision bf16
Saved tensor dtypes F32

The independent cross-entropy calculation passed and the final fixed-probe evaluation completed. train_metrics.jsonl contains the training update metrics; evaluation-summary.jsonl contains the fixed train and held-out probe summaries, including teacher-forced and free-generation measurements. These are experiment probe results, not a claim of standard benchmark performance.

Use

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "zbeeb/Qwen2.5-Math-1.5B-OpenR1-SFT"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id, dtype=torch.bfloat16, device_map="auto")
messages = [
    {"role": "system", "content": 'Please reason step by step, and put your final answer within \\boxed{}.'},
    {"role": "user", "content": "Solve 2x + 3 = 11."},
]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=3072, do_sample=True, temperature=0.6, top_p=0.95, eos_token_id=[151643, 151645], pad_token_id=151643)
print(tokenizer.decode(outputs[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Study collection and subsequent GRPO stage

This repository contains final SFT weights. Subsequent GRPO checkpoints are published separately in the same study collection after their runs finish. The subsequent RL dataset is zbeeb/Staleness-GRPO-DAPO-Math-17k.

provenance.json, training-config.json, and export-manifest.json record the source model and dataset revisions, training recipe, completion evidence, artifact sizes, and SHA-256 checksums. Optimizer states and recovery checkpoints are not included.

License and attribution

The upstream license is retained verbatim in LICENSE; attribution and modification notices are in Notice. This fine-tuned model is distributed under the same upstream model license.

Downloads last month
148
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zbeeb/Qwen2.5-Math-1.5B-OpenR1-SFT

Finetuned
(334)
this model
Finetunes
1 model

Dataset used to train zbeeb/Qwen2.5-Math-1.5B-OpenR1-SFT

Collection including zbeeb/Qwen2.5-Math-1.5B-OpenR1-SFT