Instructions to use zbeeb/Qwen2.5-Math-1.5B-Base-GRPO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zbeeb/Qwen2.5-Math-1.5B-Base-GRPO with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="zbeeb/Qwen2.5-Math-1.5B-Base-GRPO") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("zbeeb/Qwen2.5-Math-1.5B-Base-GRPO") model = AutoModelForCausalLM.from_pretrained("zbeeb/Qwen2.5-Math-1.5B-Base-GRPO", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use zbeeb/Qwen2.5-Math-1.5B-Base-GRPO with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "zbeeb/Qwen2.5-Math-1.5B-Base-GRPO" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zbeeb/Qwen2.5-Math-1.5B-Base-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/zbeeb/Qwen2.5-Math-1.5B-Base-GRPO
- SGLang
How to use zbeeb/Qwen2.5-Math-1.5B-Base-GRPO with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "zbeeb/Qwen2.5-Math-1.5B-Base-GRPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zbeeb/Qwen2.5-Math-1.5B-Base-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "zbeeb/Qwen2.5-Math-1.5B-Base-GRPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zbeeb/Qwen2.5-Math-1.5B-Base-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use zbeeb/Qwen2.5-Math-1.5B-Base-GRPO with Docker Model Runner:
docker model run hf.co/zbeeb/Qwen2.5-Math-1.5B-Base-GRPO
Qwen2.5-Math-1.5B-Base-GRPO
Built with Qwen. This is the final base → GRPO checkpoint from the OpenR1 token-level study. GRPO started directly from the original Qwen/Qwen2.5-Math-1.5B weights at revision 4a83ca6e4526a4f2da3aa259ec36c259f66b2ab2. No supervised fine-tuning stage preceded this run. The initial snapshot was checked against the public upstream file checksums. This repository contains the exact saved final weights, tokenizer and generation configuration.
Completion and learning signal
The run completed 1,000 optimizer updates, final fixed-probe evaluation, and the independent selected-token probability check. Of those updates, 863 had a nonzero gradient norm and 137 had a zero gradient norm. Completion does not establish improved reasoning performance. Zero-gradient updates can still move weights through Adam momentum or weight decay. The full recorded metrics and learning-signal counts are retained for analysis.
Training recipe
| Setting | Value |
|---|---|
| Framework | Prime RL 0.9.0 |
| Prime revision | ab5de8fff44b2c4a5c85e24b6e6e3f7d57eee7b1 |
| Dataset | zbeeb/Staleness-GRPO-DAPO-Math-17k |
| Dataset revision | 53064564abf94eac096877a61d63e92ac4217433 |
| Dataset rows | 17,005 |
| Updates | 1,000 |
| Batch | 32 responses: four prompts × eight responses |
| Learning rate | 1e-06 |
| Scheduler | 30-update warmup, then constant |
| Optimizer | AdamW, weight decay 0.01, gradient clipping at norm 1 |
| Loss | PPO ratio clipping at 0.2 |
| Advantage | Reward minus group mean, without standard-deviation normalization |
| Reference-policy KL penalty | None |
| Zero-advantage groups | Retained; they contribute zero policy-loss gradient |
| Maximum off-policy age | 8 updates |
| Sampling | Temperature 1, top-p 1 |
| Response limit | 3,072 tokens |
| Total context | 4,096 tokens |
| Seed | 42 |
| Saved tensor dtypes | F32 |
training-config.json records the resolved settings. Environment variables, headers and credential values are excluded, and cluster-specific absolute paths are replaced by their basenames.
Use
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "zbeeb/Qwen2.5-Math-1.5B-Base-GRPO"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id, dtype=torch.bfloat16, device_map="auto")
messages = [
{"role": "system", "content": "Please reason step by step, and put your final answer within \\boxed{}."},
{"role": "user", "content": "Solve 2x + 3 = 11."},
]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=3072, do_sample=True, temperature=1.0, top_p=1.0, eos_token_id=[151643, 151645], pad_token_id=151643)
print(tokenizer.decode(outputs[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Included measurements
metrics.jsonl.gz preserves the unfiltered native trainer, orchestrator, inference and benchmark metrics. study-metrics-rank-*.jsonl contains the trained-token counts, independently checked probabilities, gradient norms and memory measurements. evaluation-summary.jsonl contains the unchanged OpenR1 train/held-out reference-probe summaries at RL step 0 and every 100 updates through 1,000, using teacher forcing and free generation. Probe labels describe membership in the original SFT population; they are not a new split of the RL dataset.
These repositories publish checkpoints and compact run logs. The larger per-token Parquet exports, raw rollout traces and system logs remain preserved in the experiment's persistent storage. No full-vocabulary logit table or optimizer recovery state is included. provenance.json, learning-signal.json and export-manifest.json record the input checkpoint checksums, training arm, completion evidence, source revisions, artifact sizes and SHA-256 hashes.
License and attribution
The original model license is retained verbatim in LICENSE. Notice records upstream attribution and the modifications from SFT and GRPO. This model is distributed under the upstream model license.
- Downloads last month
- 382