DeepMath-103K-Level6-Qwen3-4B-Base-GRPO

DeepMath-103K-Level6-Qwen3-4B-Base-GRPO is a post-trained version of Qwen/Qwen3-4B-Base, trained with Group Relative Policy Optimization (GRPO) on the Level 6 subset of DeepMath-103K for mathematical reasoning.

Training used the verl framework and the training code from Thinking-Space/Rethinking-OPD, with full-parameter actor updates and rule-based math outcome rewards.

Dataset

The training data consists of 57,046 examples from the DeepMath Level 6 training file available in Keven16/G-OPD-Training-Data, originating from zwhe99/DeepMath-103K.

Training Details

Training configuration

  • Base model: Qwen/Qwen3-4B-Base
  • Training framework: verl
  • Algorithm: GRPO
  • Parameter update: Full-parameter fine-tuning
  • Rollout engine: vLLM
  • Context length: 32,768 tokens
  • Responses per prompt: 8
  • GRPO outcome weight: 1.0
  • Prompt length: 1,024 tokens
  • Response length: 7,168 tokens
  • Max model length: 32,768 tokens
  • Rollout temperature: 1.0
  • Rollout top-p / top-k: 1.0 / -1
  • Repetition penalty: 1.0
  • KL loss: Disabled
  • Format reward: Disabled
  • Learned reward model: Disabled
  • Loss aggregation: token-mean
  • Learning rate: 1e-6
  • Learning-rate schedule: Constant; no warmup
  • Weight decay: 0.01
  • PPO mini-batch size: 64
  • PPO micro-batch size per GPU: 1
  • Number of GPUs: 4
  • Number of epochs: 1
  • Save frequency: Every 20 steps
  • Test frequency: Every 20 steps
  • Validation sampling: 16 responses per prompt; temperature 1.0; top-p 0.95
  • Validation response length: 31,744 tokens

Dataset

  • Training dataset: DeepMath-103K Level 6
  • Training examples: 57,046
  • Training-time validation datasets: AIME25, AMC22–23, AIME24
  • Validation questions: 143

Validation Accuracy

  • AIME 2025 (avg@16): 21.25%
  • AMC 2022–2023 (avg@16): 60.77%
  • AIME 2024 (avg@16): 23.54%

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "ExploreXploitQ/DeepMath-103K-Level6-Qwen3-4B-Base-GRPO"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    dtype="auto",
    device_map="auto",
)
Downloads last month
40
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ExploreXploitQ/DeepMath-103K-Level6-Qwen3-4B-Base-GRPO

Finetuned
(488)
this model

Datasets used to train ExploreXploitQ/DeepMath-103K-Level6-Qwen3-4B-Base-GRPO