rlmath agentic GRPO checkpoints
Intermediate checkpoints from agentic GRPO runs on rlmath, a collection of math construction and optimization environments with deterministic verifiers and no LLM judge, packaged as Harbor tasks. The agent is Terminus-2. It works in a sandboxed terminal: it writes and runs code, then submits a construction that a programmatic grader scores.
Each folder is a standalone Hugging Face model directory with safetensors weights, config, tokenizer and chat template. It loads with from_pretrained(<repo>, subfolder="<folder>"). Optimizer states are not included.
Common setup
- Trainer: verl (fully-async policy, partial rollout, staleness 0.5) with the Alibaba agentic recipe (remote agent loop and LLM proxy). vLLM rollouts, FSDP training.
- Algorithm: GRPO, lr 1e-6, clip 0.2 / 0.28, no KL loss, token-mean loss aggregation, temperature 1.0.
- Reward: construct tasks give 1/0. Optimize tasks give 0 if the submission is invalid, otherwise 0.1 + 0.9 · clipped progress toward the target.
- Validation: 139 tasks, one per family, temperature 0.6, top-p 0.95. Caveat: that set was sampled from the training pool, and 93 of the 139 tasks are also in the 1,145-task training set used from
rlmath_async_v2onward. The validation numbers below therefore partly measure training tasks, not held-out generalization.
Checkpoints
| folder | base | run settings | validation (contaminated, see above) |
|---|---|---|---|
qwen3-4b-thinking/rlmath_q4bthink_async_v1/step_25 |
Qwen3-4B-Thinking-2507 | 16 prompts × 8 rollouts, 8 turns, 8k tokens/call, 32k total, 3,401 tasks | step 25: 0.235 |
qwen3-4b-thinking/rlmath_q4bthink_async_v1/step_50 |
〃 | 〃 | — |
qwen3-4b-thinking/rlmath_async_v2/step_25 |
v1 step 25 | 1,145 non-trivial tasks, overlong filtering, entropy bonus 0.001 | step 0: 0.235 → step 25: 0.134 |
qwen3-4b-thinking/rlmath_4b_kt_v1/step_36, step_48 |
Qwen3-4B-Thinking-2507 | 16 × 32, 12 turns, 16k/call, 64k total, 1,145 tasks | 0.054 (attempt 1, step 0) / 0.227 / 0.294 / 0.299 / 0.306 at steps 0 / 12 / 24 / 36 / 48 |
qwen3-30b-a3b-thinking/rlmath_30b_v1/step_6 |
Qwen3-30B-A3B-Thinking-2507 | 16 × 8, 8 turns, 12k/call, 32k total, 1,145 tasks | step 0: 0.341 |
qwen3-30b-a3b-thinking/rlmath_30b_mlxp_v1/step_36, step_48 |
Qwen3-30B-A3B-Thinking-2507 | 16 × 32, 12 turns, 16k/call, 64k total, 1,145 tasks | 0.331 / 0.372 / 0.447 / 0.521 / 0.503 at steps 0 / 12 / 24 / 36 / 48 |
Known issue in the earlier 4B runs
rlmath_q4bthink_async_v1 and rlmath_async_v2 were trained with a train/rollout context mismatch. Qwen3-Thinking's chat template drops the reasoning of earlier assistant turns, so turns after the first were sampled without that reasoning in context but trained with it (vLLM-vs-trainer KL ≈ 0.045). Both runs degraded after about 25 steps. All later runs (rlmath_4b_kt_v1, rlmath_30b_*) prompt every turn with exactly the token sequence the trainer sees, which brings the KL down to about 1e-3.
The 30B runs use bf16 weights with torchao's AdamW (bf16 stochastic rounding) under FSDP1.
Model tree for amphora/rlmath-agentic-grpo-checkpoints
Base model
Qwen/Qwen3-30B-A3B-Thinking-2507