rlmath agentic GRPO checkpoints

Intermediate checkpoints from agentic GRPO runs on rlmath, a collection of math construction and optimization environments with deterministic verifiers and no LLM judge, packaged as Harbor tasks. The agent is Terminus-2. It works in a sandboxed terminal: it writes and runs code, then submits a construction that a programmatic grader scores.

Each folder is a standalone Hugging Face model directory with safetensors weights, config, tokenizer and chat template. It loads with from_pretrained(<repo>, subfolder="<folder>"). Optimizer states are not included.

Common setup

  • Trainer: verl (fully-async policy, partial rollout, staleness 0.5) with the Alibaba agentic recipe (remote agent loop and LLM proxy). vLLM rollouts, FSDP training.
  • Algorithm: GRPO, lr 1e-6, clip 0.2 / 0.28, no KL loss, token-mean loss aggregation, temperature 1.0.
  • Reward: construct tasks give 1/0. Optimize tasks give 0 if the submission is invalid, otherwise 0.1 + 0.9 · clipped progress toward the target.
  • Validation: 139 tasks, one per family, temperature 0.6, top-p 0.95. Caveat: that set was sampled from the training pool, and 93 of the 139 tasks are also in the 1,145-task training set used from rlmath_async_v2 onward. The validation numbers below therefore partly measure training tasks, not held-out generalization.

Checkpoints

folder base run settings validation (contaminated, see above)
qwen3-4b-thinking/rlmath_q4bthink_async_v1/step_25 Qwen3-4B-Thinking-2507 16 prompts × 8 rollouts, 8 turns, 8k tokens/call, 32k total, 3,401 tasks step 25: 0.235
qwen3-4b-thinking/rlmath_q4bthink_async_v1/step_50 〃 〃 —
qwen3-4b-thinking/rlmath_async_v2/step_25 v1 step 25 1,145 non-trivial tasks, overlong filtering, entropy bonus 0.001 step 0: 0.235 → step 25: 0.134
qwen3-4b-thinking/rlmath_4b_kt_v1/step_36, step_48 Qwen3-4B-Thinking-2507 16 × 32, 12 turns, 16k/call, 64k total, 1,145 tasks 0.054 (attempt 1, step 0) / 0.227 / 0.294 / 0.299 / 0.306 at steps 0 / 12 / 24 / 36 / 48
qwen3-30b-a3b-thinking/rlmath_30b_v1/step_6 Qwen3-30B-A3B-Thinking-2507 16 × 8, 8 turns, 12k/call, 32k total, 1,145 tasks step 0: 0.341
qwen3-30b-a3b-thinking/rlmath_30b_mlxp_v1/step_36, step_48 Qwen3-30B-A3B-Thinking-2507 16 × 32, 12 turns, 16k/call, 64k total, 1,145 tasks 0.331 / 0.372 / 0.447 / 0.521 / 0.503 at steps 0 / 12 / 24 / 36 / 48

Known issue in the earlier 4B runs

rlmath_q4bthink_async_v1 and rlmath_async_v2 were trained with a train/rollout context mismatch. Qwen3-Thinking's chat template drops the reasoning of earlier assistant turns, so turns after the first were sampled without that reasoning in context but trained with it (vLLM-vs-trainer KL ≈ 0.045). Both runs degraded after about 25 steps. All later runs (rlmath_4b_kt_v1, rlmath_30b_*) prompt every turn with exactly the token sequence the trainer sees, which brings the KL down to about 1e-3.

The 30B runs use bf16 weights with torchao's AdamW (bf16 stochastic rounding) under FSDP1.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for amphora/rlmath-agentic-grpo-checkpoints

Finetuned
(44)
this model