TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning
Abstract
Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-gradient direction, distinguishing locally amplifying from contracting update contributions that mismatch magnitude alone cannot identify. In native NVFP4 runs, we observe an early imbalance between the two amplifying regions, favoring negative-advantage, negative-gap updates. Their tail tokens become concentrated in a small fraction of response segments before mismatch spreads globally. Motivated by these findings, we introduce TRIAGE, a direction-aware stabilization method that uses segment-level diagnosis to selectively rebalance policy-gradient updates and applies bounded repair to residual severe mismatch. TRIAGE modifies the optimization objective while retaining native NVFP4 weight-and activation 4-bit (W4A4) forward execution on both the sampler and learner. Experiments on Qwen3-4B and Qwen3-30B-A3B show stable optimization throughout the evaluated training horizon and achieve full precision level performance across five mathematical reasoning benchmarks, while native NVFP4 with TRIAGE provides up to 2.3x higher rollout throughput than BF16.
Community
In this work, we ask why native NVFP4 RL can become unstable even when learner–sampler mismatch is controlled by standard magnitude-based corrections. We find that mismatch magnitude alone is not enough to characterize the risk: what also matters is whether the policy-gradient update contracts or amplifies the existing mismatch. In both Qwen3-4B and Qwen3-30B-A3B, a directional imbalance emerges before the global mismatch distribution becomes severe, with mismatch-amplifying negative-advantage tokens becoming concentrated in a small fraction of response segments. Motivated by this observation, we introduce TRIAGE, which performs segment-level diagnosis, selectively attenuates mismatch-amplifying updates, and applies bounded repair to residual severe mismatch. Importantly, TRIAGE preserves native NVFP4 W4A4 execution on both the sampler and learner, rather than replacing it with fake quantization. Across our experiments, TRIAGE maintains stable optimization and near-BF16 performance while achieving up to 2.30× rollout throughput and 1.30× end-to-end RL speedup over BF16. We hope this work provides a useful perspective for understanding and stabilizing execution mismatch in low-precision LLM RL.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Score Centering Stabilizes Off-policy Reinforcement Learning (2026)
- Towards Full Pipeline FP8 Reinforcement Learning for LLMs (2026)
- Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It (2026)
- CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning (2026)
- RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning (2026)
- SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL (2026)
- Teach to Learn: Hint Annealing for Self-improving LLM Reasoning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.07043 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper