-
DSDR: Dual-Scale Diversity Regularization for Exploration in LLM Reasoning
Paper • 2602.19895 • Published • 14 -
SetPO: Set-Level Policy Optimization for Diversity-Preserving LLM Reasoning
Paper • 2602.01062 • Published • 2 -
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
Paper • 2602.06717 • Published • 76 -
MC-GRPO: Median-Centered Group Relative Policy Optimization for Small-Rollout Reinforcement Learning
Paper • 2601.22582 • Published
Lei Xia
cszdwxm
AI & ML interests
None yet
Recent Activity
upvoted a paper 30 days ago
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation upvoted a paper about 2 months ago
Rethinking the Divergence Regularization in LLM RL updated a collection 5 months ago
XXPO/XXRL