False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents Paper • 2609.39102 • Published 6 days ago • 470
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 7 days ago • 552
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents Paper • 2602.02474 • Published Feb 2 • 64
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL Paper • 2610.00574 • Published 6 days ago • 59
EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery Paper • 2609.40340 • Published 6 days ago • 107
ReasonFLux-Coder Collection Coding LLMs excel at both writing code and generating unit tests. • 9 items • Updated 17 days ago • 12
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 15 days ago • 222
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published Sep 1 • 567
The Verification Horizon: No Silver Bullet for Coding Agent Rewards Paper • 2606.26300 • Published Jun 24 • 53
StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback Paper • 2402.01391 • Published Feb 2, 2024 • 44
view article Article A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyond karina-zadorozhny • Jan 19 • 52
SFTvsRL Models & Data Collection This collection contains 4 initial checkpoints for https://github.com/LeslieTrue/SFTvsRL and necessary data for V-IRL training. • 6 items • Updated Mar 2 • 10
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training Paper • 2501.17161 • Published Jan 28, 2025 • 128