Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL Paper • 2610.00574 • Published 2 days ago • 20
EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery Paper • 2609.40340 • Published 2 days ago • 98
Recursive Synthesis for Long-Horizon Terminal Tasks Collection CC BY 4.0 datasets and SFT/RL checkpoints for recursive task synthesis; base-model and third-party terms still apply. • 6 items • Updated Aug 10 • 18
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published Jul 9 • 77
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL Paper • 2604.17073 • Published Apr 18 • 9