Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL Paper • 2610.00574 • Published 11 days ago • 65
Codifying the Judge: Scalable Evaluation via Program Distillation Paper • 2607.22561 • Published May 29 • 9
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published Jul 9 • 79
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents Paper • 2604.10577 • Published Apr 12 • 27