A New Role for Relevance: Guiding Corpus Interaction in Agentic Search Paper • 2607.24223 • Published 5 days ago • 90
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 18 days ago • 230
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published 9 days ago • 26
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published 10 days ago • 31
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Paper • 2607.18722 • Published 11 days ago • 35
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published 16 days ago • 142
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 23 days ago • 76
Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots Paper • 2606.28133 • Published Jun 26 • 40
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications Paper • 2509.26490 • Published Sep 30, 2025 • 22
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? Paper • 2606.24530 • Published Jun 23 • 65
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published Jun 25 • 57
OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning Paper • 2606.25757 • Published Jun 24 • 1
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 153
GrepSeek: Training Search Agents for Direct Corpus Interaction Paper • 2605.29307 • Published May 28 • 117
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 241