Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published 4 days ago • 20
DSWorld: A Data Science World Model for Efficient Autonomous Agents Paper • 2607.15901 • Published 10 days ago • 12
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 11 days ago • 102
DrugGen 2: A disease-aware language model for enhancing drug discovery Paper • 2607.08404 • Published 18 days ago • 20
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Paper • 2607.04412 • Published 22 days ago • 35
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies Paper • 2607.04434 • Published 20 days ago • 15
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use Paper • 2607.01874 • Published 25 days ago • 22
AgenticDataBench: A Comprehensive Benchmark for Data Agents Paper • 2607.01647 • Published 25 days ago • 37
AutoTrainess: Teaching Language Models to Improve Language Models Autonomously Paper • 2606.31551 • Published 27 days ago • 24
Autonomous Scientific Discovery via Iterative Meta-Reflection Paper • 2607.01131 • Published 26 days ago • 8
Evolution Fine-Tuning Collection Internalizing Discovery Capability into LLM • 10 items • Updated 25 days ago • 1
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks Paper • 2606.29082 • Published 30 days ago • 42
Agentic Abstention: Do Agents Know When to Stop Instead of Act? Paper • 2606.28733 • Published 30 days ago • 149
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks Paper • 2606.29537 • Published 29 days ago • 22
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents Paper • 2606.28480 • Published about 1 month ago • 48
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent Paper • 2606.30616 • Published 28 days ago • 103
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation Paper • 2606.28128 • Published about 1 month ago • 53
Autodata: An agentic data scientist to create high quality synthetic data Paper • 2606.25996 • Published Jun 24 • 18