DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling Paper • 2610.04933 • Published 8 days ago • 17
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation Paper • 2609.38024 • Published 13 days ago • 64
Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation Paper • 2606.02479 • Published Jun 1 • 25
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems Paper • 2609.08572 • Published Sep 8 • 86
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards Paper • 2605.10899 • Published May 11 • 78
Exploration and Exploitation Errors Are Measurable for Language Model Agents Paper • 2604.13151 • Published Apr 14 • 25
Thinking Makes LLM Agents Introverted: How Mandatory Thinking Can Backfire in User-Engaged Agents Paper • 2602.07796 • Published Feb 8 • 7
Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense Paper • 2510.07242 • Published Oct 8, 2025 • 30