CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning Paper • 2609.36820 • Published 11 days ago • 40
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI Paper • 2609.38143 • Published 11 days ago • 85
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation Paper • 2608.30730 • Published Aug 31 • 18
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations Paper • 2608.15930 • Published Aug 16 • 47
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published Aug 12 • 118
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design Paper • 2608.10299 • Published Aug 10 • 138
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design Paper • 2608.10299 • Published Aug 10 • 138
Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty Paper • 2508.08992 • Published Apr 10
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration? Paper • 2510.24505 • Published Oct 28, 2025 • 5
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents Paper • 2511.02734 • Published Nov 4, 2025 • 23
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published Jul 19 • 93
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty Paper • 2412.20251 • Published May 25, 2025
SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision Paper • 2606.01139 • Published Jun 2 • 1
AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora Paper • 2505.23628 • Published May 29, 2025
NAACL: Noise-AwAre Verbal Confidence Calibration for LLMs in RAG Systems Paper • 2601.11004 • Published Jan 16 • 31
AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility Estimation Paper • 2402.10646 • Published Feb 16, 2024
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published Jul 19 • 93