Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens Paper • 2610.01939 • Published 11 days ago • 52
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering Paper • 2609.19879 • Published 25 days ago • 33
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 27 days ago • 78
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published Sep 7 • 376
CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation Paper • 2609.06931 • Published Sep 7 • 27
Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions Paper • 2608.29109 • Published Aug 29 • 17
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Paper • 2609.09153 • Published Sep 8 • 44