Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning Paper • 2609.37915 • Published 4 days ago • 7
OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning Paper • 2610.02181 • Published 2 days ago • 8
Smaller Models, Better Rejects: Preference Distillation Scaling Paper • 2609.38987 • Published 3 days ago • 10
Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens Paper • 2610.01939 • Published 2 days ago • 26
RPTune: Learned Context Curation for LLM Catalog Search Paper • 2610.00964 • Published 2 days ago • 17
RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers Paper • 2609.34210 • Published 4 days ago • 6
Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents Paper • 2609.32536 • Published 7 days ago • 9
EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery Paper • 2609.40340 • Published 3 days ago • 104
GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis Paper • 2609.38923 • Published 3 days ago • 34
Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation Paper • 2610.00348 • Published 4 days ago • 7
AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines Paper • 2610.01108 • Published 2 days ago • 7
Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs Paper • 2610.01428 • Published 2 days ago • 6
AutoDataBench: A Data-centric Testbed for Accelerating Auto Research Paper • 2609.40097 • Published 3 days ago • 20
Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing Paper • 2609.37362 • Published 4 days ago • 9
A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review Paper • 2609.39027 • Published 3 days ago • 65
Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization Paper • 2610.01017 • Published 2 days ago • 2
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards Paper • 2610.02206 • Published 2 days ago • 5