Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX Paper • 2609.18011 • Published 22 days ago • 30
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents Paper • 2609.05903 • Published Sep 5 • 65
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation Paper • 2609.11486 • Published 28 days ago • 34
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Paper • 2609.09153 • Published 30 days ago • 44
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests Paper • 2608.27831 • Published Aug 31 • 33
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published Aug 27 • 155
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published Sep 2 • 408
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published Aug 31 • 316
Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See Paper • 2608.17744 • Published Aug 18 • 16
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 161