SAF-OPD: Stable Advantage Fusion for On-Policy Distillation Paper • 2607.29209 • Published 5 days ago • 29
Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge Paper • 2608.01862 • Published 2 days ago • 3
In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing Paper • 2607.15820 • Published 19 days ago • 4
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Paper • 2607.26326 • Published 8 days ago • 2
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents Paper • 2608.01827 • Published 2 days ago • 15
RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems Paper • 2607.29241 • Published 5 days ago • 7
To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing Paper • 2607.28887 • Published 6 days ago • 18
EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents Paper • 2607.28229 • Published 6 days ago • 7
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning Paper • 2607.28478 • Published 6 days ago • 3
Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI Paper • 2608.01462 • Published 3 days ago • 3
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning Paper • 2608.02585 • Published 2 days ago • 22
Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV Paper • 2607.23693 • Published 10 days ago • 3
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Paper • 2607.28802 • Published 6 days ago • 9