From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders Paper • 2610.02486 • Published 5 days ago • 2
SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation Paper • 2610.02304 • Published 5 days ago • 21
Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers Paper • 2610.00531 • Published 6 days ago • 15
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents Paper • 2610.03574 • Published 4 days ago • 18
Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts Paper • 2610.01153 • Published 5 days ago • 6
Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling Paper • 2609.36529 • Published 7 days ago • 5
Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems Paper • 2609.39050 • Published 6 days ago • 4
Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It Paper • 2610.03195 • Published 4 days ago • 15
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite Paper • 2610.02826 • Published 4 days ago • 69
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts Paper • 2610.00314 • Published 7 days ago • 97
Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding Paper • 2609.32019 • Published 11 days ago • 53
ROWBench: Do Video Models Render What the Program Specifies? Paper • 2610.02205 • Published 5 days ago • 65
OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories Paper • 2609.32810 • Published 10 days ago • 14
DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration Paper • 2609.33403 • Published 9 days ago • 15
The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends Paper • 2609.39661 • Published 6 days ago • 12
Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations Paper • 2609.32522 • Published 10 days ago • 81
Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior Paper • 2609.39827 • Published 6 days ago • 12
CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning Paper • 2609.36820 • Published 7 days ago • 32
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL Paper • 2610.00574 • Published 6 days ago • 59
LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models Paper • 2609.32264 • Published 10 days ago • 47