Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts Paper • 2610.01153 • Published 4 days ago • 3
Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling Paper • 2609.36529 • Published 6 days ago • 3
Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems Paper • 2609.39050 • Published 5 days ago • 3
Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It Paper • 2610.03195 • Published 3 days ago • 12
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite Paper • 2610.02826 • Published 3 days ago • 12
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts Paper • 2610.00314 • Published 6 days ago • 90
Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding Paper • 2609.32019 • Published 10 days ago • 51
ROWBench: Do Video Models Render What the Program Specifies? Paper • 2610.02205 • Published 4 days ago • 64
OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories Paper • 2609.32810 • Published 9 days ago • 14
DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration Paper • 2609.33403 • Published 8 days ago • 15
The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends Paper • 2609.39661 • Published 5 days ago • 10
Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations Paper • 2609.32522 • Published 9 days ago • 81
Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior Paper • 2609.39827 • Published 5 days ago • 11
CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning Paper • 2609.36820 • Published 6 days ago • 31
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL Paper • 2610.00574 • Published 5 days ago • 57
LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models Paper • 2609.32264 • Published 9 days ago • 46
DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation Paper • 2609.33711 • Published 8 days ago • 13
Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation Paper • 2609.39687 • Published 5 days ago • 16
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 6 days ago • 468