Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention Paper • 2605.22791 • Published May 21 • 31
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning Paper • 2604.12374 • Published Apr 14 • 39
PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents Paper • 2609.40285 • Published 9 days ago • 23
PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents Paper • 2609.40285 • Published 9 days ago • 23
PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents Paper • 2609.40285 • Published 9 days ago • 23
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention Paper • 2605.22791 • Published May 21 • 31
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data Paper • 2510.03264 • Published Sep 26, 2025 • 26