Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position Paper • 2610.10114 • Published 1 day ago • 28
ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience Paper • 2610.05303 • Published 5 days ago • 23
Self-Supervised Scaling of Terminal Environments for Scientific Domains Paper • 2610.02710 • Published 7 days ago • 12
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 22 days ago • 57
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published Sep 7 • 19