WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents Paper • 2609.36887 • Published 8 days ago • 21
Towards Looped Models Done Right, Part II: Rethinking at Fixed Points Paper • 2610.06833 • Published 2 days ago • 20
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models Paper • 2605.07721 • Published May 8 • 31
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 11 days ago • 324
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 8 days ago • 578
LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling Paper • 2604.11748 • Published Apr 15 • 15
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 23 days ago • 251
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 20 days ago • 224
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published about 1 month ago • 376
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 29 days ago • 84
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs Paper • 2609.04753 • Published Sep 4 • 16
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation Paper • 2609.05295 • Published Sep 4 • 19
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published Sep 4 • 26
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking Paper • 2403.09629 • Published Mar 14, 2024 • 81
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters Paper • 2408.03314 • Published Aug 6, 2024 • 68
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published Sep 3 • 248
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published Sep 3 • 104
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published Sep 2 • 408