Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 6 days ago • 281
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 3 days ago • 272
LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling Paper • 2604.11748 • Published Apr 15 • 15
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 18 days ago • 250
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 15 days ago • 191
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 25 days ago • 376
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 24 days ago • 84
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs Paper • 2609.04753 • Published 28 days ago • 16
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation Paper • 2609.05295 • Published 28 days ago • 19
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published 28 days ago • 26
Enoki: Efficient Multi-Level Hallucination Detection Paper • 2609.00581 • Published about 1 month ago • 29
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking Paper • 2403.09629 • Published Mar 14, 2024 • 81
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters Paper • 2408.03314 • Published Aug 6, 2024 • 68
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 29 days ago • 248
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 29 days ago • 104
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 30 days ago • 407
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 29 days ago • 85
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published about 1 month ago • 122
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Paper • 2608.31075 • Published Aug 31 • 29