Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability Paper • 2610.08448 • Published 3 days ago • 163
Using Grounded Theory for Agent Behavior Analysis at Scale Paper • 2608.30391 • Published Aug 31 • 19
CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation Paper • 2609.04083 • Published Sep 3 • 27
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published Sep 3 • 104
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States Paper • 2609.04196 • Published Sep 3 • 71
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published Sep 3 • 188
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions Paper • 2609.04199 • Published Sep 3 • 333
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published Sep 3 • 186
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published Sep 3 • 248
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning Paper • 2605.25604 • Published May 25 • 137