Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Paper • 2607.14614 • Published 12 days ago • 12
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation Paper • 2607.13124 • Published 14 days ago • 19
Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Paper • 2607.14431 • Published 13 days ago • 12
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals Paper • 2607.11505 • Published 15 days ago • 18
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 252
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents Paper • 2605.17933 • Published May 18 • 7
Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers Paper • 2605.06169 • Published May 7 • 238