Learning from the Self-future: On-policy Self-distillation for dLLMs Paper • 2606.18195 • Published Jun 16 • 147
Post-Training Leaves Behavioral Shadows on Unrelated Decisions Paper • 2609.29233 • Published 9 days ago • 271
Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue Paper • 2609.31948 • Published 8 days ago • 89
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published Sep 1 • 294
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published Aug 26 • 188
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published Aug 17 • 52