Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Paper • 2606.11042 • Published Jun 9 • 221
Watch Before You Answer: Learning from Visually Grounded Post-Training Paper • 2604.05117 • Published Apr 6 • 162
Learning from the Self-future: On-policy Self-distillation for dLLMs Paper • 2606.18195 • Published Jun 16 • 175
Post-Training Leaves Behavioral Shadows on Unrelated Decisions Paper • 2609.29233 • Published 15 days ago • 274
Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue Paper • 2609.31948 • Published 14 days ago • 89
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published Sep 1 • 569
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published Aug 26 • 337
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published Aug 17 • 52
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 149
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 50
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop Paper • 2605.18746 • Published May 18 • 8
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence Paper • 2605.12882 • Published May 13 • 65
Position: LLM Inference Should Be Evaluated as Energy-to-Token Production Paper • 2605.11733 • Published May 12 • 4
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation Paper • 2604.28196 • Published Apr 30 • 27
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music Paper • 2604.10905 • Published Apr 13 • 29