Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 24 days ago • 84
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction Paper • 2609.10715 • Published Sep 9 • 269
WorldReward: Reward Modeling for Camera-Conditioned World Models Paper • 2609.03952 • Published Sep 3 • 28
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents Paper • 2609.08149 • Published Sep 8 • 28
HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models Paper • 2608.13205 • Published Aug 13 • 3
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis Paper • 2608.18580 • Published Aug 19 • 75
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Paper • 2607.13705 • Published Jul 15 • 40
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 167
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games Paper • 2606.19338 • Published Jun 17 • 52
OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs Paper • 2606.03890 • Published Jun 2 • 32
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation Paper • 2605.31264 • Published May 29 • 132
SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Paper • 2605.20110 • Published May 19 • 4
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation Paper • 2605.10912 • Published May 11 • 38
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale Paper • 2603.25040 • Published Mar 26 • 132