WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 6 days ago • 155
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms Paper • 2609.23658 • Published 7 days ago • 29
All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts Paper • 2609.24058 • Published 6 days ago • 52
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 5 days ago • 32
Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies Paper • 2609.24682 • Published 6 days ago • 12
EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation Paper • 2609.24115 • Published 6 days ago • 5
FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations Paper • 2609.20817 • Published 10 days ago • 39
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering Paper • 2609.19879 • Published 10 days ago • 32
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation Paper • 2609.20744 • Published 10 days ago • 53
JEPA-Anything: Learning Predictive Models across Different Worlds Paper • 2609.20800 • Published 10 days ago • 73
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model Paper • 2609.18323 • Published 11 days ago • 132
WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing Paper • 2609.20423 • Published 10 days ago • 48
Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control Paper • 2609.17909 • Published 12 days ago • 46