MiniWorld: Democratizing the Training of Video World Models from Scratch Paper • 2608.01127 • Published 10 days ago • 19
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? Paper • 2608.03874 • Published 8 days ago • 14
UniWorld-Design: From Pixel Generation to Layer-Native Design Paper • 2608.03971 • Published 8 days ago • 21
Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts Paper • 2608.00574 • Published 11 days ago • 8
Flux-OPD: On-Policy Distillation with Evolving Contexts Paper • 2607.28022 • Published 13 days ago • 44
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Paper • 2605.09635 • Published 20 days ago • 64
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published 25 days ago • 142
VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery Paper • 2607.06374 • Published Jul 7 • 10
GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning Paper • 2606.17480 • Published Jun 16 • 3
DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects Paper • 2606.15133 • Published Jun 13 • 74
MotionVLA: Vision-Language-Action Model for Humanoid Motion Paper • 2606.15142 • Published Jun 13 • 5
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Paper • 2606.07433 • Published Jun 5 • 21
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Paper • 2606.06155 • Published Jun 4 • 10
LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing Paper • 2606.06042 • Published Jun 4 • 24
OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning Paper • 2605.28691 • Published May 27 • 25
RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting Paper • 2605.18263 • Published May 18 • 9
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Paper • 2605.18287 • Published May 18 • 15
GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding Paper • 2605.15250 • Published May 14 • 14
VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction Paper • 2605.15186 • Published May 14 • 26