X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published 21 days ago • 48
VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis Paper • 2609.33253 • Published 4 days ago • 4
Index-Translate Collection Index‐Translate: A Multilingual Translation Model Family Text, Speech, Controlled Dubbing, and Long‐Document Translation • 11 items • Updated about 11 hours ago • 11
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 2 days ago • 66
InternVLA-A1.5 Collection The collection includes the pretrained and finetuned checkpoints of InternVLA-A1.5 • 4 items • Updated Jul 8 • 8
InternW0-Δ Collection The collection includes the pretrained and finetuned checkpoints of InternW0-Δ • 4 items • Updated 3 days ago • 4
view article Article Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents MultiverseComputingCAI • 1 day ago • 11
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 5 days ago • 18
CompoWorld: Compositional Environment Scaling for General Agents Paper • 2609.33665 • Published 4 days ago • 30
Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models Paper • 2609.34972 • Published 3 days ago • 20
Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence Paper • 2609.35432 • Published 3 days ago • 99
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 3 days ago • 38
FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance Paper • 2609.25716 • Published 9 days ago • 20