X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published 22 days ago • 48
VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis Paper • 2609.33253 • Published 5 days ago • 4
Index-Translate Collection Index‐Translate: A Multilingual Translation Model Family Text, Speech, Controlled Dubbing, and Long‐Document Translation • 15 items • Updated about 3 hours ago • 16
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 3 days ago • 86
InternVLA-A1.5 Collection The collection includes the pretrained and finetuned checkpoints of InternVLA-A1.5 • 4 items • Updated Jul 8 • 8
InternW0-Δ Collection The collection includes the pretrained and finetuned checkpoints of InternW0-Δ • 4 items • Updated 4 days ago • 4
view article Article Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents MultiverseComputingCAI • 2 days ago • 16
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 6 days ago • 24
CompoWorld: Compositional Environment Scaling for General Agents Paper • 2609.33665 • Published 5 days ago • 36
Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models Paper • 2609.34972 • Published 4 days ago • 28
Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence Paper • 2609.35432 • Published 4 days ago • 104
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 4 days ago • 46
FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance Paper • 2609.25716 • Published 10 days ago • 22