Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing Paper • 2610.00825 • Published 4 days ago • 11
microsoft/VibeVoice-ASR-Streaming-7B Automatic Speech Recognition • 9B • Updated about 1 month ago • 6.64k • 251
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence Paper • 2608.10720 • Published Aug 11 • 16
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing Paper • 2608.04956 • Published Aug 5 • 17
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing Paper • 2608.06146 • Published Aug 6 • 24
Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval Paper • 2608.06060 • Published Aug 6 • 41