OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper • 2609.21465 • Published 23 days ago • 153
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published about 1 month ago • 656
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published Jul 13 • 78
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs Paper • 2601.08763 • Published Jan 13 • 150