Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing Paper • 2610.00825 • Published 5 days ago • 17
What Does Privileged Information Add to On-Policy Self-Distillation? Paper • 2609.20612 • Published 18 days ago • 36
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published 28 days ago • 19
QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation Paper • 2608.29253 • Published Aug 29 • 14
Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration Paper • 2609.01072 • Published Sep 2 • 13
CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers Paper • 2608.12773 • Published Aug 13 • 9
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design Paper • 2608.10299 • Published Aug 10 • 138
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published Aug 3 • 96