Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL Paper • 2609.37200 • Published 10 days ago • 138
RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling Paper • 2609.22947 • Published 20 days ago • 43
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout Paper • 2609.09123 • Published about 1 month ago • 55
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning Paper • 2606.27828 • Published Jun 26 • 26
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models Paper • 2605.25077 • Published May 24 • 22
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation Paper • 2605.23271 • Published May 22 • 31
CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition Paper • 2605.19995 • Published May 19 • 35
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors Paper • 2605.00658 • Published May 1 • 87
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video Paper • 2604.11102 • Published Apr 13 • 8
Textured 3D Regenerative Morphing with 3D Diffusion Prior Paper • 2502.14316 • Published Feb 20, 2025 • 1
Counterfactual Explanations for Face Forgery Detection via Adversarial Removal of Artifacts Paper • 2404.08341 • Published Apr 12, 2024 • 1
SeFi-IDE: Semantic-Fidelity Identity Embedding for Personalized Diffusion-Based Generation Paper • 2402.00631 • Published Jan 31, 2024 • 1
Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models Paper • 2604.10949 • Published Apr 13 • 15
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model Paper • 2603.22281 • Published Mar 23 • 22
HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning Paper • 2603.17024 • Published Mar 17 • 111
Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible Multilinguality Paper • 2603.17512 • Published Mar 18 • 4
AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science Paper • 2603.19005 • Published Mar 19 • 7
FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow Paper • 2603.19598 • Published Mar 20 • 32