ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing Paper • 2609.38541 • Published 6 days ago • 30
SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video Paper • 2609.37969 • Published 6 days ago • 39
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? Paper • 2609.38079 • Published 6 days ago • 52
PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing Paper • 2609.23784 • Published 15 days ago • 16
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Paper • 2609.20816 • Published 18 days ago • 57
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 18 days ago • 138
OpenVE inference results Collection VideoCoF OPD 的 OpenVE-Bench 推理结果,供 openve-watcher 爬取并用 Gemini 打分 • 114 items • Updated about 1 hour ago • 1
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models Paper • 2609.02886 • Published Sep 2 • 118
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models Paper • 2608.16887 • Published Aug 17 • 36
From SRA to Self-Flow: Data Augmentation or Self-Supervision? Paper • 2607.02508 • Published Jul 2 • 13
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing Paper • 2606.26740 • Published Jun 25 • 82
Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning Paper • 2606.07436 • Published Jun 5 • 27
CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives Paper • 2605.12496 • Published May 12 • 31
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer Paper • 2605.15178 • Published May 14 • 92
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation Paper • 2605.13724 • Published May 13 • 104
Lightning Unified Video Editing via In-Context Sparse Attention Paper • 2605.04569 • Published May 6 • 17
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models Paper • 2605.05204 • Published May 6 • 28
LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories Paper • 2604.15311 • Published Apr 16 • 14
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Paper • 2604.11804 • Published Apr 13 • 73
Mode Seeking meets Mean Seeking for Fast Long Video Generation Paper • 2602.24289 • Published Feb 27 • 41