TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 14 days ago • 166
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 17 days ago • 170
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer Paper • 2606.16255 • Published Jun 15 • 15
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Paper • 2606.13289 • Published Jun 11 • 30
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Paper • 2606.13289 • Published Jun 11 • 30
Representation Forcing for Bottleneck-Free Unified Multimodal Models Paper • 2605.31604 • Published May 29 • 63
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens Paper • 2603.19232 • Published Mar 19 • 33
Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy Paper • 2511.21579 • Published Nov 26, 2025 • 23
Video Generation Models Are Good Latent Reward Models Paper • 2511.21541 • Published Nov 26, 2025 • 50
Video Generation Models Are Good Latent Reward Models Paper • 2511.21541 • Published Nov 26, 2025 • 50
Video Generation Models Are Good Latent Reward Models Paper • 2511.21541 • Published Nov 26, 2025 • 50 • 6
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation Paper • 2511.19365 • Published Nov 24, 2025 • 66
Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame Interpolation Paper • 2303.00440 • Published Mar 1, 2023 • 1
DPL: Decoupled Prompt Learning for Vision-Language Models Paper • 2308.10061 • Published Aug 19, 2023 • 2
MGMAE: Motion Guided Masking for Video Masked Autoencoding Paper • 2308.10794 • Published Aug 21, 2023
StableDrag: Stable Dragging for Point-based Image Editing Paper • 2403.04437 • Published Mar 7, 2024 • 27