Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 4 days ago • 23
ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations Paper • 2606.11188 • Published Jun 9 • 27
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Paper • 2604.26694 • Published Apr 29 • 6
VQ-VA World: Towards High-Quality Visual Question-Visual Answering Paper • 2511.20573 • Published Nov 25, 2025 • 7
LightBagel: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation Paper • 2510.22946 • Published Oct 27, 2025 • 18
MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentation Paper • 2505.04656 • Published May 7, 2025 • 1
DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion Paper • 2411.04928 • Published Nov 7, 2024 • 56
GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting Paper • 2311.14521 • Published Nov 24, 2023
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 4 days ago • 23
MeshAnything V2: Artist-Created Mesh Generation With Adjacent Mesh Tokenization Paper • 2408.02555 • Published Aug 5, 2024 • 31
Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels Paper • 2405.16822 • Published May 27, 2024 • 13
V3D: Video Diffusion Models are Effective 3D Generators Paper • 2403.06738 • Published Mar 11, 2024 • 30