OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper • 2607.23855 • Published 2 days ago • 19
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 15 days ago • 84
ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-level Text-Image Alignment in Diffusion Models Paper • 2506.09740 • Published Jun 11, 2025 • 1
Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter Paper • 2309.02773 • Published Sep 6, 2023 • 1
Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter Paper • 2309.02773 • Published Sep 6, 2023 • 1