MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation Paper • 2606.26016 • Published Jun 24 • 10
view post Post 227 Hi everyone,I've created a Gradio space for embedding and extracting invisible watermarks in images:👉 eienmojiki/blind-watermark-studioIt supports hiding text, images, and bit arrays using the DWT-DCT-SVD algorithm.Credits:- Original library: https://github.com/guofei9987/blind_watermark- Author: Guo Fei:). See translation Reply
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer Paper • 2606.16255 • Published Jun 15 • 15
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer Paper • 2606.16255 • Published Jun 15 • 15
Representation Forcing for Bottleneck-Free Unified Multimodal Models Paper • 2605.31604 • Published May 29 • 63
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens Paper • 2603.19232 • Published Mar 19 • 33
ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks Paper • 2603.27862 • Published Mar 29 • 33
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens Paper • 2603.19232 • Published Mar 19 • 33
SeeingEye: Agentic Information Flow Unlocks Multimodal Reasoning In Text-only LLMs Paper • 2510.25092 • Published Oct 29, 2025 • 8
Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment Paper • 2511.22345 • Published Nov 27, 2025 • 13
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation Paper • 2511.19365 • Published Nov 24, 2025 • 66