An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models Paper • 2608.16887 • Published Aug 17 • 36
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models Paper • 2607.06553 • Published Jul 9 • 15
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models Paper • 2607.06553 • Published Jul 9 • 15
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation Paper • 2510.06139 • Published Oct 7, 2025 • 4
From SRA to Self-Flow: Data Augmentation or Self-Supervision? Paper • 2607.02508 • Published Jul 2 • 13
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models Paper • 2605.05204 • Published May 6 • 28
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning Paper • 2601.21037 • Published Jan 28 • 15