PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation Paper • 2609.38597 • Published 3 days ago • 9
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Paper • 2609.38721 • Published 2 days ago • 271
view article Article NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction nvidia • 3 days ago • 61
MLLMs Need 3D-Aware Representation Supervision for Scene Understanding Paper • 2506.01946 • Published Jun 2, 2025 • 3
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 22 days ago • 706
Elastic Token Compression for Pixel-Space Diffusion Transformers Paper • 2608.29281 • Published Aug 29 • 1
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Paper • 2608.03573 • Published Aug 6 • 61
UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation Paper • 2606.23503 • Published Jun 22 • 3
Next-Latent Prediction Transformers Learn Compact World Models Paper • 2511.05963 • Published Nov 8, 2025 • 7
MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation Paper • 2606.09056 • Published Jun 8 • 8
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models Paper • 2605.21573 • Published May 20 • 109
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion Paper • 2605.23902 • Published May 22 • 46
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs Paper • 2603.16932 • Published Mar 14 • 91
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Paper • 2508.18265 • Published Aug 25, 2025 • 222
Video Analysis and Generation via a Semantic Progress Function Paper • 2604.22554 • Published Apr 24 • 59