HDR Video Generation via Latent Alignment with Logarithmic Encoding Paper • 2604.11788 • Published Apr 13 • 15
view article Article FLUX 3 Model Overview: Multimodal Flow Models for Image, Video, Audio, and Action Prediction ResterChed • 16 days ago • 13
CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation Paper • 2607.03803 • Published Jul 4 • 19
StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration Paper • 2605.25659 • Published May 25 • 17
KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks Paper • 2606.03458 • Published Jun 2 • 69
JLT: Clean-Latent Prediction in Latent Diffusion Transformers Paper • 2605.27102 • Published May 26 • 33
Nemotron-Labs-Diffusion Collection A Tri-Mode Language Model Family Unifying Autoregressive, Diffusion, and Self-Speculation Decoding • 7 items • Updated 24 days ago • 53
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer Paper • 2509.24695 • Published Sep 29, 2025 • 54
AudioX: Diffusion Transformer for Anything-to-Audio Generation Paper • 2503.10522 • Published Mar 13, 2025 • 29
Geometric Context Transformer for Streaming 3D Reconstruction Paper • 2604.14141 • Published Apr 15 • 38
Kronos: A Foundation Model for the Language of Financial Markets Paper • 2508.02739 • Published Aug 2, 2025 • 51
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression Paper • 2604.04921 • Published Apr 6 • 117
The Y-Combinator for LLMs: Solving Long-Context Rot with λ-Calculus Paper • 2603.20105 • Published Mar 20 • 37