H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction Paper • 2508.03118 • Published Aug 5, 2025 • 1
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model Paper • 2609.18323 • Published 20 days ago • 133
H3-World: Turning Language Understanding into World Control Paper • 2609.01560 • Published Sep 1 • 53
GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation Paper • 2609.24981 • Published 15 days ago • 73
Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion Paper • 2603.15614 • Published Mar 16 • 7
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Paper • 2608.09888 • Published Aug 10 • 798
SolarWM Collection Family-specific checkpoints, released data, and paper for SolarWM long-horizon interactive video world models. • 7 items • Updated 25 days ago • 10
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution Paper • 2608.16157 • Published Aug 17 • 113
HDR Video Generation via Latent Alignment with Logarithmic Encoding Paper • 2604.11788 • Published Apr 13 • 15
view article Article FLUX 3 Model Overview: Multimodal Flow Models for Image, Video, Audio, and Action Prediction ResterChed • Jul 24 • 25
CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation Paper • 2607.03803 • Published Jul 4 • 17
StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration Paper • 2605.25659 • Published May 25 • 14
KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks Paper • 2606.03458 • Published Jun 2 • 67
JLT: Clean-Latent Prediction in Latent Diffusion Transformers Paper • 2605.27102 • Published May 26 • 30