On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 241
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 6 days ago • 37
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published 11 days ago • 138
Self-Improvements in Modern Agentic Systems: A Survey Paper • 2607.13104 • Published 15 days ago • 31
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published 26 days ago • 81
Vidu S1: A Real-Time Interactive Video Generation Model Paper • 2607.03118 • Published 26 days ago • 144
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published 21 days ago • 64
Representation Distribution Matching for One-Step Visual Generation Paper • 2607.02375 • Published 27 days ago • 11
MemLearner: Learning to Query Context memory for Video World Models Paper • 2606.31734 • Published 29 days ago • 28
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Paper • 2604.24763 • Published Apr 27 • 71
SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators Paper • 2502.06394 • Published Feb 10, 2025 • 88
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Paper • 2502.06703 • Published Feb 10, 2025 • 153