Mamba: Linear-Time Sequence Modeling with Selective State Spaces Paper • 2312.00752 • Published Dec 1, 2023 • 153
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 19 days ago • 85
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 13 days ago • 170
HuggingFace's Transformers: State-of-the-art Natural Language Processing Paper • 1910.03771 • Published Oct 9, 2019 • 26
Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models Paper • 2606.03748 • Published Jun 2 • 21
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Paper • 2506.09985 • Published Jun 11, 2025 • 33
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning Paper • 2603.14482 • Published Mar 15 • 37
Running Featured 55 Gemma 4 - Vision Token Budget 🖼 55 Resize images for visual token budgets while keeping aspect ratio
V-JEPA 2.1 – HuggingFace ports Collection HF-format conversions of Meta's V-JEPA 2.1 encoders (384px), with numerical validation against the reference implementation. • 4 items • Updated 2 days ago • 2