facebook/dinov3-vitb16-pretrain-lvd1689m Image Feature Extraction • 85.7M • Updated Aug 19, 2025 • 778k • 201
InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning Paper • 2505.13888 • Published May 20, 2025
Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models Paper • 2410.00363 • Published Oct 1, 2024 • 1
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT Paper • 2406.18583 • Published Jun 5, 2024
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers Paper • 2405.05945 • Published May 9, 2024 • 4
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation Paper • 2508.06426 • Published Aug 8, 2025 • 10
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model Paper • 2403.08350 • Published Mar 13, 2024
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation Paper • 2508.06426 • Published Aug 8, 2025 • 10
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation Paper • 2508.06426 • Published Aug 8, 2025 • 10 • 2
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation Paper • 2508.05635 • Published Aug 7, 2025 • 73
Adversarial Data Collection: Human-Collaborative Perturbations for Efficient and Robust Robotic Imitation Learning Paper • 2503.11646 • Published Mar 14, 2025 • 34
Running on Zero Agents Featured 219 Microsoft Phi-3-Vision-128k 😻 219 Chat with an image using Phi-3 Vision model
Running on CPU Upgrade Featured 969 TTS Arena V2 🗣 969 Compare and rank TTS voices by listening and voting
Channel Importance Matters in Few-Shot Image Classification Paper • 2206.08126 • Published Jun 16, 2022