merve PRO
AI & ML interests
Recent Activity
Organizations
-
PaddlePaddle/PP-OCRv6_medium_det
Image-to-Text • Updated • 80k • 31 -
PaddlePaddle/PP-OCRv6_tiny_det_safetensors
Image-to-Text • 438k • Updated • 137 • 27 -
moonshotai/Kimi-K2.7-Code
Image-Text-to-Text • 1T • Updated • 99.8k • • 1.4k -
PaddlePaddle/PP-OCRv6_small_rec_onnx
Image-to-Text • Updated • 14.5k • 22
-
internlm/Intern-S2-Preview
Image-Text-to-Text • 36B • Updated • 845 • 120 -
nvidia/nemotron-3.5-asr-streaming-0.6b
Automatic Speech Recognition • 0.6B • Updated • 828k • • 1.13k -
internlm/Intern-S2-Preview-FP8
Image-Text-to-Text • 36B • Updated • 165 • 24 -
Aratako/Irodori-TTS-500M-v3
Text-to-Speech • 0.5B • Updated • 116
-
google/translategemma-27b-it
Image-Text-to-Text • 29B • Updated • 13.8k • 400 -
kakaocorp/kanana-2-30b-a3b-mid-2601
Text Generation • 31B • Updated • 159 • 31 -
black-forest-labs/FLUX.2-klein-base-4B
Image-to-Image • 4B • Updated • 280k • • 172 -
google/translategemma-12b-it
Image-Text-to-Text • 13B • Updated • 9.05k • 345
-
facebook/metaclip-2-worldwide-s16
Zero-Shot Image Classification • 0.4B • Updated • 1.14k • 10 -
facebook/metaclip-2-worldwide-m16
Zero-Shot Image Classification • 0.5B • Updated • 35 • 4 -
facebook/metaclip-2-worldwide-l14
Zero-Shot Image Classification • 1B • Updated • 2.46k • 13 -
facebook/metaclip-2-worldwide-b32
Zero-Shot Image Classification • 0.6B • Updated • 724 • 7
-
deepseek-ai/DeepSeek-V3-0324
Text Generation • 685B • Updated • 1.31M • • 3.18k -
Qwen/Qwen2.5-Omni-7B
Any-to-Any • 11B • Updated • 325k • 1.95k -
google/txgemma-27b-chat
Text Generation • 27B • Updated • 96 • • 62 - RunningAgentsFeatured375
Qwen2.5 Omni 7B Demo
🏆375Chat with text, audio, images, and video, get spoken replies
- Running on ZeroAgents273
Qwen2-VL-7B
🔥273Answer questions about uploaded images
- RunningAgents67
UI-TARS
🌖67Predict UI click coordinates from a screenshot and instruction
- PausedAgents102
Qwen2.5-1M Demo
💻102Ask questions about your uploaded documents
-
Qwen/Qwen2.5-14B-Instruct-1M
Text Generation • 15B • Updated • 29.4k • • 340
-
ibm-granite/granite-3.0-8b-instruct
Text Generation • 8B • Updated • 20.9k • 208 -
ibm-granite/granite-3.0-2b-instruct
Text Generation • 3B • Updated • 5.2k • 48 -
CohereLabs/aya-expanse-8b
Text Generation • 8B • Updated • 14k • 451 -
CohereLabs/aya-expanse-32b
Text Generation • 32B • Updated • 4.81k • • 303
- Running on ZeroAgentsFeatured209
DepthCrafter
🦀209a super consistent video depth model
- PausedAgentsFeatured223
Depth Pro
🚀223Generate an inverse depth map from an image
- Running on ZeroAgents80
Lotus Depth
🚀80Official Demo of Lotus (https://lotus3d.github.io/)
-
apple/DepthPro
Depth Estimation • Updated • 4.96k • 554
-
microsoft/resnet-50
Image Classification • 25.6M • Updated • 590k • • 514 -
google/vit-base-patch16-224-in21k
Image Feature Extraction • 86.4M • Updated • 1.89M • 419 -
google/vit-base-patch32-224-in21k
Image Feature Extraction • 88M • Updated • 20.5k • 20 -
facebook/dinov2-large
Image Feature Extraction • 0.3B • Updated • 857k • 119
-
facebook/detr-resnet-50
Object Detection • 41.6M • Updated • 349k • • 977 -
facebook/detr-resnet-101-dc5
Object Detection • 60.7M • Updated • 3.21k • 21 -
facebook/detr-resnet-50-dc5
Object Detection • 41.6M • Updated • 2.06k • 7 -
google/owlvit-base-patch32
Zero-Shot Object Detection • 0.2B • Updated • 205k • 151
-
openai/clip-vit-large-patch14
Zero-Shot Image Classification • 0.4B • Updated • 8.65M • 2.1k -
openai/clip-vit-base-patch32
Zero-Shot Image Classification • Updated • 22.3M • 1.57k -
laion/CLIP-ViT-bigG-14-laion2B-39B-b160k
Zero-Shot Image Classification • Updated • 88.5k • 321 -
kakaobrain/align-base
Zero-Shot Image Classification • Updated • 28.8k • 31
-
microsoft/xclip-base-patch32
Video Classification • 0.2B • Updated • 114k • 119 -
facebook/timesformer-base-finetuned-k400
Video Classification • Updated • 13.7k • 44 -
facebook/timesformer-base-finetuned-k600
Video Classification • Updated • 15.3k • 12 -
google/vivit-b-16x2
Video Classification • Updated • 23k • 11
- Running on ZeroAgentsFeatured74
Draw To Search Art
🐠74Draw/upload image and search among WikiART using SigLIP
- Running on CPU UpgradeAgents23
Compare Clip Siglip
🏃23Compare strong zero-shot image classification models
- Runtime errorAgents14
Multilingual Zero Shot Image Clf
🏢14Comparing powerful multilingual zero-shot image clf models
-
BAAI/bunny-phi-2-siglip-lora
Text Generation • Updated • 175 • 48
-
google/owlvit-base-patch32
Zero-Shot Object Detection • 0.2B • Updated • 205k • 151 -
google/owlvit-base-patch16
Zero-Shot Object Detection • Updated • 31.9k • 14 -
google/owlvit-large-patch14
Zero-Shot Object Detection • Updated • 7.63k • 29 -
google/owlv2-base-patch16
Zero-Shot Object Detection • 0.2B • Updated • 87.3k • 30
-
google/owlvit-base-patch32
Zero-Shot Object Detection • 0.2B • Updated • 205k • 151 -
google/owlvit-base-patch16
Zero-Shot Object Detection • Updated • 31.9k • 14 -
google/owlvit-large-patch14
Zero-Shot Object Detection • Updated • 7.63k • 29 -
google/owlv2-base-patch16
Zero-Shot Object Detection • 0.2B • Updated • 87.3k • 30
- PausedAgents21
Video Llava
🐨21Generate descriptions by uploading images or videos
-
llava-hf/LLaVA-NeXT-Video-7B-hf
Video-Text-to-Text • 7B • Updated • 111k • 126 -
llava-hf/LLaVA-NeXT-Video-7B-DPO-hf
Video-Text-to-Text • 7B • Updated • 550 • 12 -
llava-hf/LLaVA-NeXT-Video-7B-32K-hf
Image-Text-to-Text • 8B • Updated • 703 • 9
-
NVEagle/Eagle-X5-13B
Image-Text-to-Text • 15B • Updated • 62 • 15 -
NVEagle/Eagle-X5-13B-Chat
Image-Text-to-Text • 15B • Updated • 161 • 28 -
NVEagle/Eagle-X5-7B
Image-Text-to-Text • 9B • Updated • 84 • 26 - Configuration errorAgents64
Eagle X5 13B Chat
🚀64Combine text and images to generate responses
-
thinkingmachines/Inkling
Image-Text-to-Text • 952B • Updated • 374k • • 1.79k -
Lightricks/LTX-2.3-22b-LoRA-Foley-V2A
Text-to-Audio • Updated • 373 • 32 -
genzeonplatform/healthcare-brain-diagnosis-icd-ner
Token Classification • 0.1B • Updated • 41 • 21 -
Hippotes/Ideogram4-Fal-ComfyUI
Updated • 1.84k • 33
- Running on ZeroAgentsFeatured59
RF-DETR Realtime Webcam Demo
🎯59Segment objects in live webcam and uploaded media
-
Roboflow/rf-detr-base
Object Detection • 32.2M • Updated • 38.1k • 5 -
Roboflow/rf-detr-base-2
Object Detection • 32.2M • Updated • 37 -
Roboflow/rf-detr-nano
Object Detection • 30.5M • Updated • 5.2k
-
OpenMOSS-Team/MOSS-Audio-4B-Instruct
Audio-Text-to-Text • 5B • Updated • 15.3k • 84 -
OpenMOSS-Team/MOSS-Audio-8B-Thinking
Audio-Text-to-Text • 9B • Updated • 732 • 84 -
bytedance-research/Timer-S1
Time Series Forecasting • 8B • Updated • 507 • 35 -
BugTraceAI/BugTraceAI-Apex-G4-26B-Q4
25B • Updated • 759 • 92
- Runtime errorAgents26
YOLO26
💙26Process images with advanced object detection and segmentation
- RunningFeatured69
YOLO26 WebGPU
🏆69Real-time object detection & pose estimation in your browser
-
onnx-community/yolo26x-ONNX
Updated • 68 • 5 -
openvision/yoloe26-n-seg
Zero-Shot Object Detection • Updated • 209 • 2
-
Wuli-art/Qwen-Image-2512-Turbo-LoRA
Text-to-Image • Updated • 5.23k • • 227 -
miromind-ai/MiroThinker-v1.5-235B
Text Generation • 235B • Updated • 89 • 254 -
prithivMLmods/Qwen-Image-Edit-2511-Object-Remover
Image-to-Image • Updated • 7.67k • • 77 -
tencent/Youtu-LLM-2B-Base
Text Generation • 2B • Updated • 2.24k • 43
-
facebook/sam3
Mask Generation • 0.9B • Updated • 2.1M • 3.49k - Running on ZeroAgentsFeatured116
SAM3 Video Segmentation
🐠116Track and label objects in videos using text prompts or clicks
-
onnx-community/sam3-tracker-ONNX
Mask Generation • Updated • 1.17k • 38 - Running32
SAM3 Tracker WebGPU
🎯32Segment images with click points and download cutouts
-
opendatalab/OmniDocBench
Viewer • Updated • 1.66k • 26k • 109 -
nanonets/Nanonets-OCR-s
Image-Text-to-Text • 4B • Updated • 187k • 1.6k -
echo840/MonkeyOCR
Image-Text-to-Text • Updated • 454 • 516 - Running on ZeroMCPFeatured143
Multimodal OCR2
💻143FireRed / Nanonets / Monkey / Thyme / Typhoon / SmolDocling
-
moonshotai/Kimi-VL-A3B-Thinking
Image-Text-to-Text • 16B • Updated • 16.6k • 454 -
agentica-org/DeepCoder-14B-Preview
Text Generation • 15B • Updated • 665 • • 683 -
HiDream-ai/HiDream-I1-Full
Text-to-Image • 17B • Updated • 858 • • 1k -
OpenGVLab/InternVL3-78B
Image-Text-to-Text • 78B • Updated • 13.7k • 239
-
NVLM: Open Frontier-Class Multimodal LLMs
Paper • 2409.11402 • Published • 75 -
BRAVE: Broadening the visual encoding of vision-language models
Paper • 2404.07204 • Published • 20 -
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Paper • 2403.18814 • Published • 49 -
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models
Paper • 2409.17146 • Published • 123
- Running on ZeroAgentsFeatured101
Lotus Normal
🌍101Official Demo of Lotus (https://lotus3d.github.io/)
- Running on ZeroAgents80
Lotus Depth
🚀80Official Demo of Lotus (https://lotus3d.github.io/)
-
jingheya/lotus-depth-g-v1-0
Depth Estimation • 0.9B • Updated • 1.66k • 27 -
jingheya/lotus-depth-d-v1-0
Depth Estimation • 0.9B • Updated • 403 • 5
-
facebook/dinov2-large
Image Feature Extraction • 0.3B • Updated • 857k • 119 -
google/flan-t5-xl
3B • Updated • 193k • 536 -
google/siglip-large-patch16-384
Zero-Shot Image Classification • 0.7B • Updated • 34.1k • 13 -
google/vit-huge-patch14-224-in21k
Image Feature Extraction • 0.6B • Updated • 23k • 22
-
facebook/deit-base-distilled-patch16-384
Image Classification • 87.6M • Updated • 11.4k • • 8 -
facebook/convnextv2-base-1k-224
Image Classification • 88.7M • Updated • 2.86k • • 4 -
facebook/deit-base-distilled-patch16-224
Image Classification • Updated • 19.5k • • 34 -
google/vit-base-patch32-384
Image Classification • 88.3M • Updated • 16.7k • • 23
-
facebook/maskformer-swin-large-coco
Image Segmentation • 0.2B • Updated • 313 • 27 -
nvidia/segformer-b0-finetuned-ade-512-512
Image Segmentation • 3.75M • Updated • 368k • • 201 -
facebook/detr-resnet-50-dc5-panoptic
Image Segmentation • 43M • Updated • 165 • 3 -
nvidia/segformer-b5-finetuned-cityscapes-1024-1024
Image Segmentation • Updated • 27.2k • • 45
-
timbrooks/instruct-pix2pix
Image-to-Image • 0.9B • Updated • 32.5k • 1.18k -
TencentARC/t2i-adapter-canny-sdxl-1.0
Image-to-Image • 79M • Updated • 2.7k • 54 -
TencentARC/t2i-adapter-sketch-sdxl-1.0
Image-to-Image • 79M • Updated • 2.99k • 77 -
CrucibleAI/ControlNetMediaPipeFace
Image-to-Image • 0.4B • Updated • 953 • 575
-
Salesforce/blip-image-captioning-large
Image-to-Text • 0.5B • Updated • 568k • 1.49k -
Salesforce/blip-image-captioning-base
Image-to-Text • Updated • 1.66M • 892 -
microsoft/trocr-base-handwritten
Image-to-Text • 0.3B • Updated • 185k • 523 -
microsoft/git-large-coco
Image-to-Text • 0.4B • Updated • 2.96k • 106
- RunningAgents126
Grounding DINO Demo
💻126Cutting edge open-vocabulary object detection app
- RunningAgentsFeatured106
Owlv2
👀106State-of-the-art Zero-shot Object Detection
- Configuration errorAgentsFeatured41
BLIP2 with transformers
🌖41BLIP2 (cutting edge image captioning) in 🤗transformers
- Build errorAgentsFeatured377
IDEFICS Playground
🐨377
- RunningAgentsFeatured106
Owlv2
👀106State-of-the-art Zero-shot Object Detection
- Running on ZeroAgentsFeatured64
Owl Tracking
⚡64Powerful foundation model for zero-shot object tracking
- Sleeping26
Search and Detect (CLIP/OWL-ViT)
🦉26Search and detect objects in images using text queries
- Running on ZeroAgentsFeatured110
OWLSAM
😻110State-of-the-art open-vocabulary image segmentation ⚡️
- Running on ZeroAgentsFeatured83
UDOP
🏃83Generate answers or summaries from document images with prompts
- Configuration errorAgents40
Pix2struct
📚40Play with all the pix2struct variants in this d
- RunningAgents26
Compare Docvqa Models
🦀26Compare different visual question answering
- Runtime errorAgentsFeatured289
DocQuery — Document Query Engine
🦉289
-
Improved Baselines with Visual Instruction Tuning
Paper • 2310.03744 • Published • 39 -
DeepSeek-VL: Towards Real-World Vision-Language Understanding
Paper • 2403.05525 • Published • 49 -
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
Paper • 2308.12966 • Published • 12 -
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
Paper • 2404.01331 • Published • 27
-
google/owlvit-base-patch32
Zero-Shot Object Detection • 0.2B • Updated • 205k • 151 -
google/owlvit-base-patch16
Zero-Shot Object Detection • Updated • 31.9k • 14 -
google/owlvit-large-patch14
Zero-Shot Object Detection • Updated • 7.63k • 29 -
google/owlv2-base-patch16
Zero-Shot Object Detection • 0.2B • Updated • 87.3k • 30
- RunningAgents212
Vidore Leaderboard
🥇212Browse and compare visual document retrieval model scores
- Running on CPU UpgradeAgents1.03k
Open VLM Leaderboard
🌎1.03kVLMEvalKit Evaluation Results Collection
- RunningFeatured561
Vision Arena (Testing VLMs side-by-side)
🖼561Explore Vision Arena visual AI demo online
- Build errorAgentsFeatured85
SEED-Bench Leaderboard
🏆85Submit model evaluation results to leaderboard
-
thinkingmachines/Inkling
Image-Text-to-Text • 952B • Updated • 374k • • 1.79k -
Lightricks/LTX-2.3-22b-LoRA-Foley-V2A
Text-to-Audio • Updated • 373 • 32 -
genzeonplatform/healthcare-brain-diagnosis-icd-ner
Token Classification • 0.1B • Updated • 41 • 21 -
Hippotes/Ideogram4-Fal-ComfyUI
Updated • 1.84k • 33
-
PaddlePaddle/PP-OCRv6_medium_det
Image-to-Text • Updated • 80k • 31 -
PaddlePaddle/PP-OCRv6_tiny_det_safetensors
Image-to-Text • 438k • Updated • 137 • 27 -
moonshotai/Kimi-K2.7-Code
Image-Text-to-Text • 1T • Updated • 99.8k • • 1.4k -
PaddlePaddle/PP-OCRv6_small_rec_onnx
Image-to-Text • Updated • 14.5k • 22
-
internlm/Intern-S2-Preview
Image-Text-to-Text • 36B • Updated • 845 • 120 -
nvidia/nemotron-3.5-asr-streaming-0.6b
Automatic Speech Recognition • 0.6B • Updated • 828k • • 1.13k -
internlm/Intern-S2-Preview-FP8
Image-Text-to-Text • 36B • Updated • 165 • 24 -
Aratako/Irodori-TTS-500M-v3
Text-to-Speech • 0.5B • Updated • 116
- Running on ZeroAgentsFeatured59
RF-DETR Realtime Webcam Demo
🎯59Segment objects in live webcam and uploaded media
-
Roboflow/rf-detr-base
Object Detection • 32.2M • Updated • 38.1k • 5 -
Roboflow/rf-detr-base-2
Object Detection • 32.2M • Updated • 37 -
Roboflow/rf-detr-nano
Object Detection • 30.5M • Updated • 5.2k
-
OpenMOSS-Team/MOSS-Audio-4B-Instruct
Audio-Text-to-Text • 5B • Updated • 15.3k • 84 -
OpenMOSS-Team/MOSS-Audio-8B-Thinking
Audio-Text-to-Text • 9B • Updated • 732 • 84 -
bytedance-research/Timer-S1
Time Series Forecasting • 8B • Updated • 507 • 35 -
BugTraceAI/BugTraceAI-Apex-G4-26B-Q4
25B • Updated • 759 • 92
-
google/translategemma-27b-it
Image-Text-to-Text • 29B • Updated • 13.8k • 400 -
kakaocorp/kanana-2-30b-a3b-mid-2601
Text Generation • 31B • Updated • 159 • 31 -
black-forest-labs/FLUX.2-klein-base-4B
Image-to-Image • 4B • Updated • 280k • • 172 -
google/translategemma-12b-it
Image-Text-to-Text • 13B • Updated • 9.05k • 345
- Runtime errorAgents26
YOLO26
💙26Process images with advanced object detection and segmentation
- RunningFeatured69
YOLO26 WebGPU
🏆69Real-time object detection & pose estimation in your browser
-
onnx-community/yolo26x-ONNX
Updated • 68 • 5 -
openvision/yoloe26-n-seg
Zero-Shot Object Detection • Updated • 209 • 2
-
Wuli-art/Qwen-Image-2512-Turbo-LoRA
Text-to-Image • Updated • 5.23k • • 227 -
miromind-ai/MiroThinker-v1.5-235B
Text Generation • 235B • Updated • 89 • 254 -
prithivMLmods/Qwen-Image-Edit-2511-Object-Remover
Image-to-Image • Updated • 7.67k • • 77 -
tencent/Youtu-LLM-2B-Base
Text Generation • 2B • Updated • 2.24k • 43
-
facebook/sam3
Mask Generation • 0.9B • Updated • 2.1M • 3.49k - Running on ZeroAgentsFeatured116
SAM3 Video Segmentation
🐠116Track and label objects in videos using text prompts or clicks
-
onnx-community/sam3-tracker-ONNX
Mask Generation • Updated • 1.17k • 38 - Running32
SAM3 Tracker WebGPU
🎯32Segment images with click points and download cutouts
-
facebook/metaclip-2-worldwide-s16
Zero-Shot Image Classification • 0.4B • Updated • 1.14k • 10 -
facebook/metaclip-2-worldwide-m16
Zero-Shot Image Classification • 0.5B • Updated • 35 • 4 -
facebook/metaclip-2-worldwide-l14
Zero-Shot Image Classification • 1B • Updated • 2.46k • 13 -
facebook/metaclip-2-worldwide-b32
Zero-Shot Image Classification • 0.6B • Updated • 724 • 7
-
opendatalab/OmniDocBench
Viewer • Updated • 1.66k • 26k • 109 -
nanonets/Nanonets-OCR-s
Image-Text-to-Text • 4B • Updated • 187k • 1.6k -
echo840/MonkeyOCR
Image-Text-to-Text • Updated • 454 • 516 - Running on ZeroMCPFeatured143
Multimodal OCR2
💻143FireRed / Nanonets / Monkey / Thyme / Typhoon / SmolDocling
-
moonshotai/Kimi-VL-A3B-Thinking
Image-Text-to-Text • 16B • Updated • 16.6k • 454 -
agentica-org/DeepCoder-14B-Preview
Text Generation • 15B • Updated • 665 • • 683 -
HiDream-ai/HiDream-I1-Full
Text-to-Image • 17B • Updated • 858 • • 1k -
OpenGVLab/InternVL3-78B
Image-Text-to-Text • 78B • Updated • 13.7k • 239
-
deepseek-ai/DeepSeek-V3-0324
Text Generation • 685B • Updated • 1.31M • • 3.18k -
Qwen/Qwen2.5-Omni-7B
Any-to-Any • 11B • Updated • 325k • 1.95k -
google/txgemma-27b-chat
Text Generation • 27B • Updated • 96 • • 62 - RunningAgentsFeatured375
Qwen2.5 Omni 7B Demo
🏆375Chat with text, audio, images, and video, get spoken replies
- Running on ZeroAgents273
Qwen2-VL-7B
🔥273Answer questions about uploaded images
- RunningAgents67
UI-TARS
🌖67Predict UI click coordinates from a screenshot and instruction
- PausedAgents102
Qwen2.5-1M Demo
💻102Ask questions about your uploaded documents
-
Qwen/Qwen2.5-14B-Instruct-1M
Text Generation • 15B • Updated • 29.4k • • 340
-
NVLM: Open Frontier-Class Multimodal LLMs
Paper • 2409.11402 • Published • 75 -
BRAVE: Broadening the visual encoding of vision-language models
Paper • 2404.07204 • Published • 20 -
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Paper • 2403.18814 • Published • 49 -
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models
Paper • 2409.17146 • Published • 123
-
ibm-granite/granite-3.0-8b-instruct
Text Generation • 8B • Updated • 20.9k • 208 -
ibm-granite/granite-3.0-2b-instruct
Text Generation • 3B • Updated • 5.2k • 48 -
CohereLabs/aya-expanse-8b
Text Generation • 8B • Updated • 14k • 451 -
CohereLabs/aya-expanse-32b
Text Generation • 32B • Updated • 4.81k • • 303
- Running on ZeroAgentsFeatured101
Lotus Normal
🌍101Official Demo of Lotus (https://lotus3d.github.io/)
- Running on ZeroAgents80
Lotus Depth
🚀80Official Demo of Lotus (https://lotus3d.github.io/)
-
jingheya/lotus-depth-g-v1-0
Depth Estimation • 0.9B • Updated • 1.66k • 27 -
jingheya/lotus-depth-d-v1-0
Depth Estimation • 0.9B • Updated • 403 • 5
- Running on ZeroAgentsFeatured209
DepthCrafter
🦀209a super consistent video depth model
- PausedAgentsFeatured223
Depth Pro
🚀223Generate an inverse depth map from an image
- Running on ZeroAgents80
Lotus Depth
🚀80Official Demo of Lotus (https://lotus3d.github.io/)
-
apple/DepthPro
Depth Estimation • Updated • 4.96k • 554
-
facebook/dinov2-large
Image Feature Extraction • 0.3B • Updated • 857k • 119 -
google/flan-t5-xl
3B • Updated • 193k • 536 -
google/siglip-large-patch16-384
Zero-Shot Image Classification • 0.7B • Updated • 34.1k • 13 -
google/vit-huge-patch14-224-in21k
Image Feature Extraction • 0.6B • Updated • 23k • 22
-
microsoft/resnet-50
Image Classification • 25.6M • Updated • 590k • • 514 -
google/vit-base-patch16-224-in21k
Image Feature Extraction • 86.4M • Updated • 1.89M • 419 -
google/vit-base-patch32-224-in21k
Image Feature Extraction • 88M • Updated • 20.5k • 20 -
facebook/dinov2-large
Image Feature Extraction • 0.3B • Updated • 857k • 119
-
facebook/deit-base-distilled-patch16-384
Image Classification • 87.6M • Updated • 11.4k • • 8 -
facebook/convnextv2-base-1k-224
Image Classification • 88.7M • Updated • 2.86k • • 4 -
facebook/deit-base-distilled-patch16-224
Image Classification • Updated • 19.5k • • 34 -
google/vit-base-patch32-384
Image Classification • 88.3M • Updated • 16.7k • • 23
-
facebook/detr-resnet-50
Object Detection • 41.6M • Updated • 349k • • 977 -
facebook/detr-resnet-101-dc5
Object Detection • 60.7M • Updated • 3.21k • 21 -
facebook/detr-resnet-50-dc5
Object Detection • 41.6M • Updated • 2.06k • 7 -
google/owlvit-base-patch32
Zero-Shot Object Detection • 0.2B • Updated • 205k • 151
-
facebook/maskformer-swin-large-coco
Image Segmentation • 0.2B • Updated • 313 • 27 -
nvidia/segformer-b0-finetuned-ade-512-512
Image Segmentation • 3.75M • Updated • 368k • • 201 -
facebook/detr-resnet-50-dc5-panoptic
Image Segmentation • 43M • Updated • 165 • 3 -
nvidia/segformer-b5-finetuned-cityscapes-1024-1024
Image Segmentation • Updated • 27.2k • • 45
-
openai/clip-vit-large-patch14
Zero-Shot Image Classification • 0.4B • Updated • 8.65M • 2.1k -
openai/clip-vit-base-patch32
Zero-Shot Image Classification • Updated • 22.3M • 1.57k -
laion/CLIP-ViT-bigG-14-laion2B-39B-b160k
Zero-Shot Image Classification • Updated • 88.5k • 321 -
kakaobrain/align-base
Zero-Shot Image Classification • Updated • 28.8k • 31
-
timbrooks/instruct-pix2pix
Image-to-Image • 0.9B • Updated • 32.5k • 1.18k -
TencentARC/t2i-adapter-canny-sdxl-1.0
Image-to-Image • 79M • Updated • 2.7k • 54 -
TencentARC/t2i-adapter-sketch-sdxl-1.0
Image-to-Image • 79M • Updated • 2.99k • 77 -
CrucibleAI/ControlNetMediaPipeFace
Image-to-Image • 0.4B • Updated • 953 • 575
-
microsoft/xclip-base-patch32
Video Classification • 0.2B • Updated • 114k • 119 -
facebook/timesformer-base-finetuned-k400
Video Classification • Updated • 13.7k • 44 -
facebook/timesformer-base-finetuned-k600
Video Classification • Updated • 15.3k • 12 -
google/vivit-b-16x2
Video Classification • Updated • 23k • 11
-
Salesforce/blip-image-captioning-large
Image-to-Text • 0.5B • Updated • 568k • 1.49k -
Salesforce/blip-image-captioning-base
Image-to-Text • Updated • 1.66M • 892 -
microsoft/trocr-base-handwritten
Image-to-Text • 0.3B • Updated • 185k • 523 -
microsoft/git-large-coco
Image-to-Text • 0.4B • Updated • 2.96k • 106
- RunningAgents126
Grounding DINO Demo
💻126Cutting edge open-vocabulary object detection app
- RunningAgentsFeatured106
Owlv2
👀106State-of-the-art Zero-shot Object Detection
- Configuration errorAgentsFeatured41
BLIP2 with transformers
🌖41BLIP2 (cutting edge image captioning) in 🤗transformers
- Build errorAgentsFeatured377
IDEFICS Playground
🐨377
- RunningAgentsFeatured106
Owlv2
👀106State-of-the-art Zero-shot Object Detection
- Running on ZeroAgentsFeatured64
Owl Tracking
⚡64Powerful foundation model for zero-shot object tracking
- Sleeping26
Search and Detect (CLIP/OWL-ViT)
🦉26Search and detect objects in images using text queries
- Running on ZeroAgentsFeatured110
OWLSAM
😻110State-of-the-art open-vocabulary image segmentation ⚡️
- Running on ZeroAgentsFeatured74
Draw To Search Art
🐠74Draw/upload image and search among WikiART using SigLIP
- Running on CPU UpgradeAgents23
Compare Clip Siglip
🏃23Compare strong zero-shot image classification models
- Runtime errorAgents14
Multilingual Zero Shot Image Clf
🏢14Comparing powerful multilingual zero-shot image clf models
-
BAAI/bunny-phi-2-siglip-lora
Text Generation • Updated • 175 • 48
- Running on ZeroAgentsFeatured83
UDOP
🏃83Generate answers or summaries from document images with prompts
- Configuration errorAgents40
Pix2struct
📚40Play with all the pix2struct variants in this d
- RunningAgents26
Compare Docvqa Models
🦀26Compare different visual question answering
- Runtime errorAgentsFeatured289
DocQuery — Document Query Engine
🦉289
-
Improved Baselines with Visual Instruction Tuning
Paper • 2310.03744 • Published • 39 -
DeepSeek-VL: Towards Real-World Vision-Language Understanding
Paper • 2403.05525 • Published • 49 -
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
Paper • 2308.12966 • Published • 12 -
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
Paper • 2404.01331 • Published • 27
-
google/owlvit-base-patch32
Zero-Shot Object Detection • 0.2B • Updated • 205k • 151 -
google/owlvit-base-patch16
Zero-Shot Object Detection • Updated • 31.9k • 14 -
google/owlvit-large-patch14
Zero-Shot Object Detection • Updated • 7.63k • 29 -
google/owlv2-base-patch16
Zero-Shot Object Detection • 0.2B • Updated • 87.3k • 30
-
google/owlvit-base-patch32
Zero-Shot Object Detection • 0.2B • Updated • 205k • 151 -
google/owlvit-base-patch16
Zero-Shot Object Detection • Updated • 31.9k • 14 -
google/owlvit-large-patch14
Zero-Shot Object Detection • Updated • 7.63k • 29 -
google/owlv2-base-patch16
Zero-Shot Object Detection • 0.2B • Updated • 87.3k • 30
-
google/owlvit-base-patch32
Zero-Shot Object Detection • 0.2B • Updated • 205k • 151 -
google/owlvit-base-patch16
Zero-Shot Object Detection • Updated • 31.9k • 14 -
google/owlvit-large-patch14
Zero-Shot Object Detection • Updated • 7.63k • 29 -
google/owlv2-base-patch16
Zero-Shot Object Detection • 0.2B • Updated • 87.3k • 30
- RunningAgents212
Vidore Leaderboard
🥇212Browse and compare visual document retrieval model scores
- Running on CPU UpgradeAgents1.03k
Open VLM Leaderboard
🌎1.03kVLMEvalKit Evaluation Results Collection
- RunningFeatured561
Vision Arena (Testing VLMs side-by-side)
🖼561Explore Vision Arena visual AI demo online
- Build errorAgentsFeatured85
SEED-Bench Leaderboard
🏆85Submit model evaluation results to leaderboard
- PausedAgents21
Video Llava
🐨21Generate descriptions by uploading images or videos
-
llava-hf/LLaVA-NeXT-Video-7B-hf
Video-Text-to-Text • 7B • Updated • 111k • 126 -
llava-hf/LLaVA-NeXT-Video-7B-DPO-hf
Video-Text-to-Text • 7B • Updated • 550 • 12 -
llava-hf/LLaVA-NeXT-Video-7B-32K-hf
Image-Text-to-Text • 8B • Updated • 703 • 9
-
NVEagle/Eagle-X5-13B
Image-Text-to-Text • 15B • Updated • 62 • 15 -
NVEagle/Eagle-X5-13B-Chat
Image-Text-to-Text • 15B • Updated • 161 • 28 -
NVEagle/Eagle-X5-7B
Image-Text-to-Text • 9B • Updated • 84 • 26 - Configuration errorAgents64
Eagle X5 13B Chat
🚀64Combine text and images to generate responses