vid analysis
updated
prithivMLmods/Qwen3-VL-8B-Abliterated-Caption-it
Image-Text-to-Text
• 9B • Updated • 35
mradermacher/Qwen3-VL-8B-NSFW-Caption-V4.5-GGUF
8B • Updated • 10.5k
• 73
prithivMLmods/Qwen3-VisionCaption-2B
Image-Text-to-Text
• 2B • Updated • 20
• 5
msrcam/Qwen3-VL-2B-Instruct-heretic
ghost-actual/Qwen3.5-4B-Claude-Opus-4.6-Distilled-heretic
Text Generation
• 5B • Updated • 38
• 3
Image-Text-to-Text
• 3B • Updated • 24
• 5
Image-Text-to-Text
• 5B • Updated • 6.55M
• • 796
bobber/routangseng-qwen35-0.8b-abliterated-onnx
Image-Text-to-Text
• Updated • 25
bobber/routangseng-0.8b-hottake-onnx
Image-Text-to-Text
• Updated • 7
Caplin43/multimodal-vision-language-mini
Image-to-Text
• Updated • 19
Ytgetahun/visual-narrator-llm
Image-to-Text
• Updated
Ytgetahun/visual-narrator-vlm
0.2B • Updated • 4
allenai/MolmoPoint-Vid-4B
Video-Text-to-Text
• 5B • Updated • 861
• 13
TencentARC/ARC-Qwen-Video-7B-Narrator
Video-Text-to-Text
• 9B • Updated • 50
• 11
Video-Text-to-Text
• Updated • 72
• 32
Image-Text-to-Text
• 5B • Updated • 63.3k
• 51
DAMO-NLP-SG/VideoLLaMA3-2B
Video-Text-to-Text
• 2B • Updated • 1.97k
• 21
u94fmn391j/SAVANT-scene-description-lora
Image-to-Text
• Updated • 5
VINAY-UMRETHE/SigMamba-V1-Large
Video Classification
• 0.9B • Updated • 8
• 6
qoranet/QORA-Vision-Video
Video Classification
• Updated • 8
sumit7488/TimesFormer_Baseline
Video Classification
• 0.1B • Updated • 3
StreamFormer/streamformer-timesformer
Video Classification
• 0.1B • Updated • 118
• 4
facebook/vjepa2-vitg-fpc64-256
Video Classification
• 1B • Updated • 211k
• 57
Image-Text-to-Text
• 3B • Updated • 7.41k
• 201
Video-Text-to-Text
• Updated • 2
BidirLM/BidirLM-Omni-2.5B-Embedding
Sentence Similarity
• 2B • Updated • 576
• 45
Image-Text-to-Text
• Updated • 17
• 14
Video-Text-to-Text
• 0.9B • Updated • 8
• 1
Video-Text-to-Text
• 2B • Updated • 6.62k
• 575
LongVie 2: Multimodal Controllable Ultra-Long Video World Model
Paper
• 2512.13604
• Published • 76
Image-Text-to-Text
• 4B • Updated • 464k
• 2.87k
prithivMLmods/Gemma4-BLIP3o-Captioner-5B
Image-Text-to-Text
• 5B • Updated • 608
• 3
lewiswatson/Frame2KG-LFM-2.5-450m-JSON
Image-Text-to-Text
• 0.4B • Updated • 52
69.7M • Updated • 16.4k
• 2