VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 15 days ago • 170
view article Article Multimodal Embedding & Reranker Models with Sentence Transformers tomaarsen • Apr 9 • 68
video-SALMONN 2 Collection video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions. • 11 items • Updated Mar 21 • 2
view article Article VLX-Flow: Continuous Video Understanding for Real-Time Multimodal Interaction omlab • Jun 27 • 15
view article Article Profiling in PyTorch (Part 3): Attention is all you profile +2 ariG23498, sergiopaniego, sayakpaul, ror • 21 days ago • 40
Granite 4.1 Language Models Collection Efficient language models for multilingual generation, coding, RAG, and AI assistant workflows. • 6 items • Updated Apr 29 • 63
ReaderLM-v2: Small Language Model for HTML to Markdown and JSON Paper • 2503.01151 • Published Mar 3, 2025 • 5
DiarizationLM: Speaker Diarization Post-Processing with Large Language Models Paper • 2401.03506 • Published Jan 7, 2024 • 16
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS Paper • 2604.11269 • Published Apr 13 • 2
M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Paper • 2506.14427 • Published Jun 17, 2025 • 1
MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization Paper • 2601.01554 • Published Jan 4 • 65
SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models Paper • 2508.06372 • Published Aug 8, 2025 • 4
VideoPrism Collection VideoPrism is a foundational video encoder that enables state-of-the-art performance on a large variety of video understanding tasks. • 5 items • Updated 9 days ago • 21
Gemma 4 Collection Gemma 4 is Google's new model family including including E2B, E4B, 26B-A4B, and 31B. • 43 items • Updated 12 days ago • 250