VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 11 days ago • 169
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 153
TIPSv2 Collection TIPSv2 foundational vision-language models. Webpage: https://gdm-tipsv2.github.io/ • 9 items • Updated 5 days ago • 39
RDP LoRA: Geometry-Driven Identification for Parameter-Efficient Adaptation in Large Language Models Paper • 2604.19321 • Published Apr 21 • 8
SigLino: Vision Foundation Models (SigLIP2 + DINOv3) Collection Vision encoders distilled from DINOv3 and SigLIP2 (MoE & Dense). CVPR 2026. • 6 items • Updated Apr 10 • 17
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking Paper • 2601.04720 • Published Jan 8 • 59
view article Article Fine-Tuning MetaCLIP-2 for Image Classification on Downstream Tasks prithivMLmods • Nov 15, 2025 • 7
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning Paper • 2511.02818 • Published Nov 4, 2025 • 15
SelectMix: Enhancing Label Noise Robustness through Targeted Sample Mixing Paper • 2509.11265 • Published Sep 14, 2025 • 1
Intra-Cluster Mixup: An Effective Data Augmentation Technique for Complementary-Label Learning Paper • 2509.17971 • Published Sep 22, 2025 • 1
Token Activation Map to Visually Explain Multimodal LLMs Paper • 2506.23270 • Published Jun 29, 2025 • 5
LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models Paper • 2504.14032 • Published Apr 18, 2025 • 7
E^2Rank: Your Text Embedding can Also be an Effective and Efficient Listwise Reranker Paper • 2510.22733 • Published Oct 26, 2025 • 32
Heavy Labels Out! Dataset Distillation with Label Space Lightening Paper • 2408.08201 • Published Aug 15, 2024 • 22
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders Paper • 2510.19779 • Published Oct 22, 2025 • 62