[ICLR 2026] VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
Ye Liu PRO
yeliudev
AI & ML interests
Vision & Language
Recent Activity
updated a Space 19 days ago
yeliudev/VideoMind-2B upvoted a paper about 2 months ago
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories updated a Space 2 months ago
yeliudev/VideoMind-2BOrganizations
UniPixel
[NeurIPS 2025] UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- Configuration errorAgents6
UniPixel
🔮6An MLLM for Unified Object Referring and Segmentation
-
PolyU-ChenLab/UniPixel-3B
Video-Text-to-Text • 4B • Updated • 258 • 3 -
PolyU-ChenLab/UniPixel-7B
Video-Text-to-Text • 8B • Updated • 328 • 1 -
PolyU-ChenLab/UniPixel-SFT-1M
Preview • Updated • 254 • 3
R2-Tuning
[ECCV 2024] R2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding
VideoMind
[ICLR 2026] VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
UniPixel
[NeurIPS 2025] UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- Configuration errorAgents6
UniPixel
🔮6An MLLM for Unified Object Referring and Segmentation
-
PolyU-ChenLab/UniPixel-3B
Video-Text-to-Text • 4B • Updated • 258 • 3 -
PolyU-ChenLab/UniPixel-7B
Video-Text-to-Text • 8B • Updated • 328 • 1 -
PolyU-ChenLab/UniPixel-SFT-1M
Preview • Updated • 254 • 3
E.T. Bench
[NeurIPS 2024] E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
R2-Tuning
[ECCV 2024] R2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding