Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
AI & ML interests
Computer Vision; Video Understanding; Action Recognition
Recent Activity
View all activity
Papers
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
-
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Paper • 2607.14935 • Published • 121 -
MCG-NJU/I3D-ViT
Image Feature Extraction • 0.4B • Updated • 232 • 12 -
MCG-NJU/VideoChat3-4B
Video-Text-to-Text • 4B • Updated • 2.55k • 28 -
MCG-NJU/VideoChat3-LV116k
Viewer • Updated • 8.07k • 18.7k • 15
VideoMAE Pre-trained Models
-
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
Paper • 2203.12602 • Published • 6 -
MCG-NJU/videomae-base
Video Classification • 94.2M • Updated • 326k • 58 -
MCG-NJU/videomae-base-finetuned-kinetics
Video Classification • 86.5M • Updated • 104k • 52 -
MCG-NJU/videomae-base-finetuned-ssv2
Video Classification • Updated • 3.39k • 7
-
MCG-NJU/LongVPO-Stage1-InternVL3-8B
Video-Text-to-Text • 8B • Updated • 22 -
MCG-NJU/LongVPO-Stage2-InternVL3-8B
Video-Text-to-Text • 8B • Updated • 23 -
MCG-NJU/LongVPO-Training-Data
Viewer • Updated • 14.5k • 29 -
LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization
Paper • 2602.02341 • Published • 1
Learning Human Skill Generators at Key-Step Levels
CaReBench data, CaRe models and all the contrastively trained MLLMs (including InternVL2, MiniCPM-V 2.6, LLaVA NeXT Video, Qwen2-VL and Tariser).
Generalist Video Temporal Grounding with Multimodal LLMs
-
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Paper • 2607.17423 • Published • 100 -
MCG-NJU/TimeLens2-93K
Viewer • Updated • 71.4k • 7.47k • 13 -
MCG-NJU/TimeLens2-8B
Video-Text-to-Text • 9B • Updated • 1.3k • 17 -
MCG-NJU/TimeLens2-4B
Video-Text-to-Text • 4B • Updated • 1.34k • 13
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
-
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
Paper • 2601.23224 • Published • 6 -
MCG-NJU/Video-o3_RL
Video-Text-to-Text • 8B • Updated • 33 • 2 -
MCG-NJU/Video-o3_SFT_RL
Video-Text-to-Text • 8B • Updated • 58 • 2 -
MCG-NJU/Seeker-173K
Preview • Updated • 186 • 5
Sports Video Understanding Benchmarks
-
SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation
Paper • 2511.19320 • Published • 28 -
MCG-NJU/SteadyDancer-14B
Image-to-Video • 16B • Updated • 631 • 69 -
MCG-NJU/X-Dance
Viewer • Updated • 36 • 257 • 20 -
MCG-NJU/SteadyDancer-GGUF
Image-to-Video • 16B • Updated • 1.33k • 25
Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
Generalist Video Temporal Grounding with Multimodal LLMs
-
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Paper • 2607.17423 • Published • 100 -
MCG-NJU/TimeLens2-93K
Viewer • Updated • 71.4k • 7.47k • 13 -
MCG-NJU/TimeLens2-8B
Video-Text-to-Text • 9B • Updated • 1.3k • 17 -
MCG-NJU/TimeLens2-4B
Video-Text-to-Text • 4B • Updated • 1.34k • 13
-
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Paper • 2607.14935 • Published • 121 -
MCG-NJU/I3D-ViT
Image Feature Extraction • 0.4B • Updated • 232 • 12 -
MCG-NJU/VideoChat3-4B
Video-Text-to-Text • 4B • Updated • 2.55k • 28 -
MCG-NJU/VideoChat3-LV116k
Viewer • Updated • 8.07k • 18.7k • 15
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
-
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
Paper • 2601.23224 • Published • 6 -
MCG-NJU/Video-o3_RL
Video-Text-to-Text • 8B • Updated • 33 • 2 -
MCG-NJU/Video-o3_SFT_RL
Video-Text-to-Text • 8B • Updated • 58 • 2 -
MCG-NJU/Seeker-173K
Preview • Updated • 186 • 5
VideoMAE Pre-trained Models
-
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
Paper • 2203.12602 • Published • 6 -
MCG-NJU/videomae-base
Video Classification • 94.2M • Updated • 326k • 58 -
MCG-NJU/videomae-base-finetuned-kinetics
Video Classification • 86.5M • Updated • 104k • 52 -
MCG-NJU/videomae-base-finetuned-ssv2
Video Classification • Updated • 3.39k • 7
Sports Video Understanding Benchmarks
-
MCG-NJU/LongVPO-Stage1-InternVL3-8B
Video-Text-to-Text • 8B • Updated • 22 -
MCG-NJU/LongVPO-Stage2-InternVL3-8B
Video-Text-to-Text • 8B • Updated • 23 -
MCG-NJU/LongVPO-Training-Data
Viewer • Updated • 14.5k • 29 -
LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization
Paper • 2602.02341 • Published • 1
-
SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation
Paper • 2511.19320 • Published • 28 -
MCG-NJU/SteadyDancer-14B
Image-to-Video • 16B • Updated • 631 • 69 -
MCG-NJU/X-Dance
Viewer • Updated • 36 • 257 • 20 -
MCG-NJU/SteadyDancer-GGUF
Image-to-Video • 16B • Updated • 1.33k • 25
Learning Human Skill Generators at Key-Step Levels
CaReBench data, CaRe models and all the contrastively trained MLLMs (including InternVL2, MiniCPM-V 2.6, LLaVA NeXT Video, Qwen2-VL and Tariser).