wzh
hg2wzh
AI & ML interests
None yet
Recent Activity
liked a model 4 days ago
openbmb/MiniCPM-V-4.6 liked a dataset 17 days ago
sensenova/SenseNova-Vision-Corpus-50M liked a Space about 1 month ago
HuggingFaceM4/encoder-free-vlmOrganizations
None yet
Datasets
Embedding
VLMs
-
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Paper • 2409.12191 • Published • 80 -
Multimodal Latent Language Modeling with Next-Token Diffusion
Paper • 2412.08635 • Published • 50 -
ATH-MaaS/Ovis2-2B
Image-Text-to-Text • 2B • Updated • 329 • 60 -
DAMO-NLP-SG/VideoLLaMA3-2B
Video-Text-to-Text • 2B • Updated • 2.59k • 21
Eval-bench
Text-to-Image
Datasets
Reasoning
Embedding
CLIP series
VLMs
-
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Paper • 2409.12191 • Published • 80 -
Multimodal Latent Language Modeling with Next-Token Diffusion
Paper • 2412.08635 • Published • 50 -
ATH-MaaS/Ovis2-2B
Image-Text-to-Text • 2B • Updated • 329 • 60 -
DAMO-NLP-SG/VideoLLaMA3-2B
Video-Text-to-Text • 2B • Updated • 2.59k • 21
LLMs