Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
zhangyunfeng's picture

zhangyunfeng

yunfeng
6 28
·
  • yunfengsay

AI & ML interests

None yet

Organizations

None yet

upvoted a collection 3 months ago

jina-embeddings-v5-omni

Collection
Multimodal (text + image + video + audio) embedding models aligned with jina-embeddings-v5-text-*. Two sizes, four task variants each. • 27 items • Updated 8 days ago • 36
upvoted an article over 1 year ago
view article
Article

seemore: Implement a Vision Language Model from Scratch

AviSoori1x
•
Jun 23, 2024
• 111
upvoted a collection almost 2 years ago

Multimodal RAG

Collection
9 items • Updated Mar 2 • 31
upvoted 2 collections over 2 years ago

MGM

Collection
Official model collection for the paper "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models" • 13 items • Updated May 3, 2024 • 47

From screenshots to HTML

Collection
WebSight is a dataset of 823,000 HTML/CSS codes representing synthetically generated English websites, each accompanied by a corresponding screenshot. • 4 items • Updated Apr 15, 2024 • 22
upvoted a paper over 2 years ago

Lumiere: A Space-Time Diffusion Model for Video Generation

Paper • 2401.12945 • Published Jan 23, 2024 • 86
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs