Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
blanchefort 's Collections
ML
Embedders
Medical
VLA models
Audio
Translate
OCR
OmniModels
Edge models
Video encoders
Judge
Datasets for Embodied
Ru text encoders
Text2Image
VLMs

Audio

updated about 24 hours ago
Upvote
-

  • nvidia/audio-flamingo-3-hf

    Audio-Text-to-Text • 8B • Updated Apr 13 • 56.3k • 195

  • facebook/sam-audio-large

    Updated Dec 30, 2025 • 11.7k • 455

  • google/medasr

    Automatic Speech Recognition • 0.1B • Updated May 26 • 21.1k • 368

  • FunAudioLLM/Fun-CosyVoice3-0.5B-2512

    Text-to-Speech • Updated Feb 3 • 197k • 651

  • facebook/sam-audio-large-tv

    Updated Dec 30, 2025 • 3.67k • 33

  • Qwen/Qwen3-TTS-12Hz-0.6B-Base

    Text-to-Speech • 0.9B • Updated Jan 29 • 605k • 307

  • lab260/spectra_0

    Audio Classification • 0.3B • Updated Jun 25 • 7

  • nvidia/Nemotron-3-Diarization

    Voice Activity Detection • 99.2M • Updated 7 days ago • 36.4k • 549

  • m-a-p/MERT-v2-FullSong

    Feature Extraction • 0.6B • Updated 1 day ago • 17k • 44
Upvote
-
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs