view article Article YODAS v3: A 1 Million Hour Dataset for the Next Generation of Open Voice AI Research espnet • 2 days ago • 24
view article Article tokenizers v1: encode, decode and scaling, measured +2 ArthurZ, sbrandeis, mcpotato, lysandre • 9 days ago • 76
view article Article **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization** nvidia • 6 days ago • 65
view article Article Transformers now runs llama.cpp quants +1 marcsun13, ArthurZ, lysandre • 8 days ago • 85
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published 22 days ago • 165
view article Article Making open-source AI weather forecasting models easy to run hugging-science • 21 days ago • 42
view article Article Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI nico-martin, Xenova • 29 days ago • 83
view article Article The Open ASR Leaderboard Adds Its First Global South Language +8 bezzam, Shobhitbanga, manasdhir04, bhaskarJT, manmeet-voicearena, pareek-voicearena, Amritansh8675, sagarjain268380, hanuman44420, vanshikachhabra-voicearena • Aug 28 • 59
view article Article How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code nielsr • Aug 21 • 32
view article Article Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC ibm-granite • Aug 25 • 36
view article Article Measuring benchmark optimization in speech recognition +5 tlebryk02, bezzam, aliceebaird, dayllon, jpc, jens-hume-ai, tzirakis • Aug 21 • 67
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • Aug 14 • 217
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India Paper • 2604.19151 • Published Apr 21 • 2
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Paper • 2505.07916 • Published May 12, 2025 • 139
view article Article Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident +2 hlarcher, XciD, raphael-gl, chris-rannou • Jul 27 • 507
RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems Paper • 2607.14846 • Published Jul 16 • 11