view article Article Scaling Real-Time Voice Agents with Cache-Aware Streaming ASR nvidia • Jan 5 • 92
Running on Zero Agents Featured 116 Marigold V2 🌼 116 Depth, surface normals, and albedo from a single image
Running 96 Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation 🌼 96 Depth, surface normals, and albedo from a single image
VibeVoice Collection Frontier Text-to-Speech Models https://microsoft.github.io/VibeVoice/ • 11 items • Updated 22 days ago • 264
Running on Zero Agents Featured 482 Parakeet-TDT-0.6b-V2 482 Transcribe audio files with timestamps and downloadable subtitles
Running on CPU Upgrade Featured 3.31k The Smol Training Playbook 📚 3.31k The secrets to building world-class LLMs
Running Featured 94 Parakeet STT Progressive Transcription 🎤 94 Transcribe speech to text instantly with WebGPU acceleration
openai/whisper-large-v3-turbo Automatic Speech Recognition • 0.8B • Updated Oct 4, 2024 • 6.53M • • 3.39k
SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations Paper • 2108.01073 • Published Aug 2, 2021 • 9
Running on Zero Agents Featured 148 Qwen3-ASR Demo 🎙 148 Transcribe audio to text with timestamps and visualization