Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context Paper • 2503.15338 • Published Mar 19, 2025
What Did I Just Say? Self-Listening for Full-Duplex Speech Models Paper • 2609.05592 • Published 21 days ago • 25
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words Paper • 2406.13340 • Published Jun 19, 2024
Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech Data Paper • 2203.17113 • Published Mar 31, 2022
LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERT Paper • 2203.15610 • Published Mar 29, 2022
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing Paper • 2110.07205 • Published Oct 14, 2021 • 6
EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models Paper • 2510.22758 • Published Oct 26, 2025 • 1
What Did I Just Say? Self-Listening for Full-Duplex Speech Models Paper • 2609.05592 • Published 21 days ago • 25
deepseek-ai/DeepSeek-R1-Distill-Qwen-32B Text Generation • 33B • Updated Feb 24, 2025 • 476k • • 1.62k
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing Paper • 2110.07205 • Published Oct 14, 2021 • 6
SpeechT5 Collection The SpeechT5 framework consists of a shared seq2seq and six modal-specific (speech/text) pre/post-nets that can address a few audio-related tasks. • 8 items • Updated May 1, 2025 • 28