Full-Duplex Speech Models
Conversational models that listen and speak simultaneously, including voice control, interactivity tuning, and multimodal streaming.
Audio-to-Audio • 8B • Updated • 175k • • 2.78kNote Full-duplex speech conversation with voice and role prompts, built on Moshi. English. Gated; NVIDIA Open Model License.
kyutai/moshiko-pytorch-bf16
8B • Updated • 176k • 258Note Original Moshi-family full-duplex speech conversation baseline; Moshiko voice, PyTorch BF16 checkpoint. English. CC-BY-4.0.
kyutai/personaplex-rl-seamless
Audio-to-Audio • 8B • Updated • 4.5k • 38Note PersonaPlex post-trained on Seamless Interaction to improve pause handling, turn-taking, backchanneling, and interruption behavior.
openbmb/MiniCPM-o-4_5
Any-to-Any • 9B • Updated • 735k • 1.5kNote Multimodal full-duplex model: concurrent audio/video input and text/speech output. Use its duplex inference mode; also supports separate half-duplex modes.
HIT-TMG/Lychee-FD
13B • Updated • 330 • 5Note Research full-duplex speech checkpoint for Chinese/English. Demo requires the separate Token2Wav vocoder from Step-Audio-2-mini. Apache-2.0 metadata.
BayLing-Models/BayLing-Duplex
516k • Updated • 189 • 7Note Research model with native listening/speaking and learned interruption/turn timing. Requires separate GLM-4-Voice tokenizer and decoder; see LICENSE/NOTICE for terms.
MuyeHuang/DuplexOmni
35B • Updated • 991 • 3Note Research multimodal full-duplex Thinker/Talker model. Card notes unexpected silence/speech-quality issues and recommends at least 8 H20 GPUs for low latency. Apache-2.0.
mindlogicinc/context-spanning-7b
Audio-to-Audio • UpdatedNote Recent PersonaPlex-derived full-duplex research model. Published runtime uses external router-LLM and ASR backends. NVIDIA Open Model License; not a standalone deployment.
nu-dialogue/j-moshi
8B • Updated • 59 • 15Note Japanese full-duplex Moshi adaptation with overlapping speech and backchannels. Official research checkpoint; j-moshi-ext is the companion variant trained with additional synthetic dialogue.
VITA-MLLM/Freeze-Omni
Updated • 21Note Duplex speech system with chunk-level interruption/state prediction and streaming input/output. Requires separate Qwen2-7B-Instruct weights; server uses VAD and cache scheduling.
gpt-omni/mini-omni2
Any-to-Any • Updated • 138 • 289Note Early multimodal duplex research baseline with end-to-end streaming speech output. Paper demonstrates command-based interruption; this is narrower than unrestricted conversational overlap handling.
tsinghua-ee/ELLSA
18B • Updated • 50 • 5Note End-to-end streaming full-duplex vision, speech, text and action model. Includes speaking-while-acting and action barge-in; research setup requires multiple companion checkpoints.
sbintuitions/DuplexCascade
Text Generation • Updated • 4Note Cascaded full-duplex system checkpoint: Qwen2-7B-Instruct fine-tuned for VAD-free micro-turn dialogue control. Requires external ASR and TTS; this repository is the LLM component.