Reorganize model card around training, samples, and usage
Browse files
README.md
CHANGED
|
@@ -22,17 +22,24 @@ two speakers together from an XML dialogue. It supports turn-taking,
|
|
| 22 |
backchannels, and overlapping speech, with a separate voice for each speaker.
|
| 23 |
The output is a 24 kHz stereo WAV: speaker A on the left, speaker B on the right.
|
| 24 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
## Audio samples
|
| 26 |
|
| 27 |
-
The
|
| 28 |
[Meta's Seamless Interaction dataset](https://huggingface.co/datasets/facebook/seamless-interaction),
|
| 29 |
-
produced using ASR and vocal-event annotation.
|
| 30 |
-
of their four speakers appeared in DuDE's training data.** We checked the
|
| 31 |
-
training manifests and retained training caches; see [provenance.json](provenance.json).
|
| 32 |
-
|
| 33 |
-
The audio is generated by DuDE using speaker reference clips supplied at
|
| 34 |
-
inference time. Both channels are normalized for audibility without shifting
|
| 35 |
-
their shared timeline.
|
| 36 |
|
| 37 |
### Sample 1
|
| 38 |
|
|
@@ -151,11 +158,9 @@ whether each speaker finished. Leading silence and pauses are preserved, with
|
|
| 151 |
silence padding after the shorter channel ends. CPU inference is available
|
| 152 |
with `--device cpu`, but is slow.
|
| 153 |
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
the model. Peak allocated GPU memory was 11.6 GB. Both speakers finished in
|
| 158 |
-
both samples. See [verification.json](verification.json).
|
| 159 |
|
| 160 |
## Writing a dialogue
|
| 161 |
|
|
@@ -194,8 +199,7 @@ and duration are not supported.
|
|
| 194 |
|
| 195 |
## Voice templates
|
| 196 |
|
| 197 |
-
|
| 198 |
-
these four presets; it does not accept uploaded voice clips.
|
| 199 |
|
| 200 |
| Template | Reference audio |
|
| 201 |
| --- | --- |
|
|
@@ -204,31 +208,6 @@ these four presets; it does not accept uploaded voice clips.
|
|
| 204 |
| `voice_3` | [Sample 2, speaker A](voices/voice_3.wav) |
|
| 205 |
| `voice_4` | [Sample 2, speaker B](voices/voice_4.wav) |
|
| 206 |
|
| 207 |
-
## Technical details
|
| 208 |
-
|
| 209 |
-
DuDE uses a shared Qwen3-TTS backbone for both speakers. A deterministic
|
| 210 |
-
frontend resolves the XML turn dependencies and densely interleaves the two
|
| 211 |
-
text channels. Voice embeddings condition each speaker, and the model generates
|
| 212 |
-
both audio channels autoregressively with independent end-of-speech decisions.
|
| 213 |
-
The codec decodes them onto a shared timeline.
|
| 214 |
-
|
| 215 |
-
The model was trained on approximately 1,922 hours of synchronized Seamless
|
| 216 |
-
Interaction conversations. The release contains merged model weights in FP32;
|
| 217 |
-
CUDA inference uses BF16 autocast.
|
| 218 |
-
|
| 219 |
-
## Evaluation
|
| 220 |
-
|
| 221 |
-
On ten held-out English conversations, both speakers finished in every example.
|
| 222 |
-
Automated word error rate was **15.70%** for generated speech and **12.65%**
|
| 223 |
-
for an ASR transcription of the original recordings. This is a small evaluation
|
| 224 |
-
set, without a human listening study.
|
| 225 |
-
|
| 226 |
-
Timing is less reliable: lexical timing coverage was **44.79%**, with **72.09%**
|
| 227 |
-
agreement among measured timing constraints. The timing check did not meet its
|
| 228 |
-
acceptance criteria. Outputs can contain wrong words, long pauses, missed vocal
|
| 229 |
-
events, or inaccurate overlaps. Other languages and individual vocal events
|
| 230 |
-
have not been separately evaluated.
|
| 231 |
-
|
| 232 |
## License
|
| 233 |
|
| 234 |
Model weights, voice references, and conversation examples are available for
|
|
@@ -241,18 +220,3 @@ Conversation material and reference voices derive from Meta's
|
|
| 241 |
also licensed CC BY-NC 4.0. This release includes derived transcripts, XML,
|
| 242 |
speaker embeddings, and synthesized recordings. Preserve the dataset attribution
|
| 243 |
when redistributing these assets.
|
| 244 |
-
|
| 245 |
-
## Citation
|
| 246 |
-
|
| 247 |
-
```bibtex
|
| 248 |
-
@misc{chang2026dude,
|
| 249 |
-
author = {Cheng-Kuang Chang},
|
| 250 |
-
title = {DuDE: Full-Duplex Conversation TTS},
|
| 251 |
-
year = {2026},
|
| 252 |
-
url = {https://huggingface.co/penguinfish1688/duplexdataengine}
|
| 253 |
-
}
|
| 254 |
-
```
|
| 255 |
-
|
| 256 |
-
Please also cite [Qwen3-TTS](https://github.com/QwenLM/Qwen3-TTS) and
|
| 257 |
-
[Seamless Interaction](https://huggingface.co/datasets/facebook/seamless-interaction)
|
| 258 |
-
when using this model.
|
|
|
|
| 22 |
backchannels, and overlapping speech, with a separate voice for each speaker.
|
| 23 |
The output is a 24 kHz stereo WAV: speaker A on the left, speaker B on the right.
|
| 24 |
|
| 25 |
+
**Author:** Cheng-Kuang Chang
|
| 26 |
+
|
| 27 |
+
## Training details
|
| 28 |
+
|
| 29 |
+
DuDE first self-distills Qwen3-TTS on independent utterances to adapt a shared
|
| 30 |
+
backbone to two interleaved channels. It is then fine-tuned on approximately
|
| 31 |
+
1,922 hours of synchronized Seamless Interaction conversations to learn
|
| 32 |
+
turn-taking, backchannels, and overlapping speech.
|
| 33 |
+
|
| 34 |
+
The XML frontend resolves turn dependencies into interleaved text inputs.
|
| 35 |
+
Reference-speaker embeddings condition the two voices, and each audio channel
|
| 36 |
+
has its own end-of-speech decision while sharing the same timeline.
|
| 37 |
+
|
| 38 |
## Audio samples
|
| 39 |
|
| 40 |
+
The samples use XML transcripts of conversations from
|
| 41 |
[Meta's Seamless Interaction dataset](https://huggingface.co/datasets/facebook/seamless-interaction),
|
| 42 |
+
produced using ASR and vocal-event annotation. The audio is generated by DuDE.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
|
| 44 |
### Sample 1
|
| 45 |
|
|
|
|
| 158 |
silence padding after the shorter channel ends. CPU inference is available
|
| 159 |
with `--device cpu`, but is slow.
|
| 160 |
|
| 161 |
+
Preset inference on an RTX PRO 6000 Blackwell used 11.6 GB of allocated GPU
|
| 162 |
+
memory and took about 1.1 seconds per second of stereo output. Weights are
|
| 163 |
+
stored in FP32; CUDA inference uses BF16 autocast.
|
|
|
|
|
|
|
| 164 |
|
| 165 |
## Writing a dialogue
|
| 166 |
|
|
|
|
| 199 |
|
| 200 |
## Voice templates
|
| 201 |
|
| 202 |
+
Four example voices are included for convenience.
|
|
|
|
| 203 |
|
| 204 |
| Template | Reference audio |
|
| 205 |
| --- | --- |
|
|
|
|
| 208 |
| `voice_3` | [Sample 2, speaker A](voices/voice_3.wav) |
|
| 209 |
| `voice_4` | [Sample 2, speaker B](voices/voice_4.wav) |
|
| 210 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 211 |
## License
|
| 212 |
|
| 213 |
Model weights, voice references, and conversation examples are available for
|
|
|
|
| 220 |
also licensed CC BY-NC 4.0. This release includes derived transcripts, XML,
|
| 221 |
speaker embeddings, and synthesized recordings. Preserve the dataset attribution
|
| 222 |
when redistributing these assets.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|