penguinfish1688 commited on
Commit
1d5a2fc
·
verified ·
1 Parent(s): ed957e6

Reorganize model card around training, samples, and usage

Browse files
Files changed (1) hide show
  1. README.md +19 -55
README.md CHANGED
@@ -22,17 +22,24 @@ two speakers together from an XML dialogue. It supports turn-taking,
22
  backchannels, and overlapping speech, with a separate voice for each speaker.
23
  The output is a 24 kHz stereo WAV: speaker A on the left, speaker B on the right.
24
 
 
 
 
 
 
 
 
 
 
 
 
 
 
25
  ## Audio samples
26
 
27
- The inputs below are XML transcripts of two held-out conversations from
28
  [Meta's Seamless Interaction dataset](https://huggingface.co/datasets/facebook/seamless-interaction),
29
- produced using ASR and vocal-event annotation. **Neither conversation nor any
30
- of their four speakers appeared in DuDE's training data.** We checked the
31
- training manifests and retained training caches; see [provenance.json](provenance.json).
32
-
33
- The audio is generated by DuDE using speaker reference clips supplied at
34
- inference time. Both channels are normalized for audibility without shifting
35
- their shared timeline.
36
 
37
  ### Sample 1
38
 
@@ -151,11 +158,9 @@ whether each speaker finished. Leading silence and pauses are preserved, with
151
  silence padding after the shorter channel ends. CPU inference is available
152
  with `--device cpu`, but is slow.
153
 
154
- An anonymous download was tested in a fresh environment on one RTX PRO 6000
155
- Blackwell with PyTorch 2.8.0+cu128. Generating 47.6 and 54.4 seconds of stereo
156
- audio took 50.7 and 57.4 seconds, respectively, plus about 28 seconds to load
157
- the model. Peak allocated GPU memory was 11.6 GB. Both speakers finished in
158
- both samples. See [verification.json](verification.json).
159
 
160
  ## Writing a dialogue
161
 
@@ -194,8 +199,7 @@ and duration are not supported.
194
 
195
  ## Voice templates
196
 
197
- Choose a template independently for each speaker. The current API supports
198
- these four presets; it does not accept uploaded voice clips.
199
 
200
  | Template | Reference audio |
201
  | --- | --- |
@@ -204,31 +208,6 @@ these four presets; it does not accept uploaded voice clips.
204
  | `voice_3` | [Sample 2, speaker A](voices/voice_3.wav) |
205
  | `voice_4` | [Sample 2, speaker B](voices/voice_4.wav) |
206
 
207
- ## Technical details
208
-
209
- DuDE uses a shared Qwen3-TTS backbone for both speakers. A deterministic
210
- frontend resolves the XML turn dependencies and densely interleaves the two
211
- text channels. Voice embeddings condition each speaker, and the model generates
212
- both audio channels autoregressively with independent end-of-speech decisions.
213
- The codec decodes them onto a shared timeline.
214
-
215
- The model was trained on approximately 1,922 hours of synchronized Seamless
216
- Interaction conversations. The release contains merged model weights in FP32;
217
- CUDA inference uses BF16 autocast.
218
-
219
- ## Evaluation
220
-
221
- On ten held-out English conversations, both speakers finished in every example.
222
- Automated word error rate was **15.70%** for generated speech and **12.65%**
223
- for an ASR transcription of the original recordings. This is a small evaluation
224
- set, without a human listening study.
225
-
226
- Timing is less reliable: lexical timing coverage was **44.79%**, with **72.09%**
227
- agreement among measured timing constraints. The timing check did not meet its
228
- acceptance criteria. Outputs can contain wrong words, long pauses, missed vocal
229
- events, or inaccurate overlaps. Other languages and individual vocal events
230
- have not been separately evaluated.
231
-
232
  ## License
233
 
234
  Model weights, voice references, and conversation examples are available for
@@ -241,18 +220,3 @@ Conversation material and reference voices derive from Meta's
241
  also licensed CC BY-NC 4.0. This release includes derived transcripts, XML,
242
  speaker embeddings, and synthesized recordings. Preserve the dataset attribution
243
  when redistributing these assets.
244
-
245
- ## Citation
246
-
247
- ```bibtex
248
- @misc{chang2026dude,
249
- author = {Cheng-Kuang Chang},
250
- title = {DuDE: Full-Duplex Conversation TTS},
251
- year = {2026},
252
- url = {https://huggingface.co/penguinfish1688/duplexdataengine}
253
- }
254
- ```
255
-
256
- Please also cite [Qwen3-TTS](https://github.com/QwenLM/Qwen3-TTS) and
257
- [Seamless Interaction](https://huggingface.co/datasets/facebook/seamless-interaction)
258
- when using this model.
 
22
  backchannels, and overlapping speech, with a separate voice for each speaker.
23
  The output is a 24 kHz stereo WAV: speaker A on the left, speaker B on the right.
24
 
25
+ **Author:** Cheng-Kuang Chang
26
+
27
+ ## Training details
28
+
29
+ DuDE first self-distills Qwen3-TTS on independent utterances to adapt a shared
30
+ backbone to two interleaved channels. It is then fine-tuned on approximately
31
+ 1,922 hours of synchronized Seamless Interaction conversations to learn
32
+ turn-taking, backchannels, and overlapping speech.
33
+
34
+ The XML frontend resolves turn dependencies into interleaved text inputs.
35
+ Reference-speaker embeddings condition the two voices, and each audio channel
36
+ has its own end-of-speech decision while sharing the same timeline.
37
+
38
  ## Audio samples
39
 
40
+ The samples use XML transcripts of conversations from
41
  [Meta's Seamless Interaction dataset](https://huggingface.co/datasets/facebook/seamless-interaction),
42
+ produced using ASR and vocal-event annotation. The audio is generated by DuDE.
 
 
 
 
 
 
43
 
44
  ### Sample 1
45
 
 
158
  silence padding after the shorter channel ends. CPU inference is available
159
  with `--device cpu`, but is slow.
160
 
161
+ Preset inference on an RTX PRO 6000 Blackwell used 11.6 GB of allocated GPU
162
+ memory and took about 1.1 seconds per second of stereo output. Weights are
163
+ stored in FP32; CUDA inference uses BF16 autocast.
 
 
164
 
165
  ## Writing a dialogue
166
 
 
199
 
200
  ## Voice templates
201
 
202
+ Four example voices are included for convenience.
 
203
 
204
  | Template | Reference audio |
205
  | --- | --- |
 
208
  | `voice_3` | [Sample 2, speaker A](voices/voice_3.wav) |
209
  | `voice_4` | [Sample 2, speaker B](voices/voice_4.wav) |
210
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
211
  ## License
212
 
213
  Model weights, voice references, and conversation examples are available for
 
220
  also licensed CC BY-NC 4.0. This release includes derived transcripts, XML,
221
  speaker embeddings, and synthesized recordings. Preserve the dataset attribution
222
  when redistributing these assets.