File size: 9,079 Bytes
4123b95
 
 
 
 
 
ed957e6
 
4123b95
 
 
 
 
 
 
 
 
ed957e6
4123b95
ed957e6
 
 
 
4123b95
1d5a2fc
 
 
 
 
 
 
 
 
 
 
 
 
ed957e6
4123b95
1d5a2fc
ed957e6
1d5a2fc
8d263fe
ed957e6
8d263fe
ed957e6
8d263fe
ed957e6
8d263fe
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ed957e6
8d263fe
ed957e6
8d263fe
ed957e6
8d263fe
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4123b95
 
ed957e6
 
 
4123b95
 
 
ed957e6
 
4123b95
 
ed957e6
4123b95
 
 
 
 
ed957e6
4123b95
 
 
 
 
ed957e6
 
 
 
 
4123b95
 
 
 
ed957e6
 
 
 
 
39c6967
1d5a2fc
 
 
4123b95
ed957e6
4123b95
ed957e6
 
 
4123b95
ed957e6
4123b95
 
 
 
 
 
 
 
 
ed957e6
4123b95
 
 
 
 
 
 
 
 
ed957e6
 
 
4123b95
ed957e6
 
 
 
4123b95
 
 
ac37044
 
 
 
 
 
 
 
 
 
 
 
 
 
ed957e6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
---
language:
- en
license: cc-by-nc-4.0
pipeline_tag: text-to-speech
tags:
- full-duplex
- conversation-tts
- dialogue
- stereo
- xml
- qwen3-tts
base_model: Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
datasets:
- facebook/seamless-interaction
---

# DuDE: Full-Duplex Conversation TTS

DuDE is a full-duplex conversation TTS (text-to-speech) model that generates
two speakers together from an XML dialogue. It supports turn-taking,
backchannels, and overlapping speech, with a separate voice for each speaker.
The output is a 24 kHz stereo WAV: speaker A on the left, speaker B on the right.

**Author:** Cheng-Kuang Chang

## Training details

DuDE first self-distills Qwen3-TTS on independent utterances to adapt a shared
backbone to two interleaved channels. It is then fine-tuned on approximately
1,922 hours of synchronized Seamless Interaction conversations to learn
turn-taking, backchannels, and overlapping speech.

The XML frontend resolves turn dependencies into interleaved text inputs.
Reference-speaker embeddings condition the two voices, and each audio channel
has its own end-of-speech decision while sharing the same timeline.

## Audio samples

The samples use XML transcripts of conversations from
[Meta's Seamless Interaction dataset](https://huggingface.co/datasets/facebook/seamless-interaction),
produced using ASR and vocal-event annotation. The audio is generated by DuDE.

### Sample 1

<audio controls src="https://huggingface.co/penguinfish1688/duplexdataengine/resolve/main/examples/sample_1.wav"></audio>

[Download stereo audio](examples/sample_1.wav) 路 [Input XML](examples/sample_1.xml) 路 Voices: `voice_1` / `voice_2`

<details>
<summary>Show complete input XML</summary>

```xml
<duplex>
<B>So, <e1/></B>
<B start="e1">if you had a chance to <e2/></B>
<B start="e2">send a message back to <e3/> the future, or the past, <e4/></B>
<A start="e3"><sneeze/> <cough/></A>
<B start="e4">or reverse, or whatever, however, what would you do? <e5/> Well, if I had a chance, I probably</B>
<A start="e5">Well, if I had a chance, I'd probably. <e6/></A>
<A start="e6"><e7/> Where can I do both, or I have to do one or the other? <e8/></A>
<B start="e7">would do both. Probably,</B>
<B start="e8">no, you probably will do both regardless. <e9/></B>
<A start="e9">Well, most <e10/> likely, I'd send a message back to the past and be like, "Yo, <e11/> watch out! Don't go around that corner." <cough/> <e12/></A>
<B start="e10">I most likely would</B>
<B start="e11"><cough/></B>
<B start="e12">who? <e13/> <laugh/></B>
<A start="e13">I would send a message to myself. <e14/></A>
<B start="e14">Oh, cool. <e15/></B>
<A start="e15">I mean, who else would I contact in the past? <e16/></A>
<B start="e16">Would you like yourself back then? <e17/></B>
<A start="e17"><e18/> I don't listen. I have no fucking clue. <cough/> <e19/></A>
<B start="e18"><cough/></B>
<B start="e19">Okay. That far. Oh, great. <e20/> great! <e21/></B>
<A start="e20">I don't have a clue.</A>
<B start="e21"><gasp/> <sigh/> <sneeze/> Bye.</B>
</duplex>
```

</details>

### Sample 2

<audio controls src="https://huggingface.co/penguinfish1688/duplexdataengine/resolve/main/examples/sample_2.wav"></audio>

[Download stereo audio](examples/sample_2.wav) 路 [Input XML](examples/sample_2.xml) 路 Voices: `voice_3` / `voice_4`

<details>
<summary>Show complete input XML</summary>

```xml
<duplex>
<B>If we're to organize a fitness challenge, which <e1/></B>
<B start="e1">I know you love, <e2/></B>
<B start="e2">what would we include <e3/></B>
<B start="e3">for activities? <e4/></B>
<B start="e4">How would we make it engaging? <e5/></B>
<B start="e5">What <e6/> do you think?</B>
<A start="e6">The activities number one has got to be yoga, <e7/> probably.</A>
<B start="e7">Yoga. I'm on <e8/> board. <e9/></B>
<A start="e8">Stretching and</A>
<A start="e9">you know, relaxing and feeling good. <e10/></A>
<A start="e10">I I think weights are good too. I mean, just a little everything. <e11/></A>
<A start="e11">I think one extreme to the other, maybe. <e12/></A>
<B start="e12">Yeah, I mean, I'm definitely on board with yoga. I like mobility exercises. I <e13/></B>
<A start="e13">Yeah, I think if you <e14/> do,</A>
<B start="e14">don't know if I like the word extreme, but <e15/> you know.</B>
<A start="e15">yeah. <e16/></A>
<A start="e16">So the only thing missing would be so you have yoga, weights, and we need some kind of cardio. <e17/></A>
<A start="e17"><e18/> Any suggestions? <e19/></A>
<B start="e18">Yeah.</B>
<B start="e19">Like speed walking, I guess. <e20/></B>
<A start="e20">Maybe a treadmill. <e21/></A>
<B start="e21">Yeah. Yeah, I like <e22/> the treadmill.</B>
<A start="e22">Yeah, I think that would cover everything.</A>
</duplex>
```

</details>


## Download and run

Use Python 3.12 and a CUDA-capable NVIDIA GPU. The download includes the model,
audio codec, XML frontend, four voice templates, and inference code. Requirements
pin PyTorch and torchaudio 2.8.0; use the corresponding CUDA builds for your system.

```bash
pip install "huggingface_hub>=0.34,<1"
hf download penguinfish1688/duplexdataengine --local-dir DuDE
cd DuDE
pip install -r requirements.txt
python infer.py \
  --model . --xml examples/sample_1.xml \
  --voice-a voice_1 --voice-b voice_2 \
  --seed 20261060 --max-seconds 118 \
  --output dialogue.wav
```

Or use the Python API from the downloaded directory:

```python
from pathlib import Path
from dude_tts import DuDE

model = DuDE.from_pretrained(".", device="cuda")
audio = model.generate(
    Path("examples/sample_1.xml").read_text(),
    voice_a="voice_1", voice_b="voice_2",
    seed=20261060, max_seconds=118,
)
audio.save("dialogue.wav")
```

`max_seconds` is an output limit, not a target speaking duration. Generation
finishes when both speakers emit an end-of-speech token; `audio.eos` reports
whether each speaker finished. Leading silence and pauses are preserved, with
silence padding after the shorter channel ends. CPU inference is available
with `--device cpu`, but is slow.

Preset inference on an RTX PRO 6000 Blackwell used 11.6 GB of allocated GPU
memory and took about 1.1 seconds per second of stereo output. Weights are
stored in FP32; CUDA inference uses BF16 autocast.

## Writing a dialogue

Use `<A>` and `<B>` for the speakers. Exactly one initial turn has no `start`
attribute. Every later turn references an inline anchor with `start="e1"`,
`start="e2"`, and so on. Define each anchor once. The `<duplex>` wrapper is optional.

For turn-taking, put the anchor at the preceding turn's end:

```xml
<duplex>
  <A>Hello! How was your day?<e1/></A>
  <B start="e1">Pretty good. I finished the project.<e2/></B>
  <A start="e2">That's great! <laugh/></A>
</duplex>
```

For overlap, put the anchor where the other speaker should enter:

```xml
<duplex>
  <A>I walked to the park, <e1/>and then stopped for coffee.<e2/></A>
  <B start="e1">Mm-hmm.</B>
  <B start="e2">That sounds nice.</B>
</duplex>
```

Anchors express turn relationships; the model predicts the audio timing.
Exact entry times are not guaranteed. Colliding turns for the same speaker
follow their XML order. All references must exist, and turns must not form a cycle.

Vocal events: `<laugh/>`, `<cough/>`, `<breath/>`, `<sigh/>`, `<cry/>`,
`<sneeze/>`, and `<gasp/>`. Write spoken backchannels as ordinary text and
escape literal ampersands as `&amp;`. Attributes such as `type`, timestamps,
and duration are not supported.

## Voice templates

Use your own reference audio for either speaker by passing a file path to
`--voice-a` and `--voice-b`, or to `voice_a` and `voice_b` in Python:

```bash
python infer.py --model . --xml dialogue.xml \
  --voice-a speaker_a.wav --voice-b speaker_b.wav --output dialogue.wav
```

Use a clear, single-speaker WAV or FLAC clip between 0.5 and 60 seconds long.
No transcript is needed. Audio is converted to mono and resampled automatically.
You can mix a custom reference with an included voice, such as
`--voice-a speaker_a.wav --voice-b voice_2`.

If you do not have reference audio, four example voices are included for convenience:

| Template | Reference audio |
| --- | --- |
| `voice_1` | [Sample 1, speaker A](voices/voice_1.wav) |
| `voice_2` | [Sample 1, speaker B](voices/voice_2.wav) |
| `voice_3` | [Sample 2, speaker A](voices/voice_3.wav) |
| `voice_4` | [Sample 2, speaker B](voices/voice_4.wav) |

## License

Model weights, voice references, and conversation examples are available for
noncommercial use under [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/).
DuDE inference code is Apache-2.0; see [LICENSE-CODE](LICENSE-CODE).
Qwen components retain their Apache-2.0 notices in [LICENSE-QWEN](LICENSE-QWEN).

Conversation material and reference voices derive from Meta's
[Seamless Interaction dataset](https://huggingface.co/datasets/facebook/seamless-interaction),
also licensed CC BY-NC 4.0. This release includes derived transcripts, XML,
speaker embeddings, and synthesized recordings. Preserve the dataset attribution
when redistributing these assets.