kadirnar commited on
Commit
a2f53be
·
verified ·
1 Parent(s): 0d35703

showcase: step 100k: 5 listening samples + the prompt-choice comparison (3 items with their old prompt) (AudioSeal-watermarked), delete 85 files the manifest no longer lists (an older listening set)

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. .gitattributes +6 -0
  2. README.md +79 -271
  3. prompts/02-question.wav +2 -2
  4. prompts/03-numbers.wav +1 -1
  5. prompts/05-long.wav +2 -2
  6. samples/{step_0030000 → prompt_choice/old_output}/02-question.wav +1 -1
  7. samples/{step_0030000 → prompt_choice/old_output}/03-numbers.wav +1 -1
  8. samples/{step_0020000 → prompt_choice/old_output}/05-long.wav +1 -1
  9. samples/{step_0020000 → prompt_choice/old_prompt}/02-question.wav +2 -2
  10. samples/{step_0040000 → prompt_choice/old_prompt}/03-numbers.wav +2 -2
  11. samples/{step_0020000/01-short.wav → prompt_choice/old_prompt/05-long.wav} +2 -2
  12. samples/step_0020000/03-numbers.wav +0 -3
  13. samples/step_0020000/04-conversational.wav +0 -3
  14. samples/step_0030000/01-short.wav +0 -3
  15. samples/step_0030000/04-conversational.wav +0 -3
  16. samples/step_0030000/05-long.wav +0 -3
  17. samples/step_0040000/01-short.wav +0 -3
  18. samples/step_0040000/02-question.wav +0 -3
  19. samples/step_0040000/04-conversational.wav +0 -3
  20. samples/step_0040000/05-long.wav +0 -3
  21. samples/step_0050000/01-short.wav +0 -3
  22. samples/step_0050000/02-question.wav +0 -3
  23. samples/step_0050000/03-numbers.wav +0 -3
  24. samples/step_0050000/04-conversational.wav +0 -3
  25. samples/step_0050000/05-long.wav +0 -3
  26. samples/step_0060000/01-short.wav +0 -3
  27. samples/step_0060000/02-question.wav +0 -3
  28. samples/step_0060000/03-numbers.wav +0 -3
  29. samples/step_0060000/04-conversational.wav +0 -3
  30. samples/step_0060000/05-long.wav +0 -3
  31. samples/step_0070000/01-short.wav +0 -3
  32. samples/step_0070000/02-question.wav +0 -3
  33. samples/step_0070000/03-numbers.wav +0 -3
  34. samples/step_0070000/04-conversational.wav +0 -3
  35. samples/step_0070000/05-long.wav +0 -3
  36. samples/step_0080000/01-short.wav +0 -3
  37. samples/step_0080000/02-question.wav +0 -3
  38. samples/step_0080000/03-numbers.wav +0 -3
  39. samples/step_0080000/04-conversational.wav +0 -3
  40. samples/step_0080000/05-long.wav +0 -3
  41. samples/step_0090000/01-short.wav +0 -3
  42. samples/step_0090000/02-question.wav +0 -3
  43. samples/step_0090000/03-numbers.wav +0 -3
  44. samples/step_0090000/04-conversational.wav +0 -3
  45. samples/step_0090000/05-long.wav +0 -3
  46. samples/step_0100000/01-short.wav +1 -1
  47. samples/step_0100000/02-question.wav +2 -2
  48. samples/step_0100000/03-numbers.wav +2 -2
  49. samples/step_0100000/04-conversational.wav +1 -1
  50. samples/step_0100000/05-long.wav +2 -2
.gitattributes CHANGED
@@ -128,3 +128,9 @@ samples/step_0190000/02-question.wav filter=lfs diff=lfs merge=lfs -text
128
  samples/step_0190000/03-numbers.wav filter=lfs diff=lfs merge=lfs -text
129
  samples/step_0190000/04-conversational.wav filter=lfs diff=lfs merge=lfs -text
130
  samples/step_0190000/05-long.wav filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
128
  samples/step_0190000/03-numbers.wav filter=lfs diff=lfs merge=lfs -text
129
  samples/step_0190000/04-conversational.wav filter=lfs diff=lfs merge=lfs -text
130
  samples/step_0190000/05-long.wav filter=lfs diff=lfs merge=lfs -text
131
+ samples/prompt_choice/old_output/02-question.wav filter=lfs diff=lfs merge=lfs -text
132
+ samples/prompt_choice/old_output/03-numbers.wav filter=lfs diff=lfs merge=lfs -text
133
+ samples/prompt_choice/old_output/05-long.wav filter=lfs diff=lfs merge=lfs -text
134
+ samples/prompt_choice/old_prompt/02-question.wav filter=lfs diff=lfs merge=lfs -text
135
+ samples/prompt_choice/old_prompt/03-numbers.wav filter=lfs diff=lfs merge=lfs -text
136
+ samples/prompt_choice/old_prompt/05-long.wav filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -31,295 +31,93 @@ The name: **DAC** for the Semantic-DACVAE audio codec whose latents it generates
31
  English, **10k** for its training data, about ten thousand hours of speech.
32
  The code: [kadirnar/dacvae-next](https://github.com/kadirnar/dacvae-next/tree/roadmap/en-echo), Python package `mytts`.
33
 
34
- > **Status: training, step 190k of 200k.** A new checkpoint (a copy of the model saved at that point of its training)
35
- > appears here every 10k steps (from step 20k on), with its samples and scores: this page updates itself. Early checkpoints sound rougher than
36
- > later ones.
 
 
 
37
 
38
  **On this page:** [Listen](#listen) · [Results](#results) · [How to use](#how-to-use) · [Training](#training) · [Data](#data) · [Licence](#licence)
39
 
40
  ## Listen
41
 
42
- The newest checkpoint, **step 190k**, reads five texts. Each text is spoken in a different voice, copied from
43
- the short voice prompt next to it. The model never heard these voices in training (held-out voices of the provided data).
44
 
45
  <table>
46
- <tr><th>What the model reads</th><th>Voice prompt (the input)</th><th>Model output, step 190k</th></tr>
47
- <tr><td><b>Short sentence</b><br>I left my umbrella at the office again, so I&#x27;m definitely getting soaked on the way home.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/01-short.wav"></audio><br><sub>A higher voice (~202 Hz). It says: “I keep telling myself I just need to power through, like, push a little harder, and then it&#x27;ll be fine. But it&#x27;s not getting fine. It&#x27;s getting— it&#x27;s getting worse.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0190000/01-short.wav"></audio></td></tr>
48
- <tr><td><b>Question</b><br>Have you ever noticed that the quietest person in the room usually has the most interesting story to tell?</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/02-question.wav"></audio><br><sub>A lower voice (~149 Hz). It says: “Wait, that mole— has it always looked like that? No, stop.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0190000/02-question.wav"></audio></td></tr>
49
- <tr><td><b>Numbers, dates, abbreviations</b><br>Dr. Patel moved my appointment to Tuesday, March 3rd, at 4:15 p.m., and the co-pay went up from $20 to $35.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/03-numbers.wav"></audio><br><sub>A higher voice (~181 Hz). It says: “this woman on the train was literally eating a whole bag of chips, like, crunching so loud, and I&#x27;m sitting there like, okay, do I move? whatever.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0190000/03-numbers.wav"></audio></td></tr>
50
- <tr><td><b>Conversation (~10 s)</b><br>So I finally tried that new ramen place downtown, and honestly, it was worth the wait. The broth was amazing, but next time I&#x27;m definitely skipping the extra spicy option.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/04-conversational.wav"></audio><br><sub>A lower voice (~115 Hz). It says: “I&#x27;m sorry, but the subway doors closed on my bag. Again.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0190000/04-conversational.wav"></audio></td></tr>
51
- <tr><td><b>Long passage (~20 s)</b><br>When the storm passed, the whole neighborhood came outside to look at the damage, and although a few fences had fallen and the old oak tree had lost its biggest branch, everyone was relieved that nobody was hurt. They spent the rest of the afternoon clearing the street and sharing whatever food they had left.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/05-long.wav"></audio><br><sub>The deepest voice (~108 Hz). It says: “Honestly I think we need to just call IT and have them... you know, actually fix it this time. I&#x27;m too old for this.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0190000/05-long.wav"></audio></td></tr>
52
  </table>
53
 
54
- Older checkpoints (17): the same five texts and voices. Click one to open it.
55
-
56
- <details>
57
- <summary><b>Step 180k</b>: echo-dev WER 1.04 % · UTMOS 3.98</summary>
58
-
59
- <table>
60
- <tr><th>What the model reads</th><th>Model output, step 180k</th></tr>
61
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0180000/01-short.wav"></audio></td></tr>
62
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0180000/02-question.wav"></audio></td></tr>
63
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0180000/03-numbers.wav"></audio></td></tr>
64
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0180000/04-conversational.wav"></audio></td></tr>
65
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0180000/05-long.wav"></audio></td></tr>
66
- </table>
67
-
68
- </details>
69
-
70
- <details>
71
- <summary><b>Step 170k</b>: echo-dev WER 1.01 % · UTMOS 3.97</summary>
72
-
73
- <table>
74
- <tr><th>What the model reads</th><th>Model output, step 170k</th></tr>
75
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0170000/01-short.wav"></audio></td></tr>
76
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0170000/02-question.wav"></audio></td></tr>
77
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0170000/03-numbers.wav"></audio></td></tr>
78
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0170000/04-conversational.wav"></audio></td></tr>
79
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0170000/05-long.wav"></audio></td></tr>
80
- </table>
81
-
82
- </details>
83
-
84
- <details>
85
- <summary><b>Step 160k</b>: echo-dev WER 0.81 % · UTMOS 3.98</summary>
86
-
87
- <table>
88
- <tr><th>What the model reads</th><th>Model output, step 160k</th></tr>
89
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0160000/01-short.wav"></audio></td></tr>
90
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0160000/02-question.wav"></audio></td></tr>
91
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0160000/03-numbers.wav"></audio></td></tr>
92
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0160000/04-conversational.wav"></audio></td></tr>
93
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0160000/05-long.wav"></audio></td></tr>
94
- </table>
95
-
96
- </details>
97
-
98
- <details>
99
- <summary><b>Step 150k</b>: echo-dev WER 0.53 % · UTMOS 3.93</summary>
100
-
101
- <table>
102
- <tr><th>What the model reads</th><th>Model output, step 150k</th></tr>
103
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0150000/01-short.wav"></audio></td></tr>
104
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0150000/02-question.wav"></audio></td></tr>
105
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0150000/03-numbers.wav"></audio></td></tr>
106
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0150000/04-conversational.wav"></audio></td></tr>
107
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0150000/05-long.wav"></audio></td></tr>
108
- </table>
109
-
110
- </details>
111
-
112
- <details>
113
- <summary><b>Step 140k</b>: echo-dev WER 0.58 % · UTMOS 3.92</summary>
114
-
115
- <table>
116
- <tr><th>What the model reads</th><th>Model output, step 140k</th></tr>
117
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0140000/01-short.wav"></audio></td></tr>
118
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0140000/02-question.wav"></audio></td></tr>
119
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0140000/03-numbers.wav"></audio></td></tr>
120
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0140000/04-conversational.wav"></audio></td></tr>
121
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0140000/05-long.wav"></audio></td></tr>
122
- </table>
123
-
124
- </details>
125
-
126
- <details>
127
- <summary><b>Step 130k</b>: echo-dev WER 0.76 % · UTMOS 3.93</summary>
128
-
129
- <table>
130
- <tr><th>What the model reads</th><th>Model output, step 130k</th></tr>
131
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0130000/01-short.wav"></audio></td></tr>
132
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0130000/02-question.wav"></audio></td></tr>
133
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0130000/03-numbers.wav"></audio></td></tr>
134
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0130000/04-conversational.wav"></audio></td></tr>
135
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0130000/05-long.wav"></audio></td></tr>
136
- </table>
137
-
138
- </details>
139
-
140
- <details>
141
- <summary><b>Step 120k</b>: echo-dev WER 0.53 % · UTMOS 3.89</summary>
142
-
143
- <table>
144
- <tr><th>What the model reads</th><th>Model output, step 120k</th></tr>
145
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0120000/01-short.wav"></audio></td></tr>
146
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0120000/02-question.wav"></audio></td></tr>
147
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0120000/03-numbers.wav"></audio></td></tr>
148
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0120000/04-conversational.wav"></audio></td></tr>
149
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0120000/05-long.wav"></audio></td></tr>
150
- </table>
151
-
152
- </details>
153
-
154
- <details>
155
- <summary><b>Step 110k</b>: echo-dev WER 1.10 % · UTMOS 3.86</summary>
156
-
157
- <table>
158
- <tr><th>What the model reads</th><th>Model output, step 110k</th></tr>
159
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0110000/01-short.wav"></audio></td></tr>
160
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0110000/02-question.wav"></audio></td></tr>
161
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0110000/03-numbers.wav"></audio></td></tr>
162
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0110000/04-conversational.wav"></audio></td></tr>
163
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0110000/05-long.wav"></audio></td></tr>
164
- </table>
165
-
166
- </details>
167
-
168
- <details>
169
- <summary><b>Step 100k</b>: echo-dev WER 0.69 % · UTMOS 3.85</summary>
170
-
171
- <table>
172
- <tr><th>What the model reads</th><th>Model output, step 100k</th></tr>
173
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/01-short.wav"></audio></td></tr>
174
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/02-question.wav"></audio></td></tr>
175
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/03-numbers.wav"></audio></td></tr>
176
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/04-conversational.wav"></audio></td></tr>
177
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/05-long.wav"></audio></td></tr>
178
- </table>
179
-
180
- </details>
181
-
182
- <details>
183
- <summary><b>Step 90k</b>: echo-dev WER 0.74 % · UTMOS 3.78</summary>
184
-
185
- <table>
186
- <tr><th>What the model reads</th><th>Model output, step 90k</th></tr>
187
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0090000/01-short.wav"></audio></td></tr>
188
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0090000/02-question.wav"></audio></td></tr>
189
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0090000/03-numbers.wav"></audio></td></tr>
190
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0090000/04-conversational.wav"></audio></td></tr>
191
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0090000/05-long.wav"></audio></td></tr>
192
- </table>
193
-
194
- </details>
195
-
196
- <details>
197
- <summary><b>Step 80k</b>: echo-dev WER 0.78 % · UTMOS 3.73</summary>
198
-
199
- <table>
200
- <tr><th>What the model reads</th><th>Model output, step 80k</th></tr>
201
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0080000/01-short.wav"></audio></td></tr>
202
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0080000/02-question.wav"></audio></td></tr>
203
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0080000/03-numbers.wav"></audio></td></tr>
204
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0080000/04-conversational.wav"></audio></td></tr>
205
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0080000/05-long.wav"></audio></td></tr>
206
- </table>
207
-
208
- </details>
209
-
210
- <details>
211
- <summary><b>Step 70k</b>: echo-dev WER 0.67 % · UTMOS 3.69</summary>
212
-
213
- <table>
214
- <tr><th>What the model reads</th><th>Model output, step 70k</th></tr>
215
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0070000/01-short.wav"></audio></td></tr>
216
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0070000/02-question.wav"></audio></td></tr>
217
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0070000/03-numbers.wav"></audio></td></tr>
218
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0070000/04-conversational.wav"></audio></td></tr>
219
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0070000/05-long.wav"></audio></td></tr>
220
- </table>
221
-
222
- </details>
223
-
224
- <details>
225
- <summary><b>Step 60k</b>: echo-dev WER 0.69 % · UTMOS 3.62</summary>
226
-
227
- <table>
228
- <tr><th>What the model reads</th><th>Model output, step 60k</th></tr>
229
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0060000/01-short.wav"></audio></td></tr>
230
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0060000/02-question.wav"></audio></td></tr>
231
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0060000/03-numbers.wav"></audio></td></tr>
232
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0060000/04-conversational.wav"></audio></td></tr>
233
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0060000/05-long.wav"></audio></td></tr>
234
- </table>
235
-
236
- </details>
237
-
238
- <details>
239
- <summary><b>Step 50k</b>: echo-dev WER 0.74 % · UTMOS 3.56</summary>
240
-
241
- <table>
242
- <tr><th>What the model reads</th><th>Model output, step 50k</th></tr>
243
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0050000/01-short.wav"></audio></td></tr>
244
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0050000/02-question.wav"></audio></td></tr>
245
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0050000/03-numbers.wav"></audio></td></tr>
246
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0050000/04-conversational.wav"></audio></td></tr>
247
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0050000/05-long.wav"></audio></td></tr>
248
- </table>
249
 
250
- </details>
251
 
252
- <details>
253
- <summary><b>Step 40k</b>: echo-dev WER 0.74 % · UTMOS 3.50</summary>
 
 
 
 
 
254
 
255
  <table>
256
- <tr><th>What the model reads</th><th>Model output, step 40k</th></tr>
257
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0040000/01-short.wav"></audio></td></tr>
258
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0040000/02-question.wav"></audio></td></tr>
259
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0040000/03-numbers.wav"></audio></td></tr>
260
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0040000/04-conversational.wav"></audio></td></tr>
261
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0040000/05-long.wav"></audio></td></tr>
262
  </table>
263
 
264
- </details>
265
-
266
- <details>
267
- <summary><b>Step 30k</b>: echo-dev WER 1.06 % · UTMOS 3.43</summary>
268
 
269
- <table>
270
- <tr><th>What the model reads</th><th>Model output, step 30k</th></tr>
271
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0030000/01-short.wav"></audio></td></tr>
272
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0030000/02-question.wav"></audio></td></tr>
273
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0030000/03-numbers.wav"></audio></td></tr>
274
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0030000/04-conversational.wav"></audio></td></tr>
275
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0030000/05-long.wav"></audio></td></tr>
276
- </table>
277
-
278
- </details>
279
 
280
- <details>
281
- <summary><b>Step 20k</b>: echo-dev WER 1.20 % · UTMOS 3.33</summary>
282
 
283
- <table>
284
- <tr><th>What the model reads</th><th>Model output, step 20k</th></tr>
285
- <tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0020000/01-short.wav"></audio></td></tr>
286
- <tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0020000/02-question.wav"></audio></td></tr>
287
- <tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0020000/03-numbers.wav"></audio></td></tr>
288
- <tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0020000/04-conversational.wav"></audio></td></tr>
289
- <tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0020000/05-long.wav"></audio></td></tr>
290
- </table>
291
 
292
- </details>
 
 
 
293
 
294
- How the samples are made: one take per text, no cherry-picking, with the settings of [How to use](#how-to-use) (32 Euler steps + sway, joint CFG w 4.0, initial noise 0.9, at most -16 LUFS (peak-limited), duration by the band rule).
295
- Every output carries an inaudible [AudioSeal](https://huggingface.co/facebook/audioseal) watermark, checked before upload; the prompts
296
- are the original recordings.
297
 
298
- ## Results
299
 
300
- How well each checkpoint does on voices and sentences it never saw in training (**lower WER** is better, **higher SIM-o and
301
- UTMOS** are better):
302
 
303
  | Checkpoint | echo-dev WER ↓ | echo-dev SIM-o ↑ | echo-dev UTMOS ↑ (real speech: 4.20) | seed-dev WER ↓ | seed-dev SIM-o ↑ | seed-dev UTMOS ↑ (real speech: 3.52) | Samples | Weights |
304
  |---|---:|---:|---:|---:|---:|---:|---|---|
305
- | **step 190k** (newest) | 0.64 % | 0.790 | 3.98 | - | - | - | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0190000) |
306
- | step 180k | 1.04 % | 0.792 | 3.98 | 1.65 % | 0.261 | 3.29 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0180000) |
307
- | step 170k | 1.01 % | 0.796 | 3.97 | - | - | - | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0170000) |
308
- | step 160k | 0.81 % | 0.796 | 3.98 | 1.58 % | 0.349 | 3.66 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0160000) |
309
- | step 150k | 0.53 % | 0.794 | 3.93 | - | - | - | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0150000) |
310
- | step 140k | 0.58 % | 0.793 | 3.92 | 1.60 % | 0.404 | 3.81 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0140000) |
311
- | step 130k | 0.76 % | 0.792 | 3.93 | - | - | - | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0130000) |
312
- | step 120k | 0.53 % | 0.790 | 3.89 | 1.45 % | 0.451 | 3.86 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0120000) |
313
- | step 110k | 1.10 % | 0.789 | 3.86 | - | - | - | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0110000) |
314
- | step 100k | 0.69 % | 0.787 | 3.85 | 1.54 % | 0.499 | 3.78 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0100000) |
315
- | step 90k | 0.74 % | 0.782 | 3.78 | - | - | - | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0090000) |
316
- | step 80k | 0.78 % | 0.776 | 3.73 | 1.58 % | 0.507 | 3.68 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0080000) |
317
- | step 70k | 0.67 % | 0.766 | 3.69 | - | - | - | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0070000) |
318
- | step 60k | 0.69 % | 0.759 | 3.62 | 1.61 % | 0.537 | 3.56 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0060000) |
319
- | step 50k | 0.74 % | 0.754 | 3.56 | - | - | - | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0050000) |
320
- | step 40k | 0.74 % | 0.748 | 3.50 | 1.85 % | 0.536 | 3.48 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0040000) |
321
- | step 30k | 1.06 % | 0.741 | 3.43 | - | - | - | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0030000) |
322
- | step 20k | 1.20 % | 0.722 | 3.33 | 2.47 % | 0.506 | 3.27 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0020000) |
323
 
324
  - **WER** (word error rate): Whisper-large-v3 writes down what it hears and this is compared with the text; 2 % means about
325
  one word in fifty is wrong.
@@ -329,8 +127,10 @@ UTMOS** are better):
329
  - **echo-dev**: held-out voices of the training data's kind (synthetic EchoTTS voices, none of them trained on). **seed-dev**:
330
  real people's voices (the dev half of Seed-TTS test-en, Common Voice recordings). seed-dev is harder: the model learned
331
  only from synthetic voices.
 
 
332
  - `-`: not measured at that step (echo-dev is scored every 10k steps, seed-dev every 20k). Every number averages two takes
333
- per sentence.
334
 
335
  ## How to use
336
 
@@ -339,6 +139,7 @@ UTMOS** are better):
339
  ```bash
340
  git clone -b roadmap/en-echo https://github.com/kadirnar/dacvae-next
341
  cd dacvae-next
 
342
  pip install -e ".[codec]"
343
  ```
344
 
@@ -352,9 +153,9 @@ from huggingface_hub import snapshot_download
352
  from mytts.flow.sampler import SamplerConfig
353
  from mytts.infer import Synthesizer
354
 
355
- ckpt = "checkpoints/step_0190000" # the newest checkpoint; any row of the Results table works
356
  local = snapshot_download("VoiceHub/DACFlow-EN-10k", allow_patterns=[f"{ckpt}/*"]) # downloads only this checkpoint
357
- tts = Synthesizer.from_export(f"{local}/{ckpt}", device="cuda")
358
  sc = SamplerConfig(steps=32, cfg_mode="joint", cfg_w=4.0, noise_scale=0.9) # the settings of the samples above
359
  wav, sr = tts.synthesize(
360
  "Any English text you like.",
@@ -362,21 +163,27 @@ wav, sr = tts.synthesize(
362
  sc=sc,
363
  prompt_audio="my_voice.wav", # the voice to copy
364
  prompt_text="The exact words spoken in my_voice.wav.",
 
365
  out_lufs=-16.0,
366
  seed=0,
367
  )
368
  sf.write("output.wav", wav, sr) # 48 kHz, with the AudioSeal watermark
369
  ```
370
 
371
- **Or from the command line** (these are its default settings):
372
 
373
  ```bash
374
- hf download VoiceHub/DACFlow-EN-10k --include "checkpoints/step_0190000/*" --local-dir DACFlow-EN-10k
375
- python scripts/synthesize.py --model DACFlow-EN-10k/checkpoints/step_0190000 \
376
  --text "Any English text you like." --prompt-audio my_voice.wav \
377
  --prompt-text "The exact words spoken in my_voice.wav." --out output.wav
378
  ```
379
 
 
 
 
 
 
380
  Tips: a clean prompt with an exact transcript works best. A long text is split at sentence ends and read piece by piece.
381
  `python scripts/watermark_check.py detect output.wav` checks the watermark.
382
 
@@ -386,7 +193,8 @@ Tips: a clean prompt with an exact transcript works best. A long text is split a
386
  generates the latents of the [Semantic-DACVAE](https://huggingface.co/Aratako/Semantic-DACVAE-Japanese) audio codec (48 kHz,
387
  25 frames per second) from the text, continuing the voice prompt in context (no separate speaker encoder).
388
  - **Data**: the curated training catalog `echo_en_v1` of the provided data ([EN-10, #34](https://github.com/kadirnar/dacvae-next/issues/34) curation; see [Data](#data)).
389
- - **Recipe**: 200k steps on one RTX 5090, about 27 minutes of speech per step (40,000 latent frames); learning rate 2.5e-4 after 5k warm-up steps, held, then lowered over the last 20 % of the steps (from step 160k).
 
390
  The settings were chosen with small screening runs: the flow-matching noise schedule t_mean -0.8 / t_std 0.8 ([EN-55, #100](https://github.com/kadirnar/dacvae-next/issues/100))
391
  and cross-utterance voice prompts with p_cross 0.6 ([EN-29, #53](https://github.com/kadirnar/dacvae-next/issues/53)).
392
  - **Code and history**: the recipe [configs/train/en_full.yaml](https://github.com/kadirnar/dacvae-next/blob/roadmap/en-echo/configs/train/en_full.yaml)
 
31
  English, **10k** for its training data, about ten thousand hours of speech.
32
  The code: [kadirnar/dacvae-next](https://github.com/kadirnar/dacvae-next/tree/roadmap/en-echo), Python package `mytts`.
33
 
34
+ > **Status: training is finished. The model is step 100k** (the EMA weights saved at that step). The run was planned for 200k steps
35
+ > and stopped at step 194,271 on 2026-10-04; this page no longer changes by itself. Why step 100k: later checkpoints sound a little
36
+ > better on held-out voices of the training data's kind, but copy real people's voices much worse. Speaker similarity on real voices
37
+ > (seed-dev SIM-o) falls from 0.50 at step 100k to 0.22 at step 190k, where 51 % of the outputs score below 0.2 (5.7 % at step
38
+ > 100k), with the same settings. Step 100k, with each prompt's own bandwidth as the condition (auto), is the best balance. Details:
39
+ > [#177](https://github.com/kadirnar/dacvae-next/issues/177).
40
 
41
  **On this page:** [Listen](#listen) · [Results](#results) · [How to use](#how-to-use) · [Training](#training) · [Data](#data) · [Licence](#licence)
42
 
43
  ## Listen
44
 
45
+ The model, **step 100k**, reads five texts. Each text is spoken in a different voice, copied from the short voice prompt next
46
+ to it. The model never heard these voices in training (held-out voices of the provided data).
47
 
48
  <table>
49
+ <tr><th>What the model reads</th><th>Voice prompt (the input)</th><th>Model output (step 100k)</th></tr>
50
+ <tr><td><b>Short sentence</b><br>I left my umbrella at the office again, so I&#x27;m definitely getting soaked on the way home.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/01-short.wav"></audio><br><sub>A higher voice (~202 Hz). It says: “I keep telling myself I just need to power through, like, push a little harder, and then it&#x27;ll be fine. But it&#x27;s not getting fine. It&#x27;s getting— it&#x27;s getting worse.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/01-short.wav"></audio></td></tr>
51
+ <tr><td><b>Question</b><br>Have you ever noticed that the quietest person in the room usually has the most interesting story to tell?</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/02-question.wav"></audio><br><sub>A lower voice (~119 Hz). It says: “Look, I&#x27;m not saying it&#x27;s easy, but you always pull it together. Just take it step by step, you know?”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/02-question.wav"></audio></td></tr>
52
+ <tr><td><b>Numbers, dates, abbreviations</b><br>Dr. Patel moved my appointment to Tuesday, March 3rd, at 4:15 p.m., and the co-pay went up from $20 to $35.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/03-numbers.wav"></audio><br><sub>A higher voice (~201 Hz). It says: “Would you just— I know you mean well, but every time you mention it I feel like an idiot. It&#x27;s probably nothing.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/03-numbers.wav"></audio></td></tr>
53
+ <tr><td><b>Conversation (~10 s)</b><br>So I finally tried that new ramen place downtown, and honestly, it was worth the wait. The broth was amazing, but next time I&#x27;m definitely skipping the extra spicy option.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/04-conversational.wav"></audio><br><sub>A lower voice (~115 Hz). It says: “I&#x27;m sorry, but the subway doors closed on my bag. Again.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/04-conversational.wav"></audio></td></tr>
54
+ <tr><td><b>Long passage (~20 s)</b><br>When the storm passed, the whole neighborhood came outside to look at the damage, and although a few fences had fallen and the old oak tree had lost its biggest branch, everyone was relieved that nobody was hurt. They spent the rest of the afternoon clearing the street and sharing whatever food they had left.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/05-long.wav"></audio><br><sub>The deepest voice (~102 Hz). It says: “I mean, come on, it&#x27;s not like I woke up and decided to have the worst day ever. Things just went wrong one after another, like dominoes!”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/05-long.wav"></audio></td></tr>
55
  </table>
56
 
57
+ How the samples are made: one take per text, no cherry-picking, with the settings of [How to use](#how-to-use) (32 Euler steps + sway, joint CFG w 4.0, initial noise 0.9, at most -16 LUFS (peak-limited), duration by the band rule, each prompt's own bandwidth as the condition (auto, clamped to 9,798-16,839 Hz)).
58
+ Every output carries an inaudible [AudioSeal](https://huggingface.co/facebook/audioseal) watermark, checked before upload; the prompts
59
+ are the original recordings.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
60
 
61
+ ### How the sample voices were chosen
62
 
63
+ The voice prompt decides most of how an output sounds. In a test of this model on 45 held-out voices, 10 texts each, the choice of
64
+ prompt explained about 63 % of the differences in predicted naturalness (UTMOS), the text about 4 %. The model also copies the
65
+ recording: a prompt with a narrow frequency band (a dull, muffled recording) or with background noise gives an output that sounds
66
+ the same. So three of the five prompts above were replaced by clips that were checked first: held-out voices of the provided data
67
+ that pass the same prompt rules, recorded with a wide band and no background noise (measured on the prompts). Below, the model
68
+ (step 100k) reads the same text with the old and the new prompt (same settings and seed), with each output's band and predicted
69
+ naturalness. Details: [#177](https://github.com/kadirnar/dacvae-next/issues/177).
70
 
71
  <table>
72
+ <tr><th>What the model reads</th><th>Old voice prompt</th><th>Output with it</th><th>New voice prompt</th><th>Output with it</th></tr>
73
+ <tr><td><b>Question</b><br><sub>Replaced: a dull recording (band 8.9 kHz)</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_prompt/02-question.wav"></audio><br><sub><code>spk_0342</code> · band 8.9 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_output/02-question.wav"></audio><br><sub>band 8.2 kHz · UTMOS 3.84</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/02-question.wav"></audio><br><sub><code>spk_0327</code> · band 17.0 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/02-question.wav"></audio><br><sub>band 15.0 kHz · UTMOS 4.25</sub></td></tr>
74
+ <tr><td><b>Numbers, dates, abbreviations</b><br><sub>Replaced: background noise (signal-to-noise 38 dB)</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_prompt/03-numbers.wav"></audio><br><sub><code>spk_2669</code> · band 17.7 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_output/03-numbers.wav"></audio><br><sub>band 17.5 kHz · UTMOS 3.70</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/03-numbers.wav"></audio><br><sub><code>spk_0409</code> · band 15.5 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/03-numbers.wav"></audio><br><sub>band 15.3 kHz · UTMOS 4.38</sub></td></tr>
75
+ <tr><td><b>Long passage (~20 s)</b><br><sub>Replaced: a dull recording (band 10.9 kHz)</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_prompt/05-long.wav"></audio><br><sub><code>spk_1437</code> · band 10.9 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_output/05-long.wav"></audio><br><sub>band 8.3 kHz · UTMOS 4.15</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/05-long.wav"></audio><br><sub><code>spk_1964</code> · band 15.7 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/05-long.wav"></audio><br><sub>band 14.5 kHz · UTMOS 4.38</sub></td></tr>
 
 
76
  </table>
77
 
78
+ Band: the highest frequency with real content in the recording (the estimator of the training data). UTMOS: predicted
79
+ naturalness, 1 to 5. Both measured on the take of the same text in the sweep of [#177](https://github.com/kadirnar/dacvae-next/issues/177) (same model, settings
80
+ and seed); the players hold this page's own takes, which can differ slightly.
 
81
 
82
+ ## Results
 
 
 
 
 
 
 
 
 
83
 
84
+ ### The model: step 100k
 
85
 
86
+ How well the model does on voices and sentences it never saw in training (**lower WER** is better, **higher SIM-o and UTMOS** are
87
+ better). The settings are those of [How to use](#how-to-use): each prompt's own bandwidth as the condition (auto, clamped to 9,798-16,839 Hz).
 
 
 
 
 
 
88
 
89
+ | Test set | WER ↓ | SIM-o ↑ | UTMOS ↑ |
90
+ |---|---:|---:|---:|
91
+ | **seed-dev**: real people's voices (545 sentences, two takes each) | 1.42 % | 0.526 | 3.80 (real speech: 3.52) |
92
+ | **echo-dev v2**: held-out voices of the training data's kind (198 voices, 1,000 sentences, one take each) | 0.53 % | 0.816 | 3.99 |
93
 
94
+ 2.4 % of the seed-dev outputs score SIM-o below 0.2 (the voice is not copied).
95
+ For comparison, step 190k with the default bandwidth condition (auto was not measured there): seed-dev WER 2.06 % · SIM-o 0.223 · UTMOS 3.00, 51 % below 0.2; echo-dev v2 WER 0.46 % · SIM-o 0.820 · UTMOS 4.14.
 
96
 
97
+ ### All checkpoints of the run
98
 
99
+ Every checkpoint of the training run, scored with the default bandwidth condition; the model is step 100k:
 
100
 
101
  | Checkpoint | echo-dev WER ↓ | echo-dev SIM-o ↑ | echo-dev UTMOS ↑ (real speech: 4.20) | seed-dev WER ↓ | seed-dev SIM-o ↑ | seed-dev UTMOS ↑ (real speech: 3.52) | Samples | Weights |
102
  |---|---:|---:|---:|---:|---:|---:|---|---|
103
+ | step 190k | 0.64 % | 0.790 | 3.98 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0190000) |
104
+ | step 180k | 1.04 % | 0.792 | 3.98 | 1.65 % | 0.261 | 3.29 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0180000) |
105
+ | step 170k | 1.01 % | 0.796 | 3.97 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0170000) |
106
+ | step 160k | 0.81 % | 0.796 | 3.98 | 1.58 % | 0.349 | 3.66 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0160000) |
107
+ | step 150k | 0.53 % | 0.794 | 3.93 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0150000) |
108
+ | step 140k | 0.58 % | 0.793 | 3.92 | 1.60 % | 0.404 | 3.81 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0140000) |
109
+ | step 130k | 0.76 % | 0.792 | 3.93 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0130000) |
110
+ | step 120k | 0.53 % | 0.790 | 3.89 | 1.45 % | 0.451 | 3.86 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0120000) |
111
+ | step 110k | 1.10 % | 0.789 | 3.86 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0110000) |
112
+ | **step 100k** (the model) | 0.69 % | 0.787 | 3.85 | 1.54 % | 0.499 | 3.78 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0100000) |
113
+ | step 90k | 0.74 % | 0.782 | 3.78 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0090000) |
114
+ | step 80k | 0.78 % | 0.776 | 3.73 | 1.58 % | 0.507 | 3.68 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0080000) |
115
+ | step 70k | 0.67 % | 0.766 | 3.69 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0070000) |
116
+ | step 60k | 0.69 % | 0.759 | 3.62 | 1.61 % | 0.537 | 3.56 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0060000) |
117
+ | step 50k | 0.74 % | 0.754 | 3.56 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0050000) |
118
+ | step 40k | 0.74 % | 0.748 | 3.50 | 1.85 % | 0.536 | 3.48 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0040000) |
119
+ | step 30k | 1.06 % | 0.741 | 3.43 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0030000) |
120
+ | step 20k | 1.20 % | 0.722 | 3.33 | 2.47 % | 0.506 | 3.27 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0020000) |
121
 
122
  - **WER** (word error rate): Whisper-large-v3 writes down what it hears and this is compared with the text; 2 % means about
123
  one word in fifty is wrong.
 
127
  - **echo-dev**: held-out voices of the training data's kind (synthetic EchoTTS voices, none of them trained on). **seed-dev**:
128
  real people's voices (the dev half of Seed-TTS test-en, Common Voice recordings). seed-dev is harder: the model learned
129
  only from synthetic voices.
130
+ - **echo-dev v2**: a larger set of the same kind (198 held-out voices), the echo set the model was chosen on; the table of all
131
+ checkpoints uses the first echo-dev set.
132
  - `-`: not measured at that step (echo-dev is scored every 10k steps, seed-dev every 20k). Every number averages two takes
133
+ per sentence (echo-dev v2: one take per sentence).
134
 
135
  ## How to use
136
 
 
139
  ```bash
140
  git clone -b roadmap/en-echo https://github.com/kadirnar/dacvae-next
141
  cd dacvae-next
142
+ git checkout 029dec6 # the code that made the samples above
143
  pip install -e ".[codec]"
144
  ```
145
 
 
153
  from mytts.flow.sampler import SamplerConfig
154
  from mytts.infer import Synthesizer
155
 
156
+ ckpt = "checkpoints/step_0100000" # the model (step 100k)
157
  local = snapshot_download("VoiceHub/DACFlow-EN-10k", allow_patterns=[f"{ckpt}/*"]) # downloads only this checkpoint
158
+ tts = Synthesizer.from_export(f"{local}/{ckpt}", device="cuda", bandwidth_range=(9797.6, 16839.0))
159
  sc = SamplerConfig(steps=32, cfg_mode="joint", cfg_w=4.0, noise_scale=0.9) # the settings of the samples above
160
  wav, sr = tts.synthesize(
161
  "Any English text you like.",
 
163
  sc=sc,
164
  prompt_audio="my_voice.wav", # the voice to copy
165
  prompt_text="The exact words spoken in my_voice.wav.",
166
+ bandwidth_hz="auto", # condition on the prompt's own bandwidth (clamped to the range above)
167
  out_lufs=-16.0,
168
  seed=0,
169
  )
170
  sf.write("output.wav", wav, sr) # 48 kHz, with the AudioSeal watermark
171
  ```
172
 
173
+ **Or from the command line**:
174
 
175
  ```bash
176
+ hf download VoiceHub/DACFlow-EN-10k --include "checkpoints/step_0100000/*" --local-dir DACFlow-EN-10k
177
+ python scripts/synthesize.py --model DACFlow-EN-10k/checkpoints/step_0100000 --bandwidth auto --bandwidth-range 9797.6 16839.0 \
178
  --text "Any English text you like." --prompt-audio my_voice.wav \
179
  --prompt-text "The exact words spoken in my_voice.wav." --out output.wav
180
  ```
181
 
182
+ `auto` measures the voice prompt's frequency band and asks the model for an output of the same band (a dull prompt
183
+ gives a dull output; see [How the sample voices were chosen](#how-the-sample-voices-were-chosen)).
184
+ The range is the 5th to 95th percentile of the training data's band; it is passed explicitly because a released
185
+ `config.json` records only the top of it. `auto` needs the code of [PR #180](https://github.com/kadirnar/dacvae-next/pull/180) or later (the commit above has it).
186
+
187
  Tips: a clean prompt with an exact transcript works best. A long text is split at sentence ends and read piece by piece.
188
  `python scripts/watermark_check.py detect output.wav` checks the watermark.
189
 
 
193
  generates the latents of the [Semantic-DACVAE](https://huggingface.co/Aratako/Semantic-DACVAE-Japanese) audio codec (48 kHz,
194
  25 frames per second) from the text, continuing the voice prompt in context (no separate speaker encoder).
195
  - **Data**: the curated training catalog `echo_en_v1` of the provided data ([EN-10, #34](https://github.com/kadirnar/dacvae-next/issues/34) curation; see [Data](#data)).
196
+ - **Recipe**: planned 200k steps on one RTX 5090, about 27 minutes of speech per step (40,000 latent frames); learning rate 2.5e-4 after 5k warm-up steps, held, then lowered over the last 20 % of the steps (from step 160k).
197
+ Training was stopped at step 194,271 on 2026-10-04. The model is the EMA export of step 100k, before the learning rate was lowered.
198
  The settings were chosen with small screening runs: the flow-matching noise schedule t_mean -0.8 / t_std 0.8 ([EN-55, #100](https://github.com/kadirnar/dacvae-next/issues/100))
199
  and cross-utterance voice prompts with p_cross 0.6 ([EN-29, #53](https://github.com/kadirnar/dacvae-next/issues/53)).
200
  - **Code and history**: the recipe [configs/train/en_full.yaml](https://github.com/kadirnar/dacvae-next/blob/roadmap/en-echo/configs/train/en_full.yaml)
prompts/02-question.wav CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:337e2af17f4096f6afa7beb9a151859157c3ee8ec3cbfeca9e63aa2a61529cab
3
- size 380972
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ec5010cbb72864e62cd95a87840c6f74a750f4ac135ddcda43ce34417b63c346
3
+ size 503852
prompts/03-numbers.wav CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:c7949d82713134b3b4480cfe85cd599e5ddfc73e860a96de1b09bee04b9102fd
3
  size 634924
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4f1d94263fbe54bdcd239706933b97365dac4d34ea62f75e319da06821cdeff7
3
  size 634924
prompts/05-long.wav CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:b0047e2d525a230a7d05e70ac838100b92c425687e7cc0a3c5ab0e065c04af95
3
- size 622636
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9f6966023f8946ea8d9d915dbd3a8ac00c242217b131dcc979eed1063df2b9e0
3
+ size 708652
samples/{step_0030000 → prompt_choice/old_output}/02-question.wav RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:30b85591e417f62ebe5ee75787e647ee7af5926de93184ad6e5679b9d29efe81
3
  size 675884
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0d251f50641106b27a64844cdaae6ddf68539441021cea535a1419ff1595be34
3
  size 675884
samples/{step_0030000 → prompt_choice/old_output}/03-numbers.wav RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:78d2072e93a23915d2693d7cc7aa86885e3aaf784136826fe4a09267607cf075
3
  size 802604
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:45490e02ea584a23db51da972998d91353bd6204fc8ebab5adb609e28f6befc1
3
  size 802604
samples/{step_0020000 → prompt_choice/old_output}/05-long.wav RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:4841a5825a32e346403b0300b10ca8c52810384be26d19fa76e8bf79778a18f2
3
  size 1699244
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:20652f389b6b9a5b7424ee8dd42765e4aa33eaa6d0176812698226990fdba49e
3
  size 1699244
samples/{step_0020000 → prompt_choice/old_prompt}/02-question.wav RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:297d3d2284db20606a1c592847ec05c05c0cb6c3ac5eb7553762bd5ea64bcaec
3
- size 675884
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:337e2af17f4096f6afa7beb9a151859157c3ee8ec3cbfeca9e63aa2a61529cab
3
+ size 380972
samples/{step_0040000 → prompt_choice/old_prompt}/03-numbers.wav RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:8fb61e22ef35366165589193edf7e23b620652cd19b2ae646695c5f964f84267
3
- size 802604
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c7949d82713134b3b4480cfe85cd599e5ddfc73e860a96de1b09bee04b9102fd
3
+ size 634924
samples/{step_0020000/01-short.wav → prompt_choice/old_prompt/05-long.wav} RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:50735a6bf3454a57b37fb3ec0e543af1f0c3e2eb46b2fbd13e17619081ecf6f1
3
- size 491564
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b0047e2d525a230a7d05e70ac838100b92c425687e7cc0a3c5ab0e065c04af95
3
+ size 622636
samples/step_0020000/03-numbers.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:c125ca86b2c4234331edff603762bff5c2560c0279daa8ed0b3864acb1ba234e
3
- size 802604
 
 
 
 
samples/step_0020000/04-conversational.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:e4bd481b97e04d9663bac0b8378c1aabf60a4876f2aebe2d23e30327049af680
3
- size 1129004
 
 
 
 
samples/step_0030000/01-short.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:ef8a4af218c3446f88cc66c5a1685af66d3287a1dd4023d5a286e79cf8f5bfa0
3
- size 491564
 
 
 
 
samples/step_0030000/04-conversational.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:268121f68f855288f80d35f0dac735393ac2b7aeab9fe3c3239b1fd6e7031711
3
- size 1129004
 
 
 
 
samples/step_0030000/05-long.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:7d7d7dec305fcd3195d9f0d77c5aea0c033080778de8ef4edbe7e63880f26e84
3
- size 1699244
 
 
 
 
samples/step_0040000/01-short.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:49a9ad994685c295924654ab4bb913396c512f83fe8c61d611af31d41a109e2a
3
- size 491564
 
 
 
 
samples/step_0040000/02-question.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:191d8b96c3c567b1fe15eaa95505db8e15d06cf3f554063fbeecd855b006e53e
3
- size 675884
 
 
 
 
samples/step_0040000/04-conversational.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:7ec715e56ca0035b5a992d05e424eee61738b3f4565d84e05cde53e3b94f1070
3
- size 1129004
 
 
 
 
samples/step_0040000/05-long.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:5525bb18667843cb8fec70df78f4f5478b79faef1a753e7ec7a02a11122a6691
3
- size 1699244
 
 
 
 
samples/step_0050000/01-short.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:90977dc5598fed3c02d897870e0661c045b25375ea5594e11991564d0c6112e2
3
- size 491564
 
 
 
 
samples/step_0050000/02-question.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:c11a29b4306ff85b6a9f1834389201a3a8eaa6dc4084dcac4a2130bb28c89eb9
3
- size 675884
 
 
 
 
samples/step_0050000/03-numbers.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:cd6993d65d2ef0ee2e695bed119d56fb0a3ec261dde9b126244814d0a0cb49c1
3
- size 802604
 
 
 
 
samples/step_0050000/04-conversational.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:9f8c8b9578823d51830602bb7f3dafa2ac0639a4c647d43871d5ef7b6ed8c433
3
- size 1129004
 
 
 
 
samples/step_0050000/05-long.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:3253fea011132cb064196f174d8856a924ecae7f3dedce0c75ab91192f0f7de5
3
- size 1699244
 
 
 
 
samples/step_0060000/01-short.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:aa41e9039c23632e6e23dc63defe11549bf664c606ad8dd02a17de95f44a5954
3
- size 491564
 
 
 
 
samples/step_0060000/02-question.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:0302c8685d9e47b71c2fc891e7ae18475ce287747186d1f061a82c6b626d4ad6
3
- size 675884
 
 
 
 
samples/step_0060000/03-numbers.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:21cd0b3472609ece712c22c0c12d72134387b17ac90ef5d9084f4b8725fd25e8
3
- size 802604
 
 
 
 
samples/step_0060000/04-conversational.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:1c935be979bca4df237b0dc95c0f2188150c5e3b04324c4e702a02dbb1977d6b
3
- size 1129004
 
 
 
 
samples/step_0060000/05-long.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:4e2199443c386340cefcb37edc19c71a9ee2e5c10ffc97b4d5b6d0d994511165
3
- size 1699244
 
 
 
 
samples/step_0070000/01-short.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:abec54dbe34fda663904f8c0266a7cb02519b557a6b64cd106ee7ad0f95916b1
3
- size 491564
 
 
 
 
samples/step_0070000/02-question.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:55da8e9b91b52837b0e508a0aef7e7de32ed47add1cf5faf6e6c63a1434932da
3
- size 675884
 
 
 
 
samples/step_0070000/03-numbers.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:4c68b434f4d130f86610bfb7ed090a91595a7fe16072f7c6703ba7c032efd346
3
- size 802604
 
 
 
 
samples/step_0070000/04-conversational.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:6437479879b21c243c6360a6270f7eb662d593315c8ebe79e07ad31190889435
3
- size 1129004
 
 
 
 
samples/step_0070000/05-long.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:be36c40ab725b3638cfdd159f68c283e0a38e19636335255ee5e3a8c1e819693
3
- size 1699244
 
 
 
 
samples/step_0080000/01-short.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:220a06fd91bc5b539bfab4cc4cf6d7fa8bb207b1bd4a8136b0e9fa0b820ef99e
3
- size 491564
 
 
 
 
samples/step_0080000/02-question.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:fdf40c4150f1063d6b141721f574208fcb3b4b4a218044880e31cd1b9801b666
3
- size 675884
 
 
 
 
samples/step_0080000/03-numbers.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:f508b564c626e54bf17a85b9d1cde757a7865b63bb2c6438006cbf314422d9ce
3
- size 802604
 
 
 
 
samples/step_0080000/04-conversational.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:297a1d51596503186a06986a690e9b899e360df647bcfd970b839a4636fcfcf4
3
- size 1129004
 
 
 
 
samples/step_0080000/05-long.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:7204173c814b5924fdd90afa379c8a42e31b3c54b4068f9a0063ea3393ea7539
3
- size 1699244
 
 
 
 
samples/step_0090000/01-short.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:912169451f2095e83992c48b83187a6e6eaf20224fe36d55866ae4b339eaec16
3
- size 491564
 
 
 
 
samples/step_0090000/02-question.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:331f3aa5925718dfd401626289b343560d11f02f84fb231c9cc884b783308162
3
- size 675884
 
 
 
 
samples/step_0090000/03-numbers.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:eec195dfc8db1a8fa38e2c5e2e46f971aafdfe66eed096dce2f12d1bd7e7f87e
3
- size 802604
 
 
 
 
samples/step_0090000/04-conversational.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:9fad8387f13bb3e23258cb435f0aa0f08d8303929b683a559c83b88915fef1cb
3
- size 1129004
 
 
 
 
samples/step_0090000/05-long.wav DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:b372727c0000be7091c5c9d29ed10ca65c116edf98b81fc4099df47ca103d9ff
3
- size 1699244
 
 
 
 
samples/step_0100000/01-short.wav CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:07e3134dbff8e6f457985e581a86d768b8a87f109c0428148f97a8d7a0094f21
3
  size 491564
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:be50f96dc34d32d24739068756b6f52e88e3b0234396ee501502ba6732c548f6
3
  size 491564
samples/step_0100000/02-question.wav CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:2e391b9baba3493423e56561992e5a267a62b86d8844a7272d6e1ce43e661ab4
3
- size 675884
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1a64e3bbded645444bab1abce831686febe10ad2efc5a45d72227d2d2821fcb8
3
+ size 583724
samples/step_0100000/03-numbers.wav CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:45490e02ea584a23db51da972998d91353bd6204fc8ebab5adb609e28f6befc1
3
- size 802604
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:81225f9e0b8b42522ba01e69e5a7207495916da59b101e694a437734ef9a3440
3
+ size 841004
samples/step_0100000/04-conversational.wav CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:b01efa57885bd9667accb51a35508d42f663a3ee5d13d292f9060c8f34d6c674
3
  size 1129004
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9cdf15771b9bc38db03573d062ef083d2c6164f8189977ed34cca7c93e365c0a
3
  size 1129004
samples/step_0100000/05-long.wav CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:25713c69605a9d4727ed81d4c04d3fe6f695739f9f43a1b7eb8b86a2f0881fda
3
- size 1699244
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:252c18556b414370fdd7330dc92e07094b33533494e4cc87a3a6bf1a111d8524
3
+ size 1691564