File size: 12,211 Bytes
63e7bc2
 
 
 
 
 
 
 
 
 
 
 
6fbb52c
63e7bc2
4e9e0e4
71084e9
9f63115
 
4e9e0e4
63e7bc2
 
71084e9
 
4e9e0e4
63e7bc2
4e9e0e4
63e7bc2
 
 
 
 
6fbb52c
63e7bc2
4e9e0e4
 
63e7bc2
4e9e0e4
71084e9
 
6fbb52c
 
63e7bc2
4e9e0e4
 
 
63e7bc2
6fbb52c
71084e9
6fbb52c
71084e9
63e7bc2
6fbb52c
71084e9
 
 
6fbb52c
63e7bc2
71084e9
6fbb52c
63e7bc2
4c2c5bc
71084e9
 
4e9e0e4
 
63e7bc2
6fbb52c
 
4e11051
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4c2c5bc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
63e7bc2
4e9e0e4
 
63e7bc2
4e9e0e4
63e7bc2
 
 
 
71084e9
 
6fbb52c
 
71084e9
 
 
 
6fbb52c
63e7bc2
71084e9
6fbb52c
63e7bc2
71084e9
 
6fbb52c
 
 
 
63e7bc2
 
71084e9
 
 
78b3cbe
71084e9
78b3cbe
71084e9
 
 
78b3cbe
71084e9
5fb8b3f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
78b3cbe
 
8909649
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
63e7bc2
4e9e0e4
63e7bc2
 
aff2f95
aaf1a08
4e11051
 
 
 
aaf1a08
 
 
 
 
 
9f63115
63e7bc2
9f63115
 
 
 
 
 
 
 
 
 
8909649
 
 
 
 
 
 
 
 
 
 
5fb8b3f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
71084e9
8909649
 
 
 
 
 
63e7bc2
4e9e0e4
63e7bc2
 
 
 
9f63115
 
 
 
63e7bc2
8909649
 
5fb8b3f
63e7bc2
 
4e9e0e4
63e7bc2
4e9e0e4
63e7bc2
 
 
 
71084e9
 
63e7bc2
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
---
language:
- zh
- en
tags:
- text-to-speech
- speech-synthesis
- code-switching
library_name: bluemagpie
pipeline_tag: text-to-speech
license: other
---

# BlueMagpie-TTS

BlueMagpie-TTS checkpoint for research and evaluation. The files in this Hub
repository are the selected two-mode recovery release. Experimental branches
are not promoted when they regress the production speaker-reference path.

This repository contains the inference artifact only: model weights, AudioVAE
weights, tokenizer files, config, and usage documentation. It does not include
optimizer state, scheduler state, training logs, private speaker registries,
local configs, or training-data metadata.

## Intended Use

- Mandarin and mixed Mandarin/English text-to-speech evaluation.
- Experiments with two mutually exclusive generation modes:
  centroid-only and speaker-reference-audio-only.
- Do not redistribute the checkpoint or generated speech unless rights and
  consent are cleared for the intended use.

## Install

```bash
git clone https://github.com/voidful/BlueMagpie-TTS
cd BlueMagpie-TTS
pip install -e ".[speaker]"
pip install soundfile
```

## Quick Start

```python
from huggingface_hub import snapshot_download
import soundfile as sf

from bluemagpie import BlueMagpieModel, speaker_embedding_from_wav

model_dir = snapshot_download("OpenFormosa/BlueMagpie-TTS")
model = BlueMagpieModel.from_local(model_dir, training=False, device="cuda")

reference_embedding = speaker_embedding_from_wav(
    "reference_speaker.wav", window_sec=3.0, hop_sec=1.5
)
audio = model.generate(
    target_text="這是 AI TTS code switching 測試。",
    speaker_centroid=reference_embedding,
    cfg_value=2.0,
    inference_timesteps=10,
    retry_badcase=False,
    stop_threshold=0.65,
    stop_consecutive=2,
)

sf.write("sample.wav", audio.detach().cpu().numpy(), model.sample_rate)
```

## Interactive Demo Endpoint Profile

The hosted [BlueMagpie-TTS Demo](https://huggingface.co/spaces/voidful/BlueMagpie-TTS-Demo)
uses the same checkpoint weights with an inference-only endpoint guard. It
starts with a 0.50 stop threshold, lowers the threshold only after 75% of the
native-rate duration estimate, reaches 0.05 at 95%, and accepts one stop hit.
Generation is still capped at the native-rate estimate plus one latent step.
Playback pace is corrected only after generation has stopped, so pace control
does not hold the model open and create extra speech.

This interactive policy is intentionally separate from the reproducible
offline evaluation default of `stop_threshold=0.65` and
`stop_consecutive=2`. Neither policy uses retry selection or best-of-N
reranking in its standard path.

## Hosted Speaker B Embedding

The hosted Space keeps Speaker A as the default and loads its fixed Speaker B
voice from `checkpoints/speaker_b_embedding.pt`. The artifact is a 192-dimensional,
L2-normalized embedding extracted from the clean Speaker B reference set. Its
payload uses the `bluemagpie-speaker-embedding-v1` format and includes
`speaker_id="female_voice"`, `generation_seed=56789`, and the Space runtime
profile `speaker_projector_scale=1.125`.

The deployed Space pins model revision
`6f7cab914a1e27c56b504ec663c0144dc25cc0a3` and validates the embedding SHA-256
`e9556e14723c140985a104c1659d1ff8a5078d2fa28ce2fb756f04906641a8a7` at
startup. Each text chunk is generated exactly once with `retry_badcase=False`;
the request path does not run ASR, speaker verification, candidate reranking,
or output-validation retries.

```python
import os
import torch

payload = torch.load(
    os.path.join(model_dir, "checkpoints", "speaker_b_embedding.pt"),
    map_location="cpu",
    weights_only=True,
)
speaker_b = torch.nn.functional.normalize(payload["embedding"].float(), dim=0)

audio = model.generate(
    target_text="今天的會議改到下午三點。",
    speaker_centroid=speaker_b,
    cfg_value=2.0,
    inference_timesteps=10,
    retry_badcase=False,
    stop_threshold=0.65,
    stop_consecutive=2,
)
```

`generation_seed` and `speaker_projector_scale` are metadata for the hosted
Space's fixed Speaker B runtime profile; `speaker_projector_scale` is not a
public `generate()` keyword argument.

## Conditioning Modes

```python
# centroid-only
audio = model.generate(
    target_text="今天的會議改到下午三點。",
    speaker_centroid=centroid_tensor,
    cfg_value=2.0,
    inference_timesteps=10,
    stop_threshold=0.65,
    stop_consecutive=2,
)

# speaker-reference-audio-only: extract once; no transcript required
reference_embedding = speaker_embedding_from_wav(
    "reference_speaker.wav", window_sec=3.0, hop_sec=1.5
)
audio = model.generate(
    target_text="今天的會議改到下午三點。",
    speaker_centroid=reference_embedding,
    cfg_value=2.0,
    inference_timesteps=10,
    stop_threshold=0.65,
    stop_consecutive=2,
)

```

Only use reference audio or speaker centroids from speakers you have permission
to synthesize. Reference clips should be at least 3 seconds. The
speaker-reference-audio-only mode uses one window-averaged speaker embedding
for the full request and does not require a reference transcript. Raw reference
tokens remain a separately evaluated research path.

## Long Text

Use the CLI for long-form synthesis. It preserves punctuation, reuses one
speaker embedding for every chunk, applies punctuation-aware pauses, and can
set a minimum duration from the target speaking rate:

```bash
python scripts/generate_tts.py \
  --checkpoint /path/to/model-snapshot \
  --mode speaker-reference-audio-only \
  --speaker-reference-wav /path/to/rights-cleared-reference.wav \
  --reference-audio-conditioning speaker-embedding \
  --text-file long_text.txt \
  --chunk-chars 80 \
  --min-chunk-chars 12 \
  --target-chars-per-sec 4.0 \
  --stop-threshold 0.65 \
  --stop-consecutive 2 \
  --crossfade-ms 80 \
  --chunk-rms-match-db 4 \
  --chunk-edge-fade-ms 80 \
  --continuation-context-sec 0 \
  --no-retry-badcase \
  --out long.wav
```

## Quality-First Short and Medium Text

The optional offline quality path generates multiple candidates and uses ASR
plus speaker verification to select one. It is slower than ordinary generation
and does not change the checkpoint weights:

```bash
python scripts/generate_tts_quality.py \
  --checkpoint /path/to/model-snapshot \
  --mode speaker-reference-audio-only \
  --speaker-reference-wav /path/to/rights-cleared-reference.wav \
  --text "這是離線品質優先的語音合成測試。" \
  --candidates 10 \
  --asr-backend whisper \
  --asr-model /path/to/compatible-asr-checkpoint \
  --candidate-speaker-weight 0.05 \
  --candidate-boundary-speaker-weight 0.1 \
  --candidate-max-boundary-speaker-drop 0.03 \
  --target-chars-per-sec 4.2 \
  --out quality.wav
```

The candidate count and score weights are measured project heuristics, not
universal TTS defaults. The command fails closed when no candidate passes the
speaker-boundary gate. `--allow-boundary-fallback` is available for research,
but is outside the measured contract. This quality policy is promoted only for
speaker-reference short/medium synthesis; centroid and strict long-form gates
did not pass.

## Evaluation

Numbers below are from an internal held-out evaluation set. The eval set and
training data are intentionally not described in this model card.

The residual-speaker pilot improved natural and long-form aggregate
metrics, but every tested gate scale failed at least one stress endpoint,
intelligibility, or speaker-boundary criterion. Its weights were not promoted;
the selected model fingerprint remains unchanged.

A later frozen-parent endpoint-delta pilot changed only its isolated stop
residual. All 24 outputs at every evaluated checkpoint remained byte-identical
to the matched parent, so no endpoint-delta checkpoint was promoted. The
follow-up stop-margin diagnosis is retained as research evidence and does not
change the model artifact or production inference contract.

| Mode | Evaluation profile | CER | WER |
| --- | --- | ---: | ---: |
| centroid-only | 30 texts x 5 seeds, controlled pace | 15.37% | 41.48% |
| speaker-reference-embedding | 30 texts x 5 seeds, controlled pace | 13.39% | 33.05% |

The controlled profile uses 4.2 target chars/sec, stop threshold 0.65, two-hit
hysteresis, and no retry selection. Its median/P95 pace is 4.04/4.17 chars/sec.
The stricter natural-stop profile has lower CER but fails the pace gate: P95
reaches 6.25 chars/sec and early-stop classification exceeds 50%. Controlled
pace is therefore part of the current production contract, not an optional
benchmark optimization.

The opt-in short/medium quality policy was measured on 30 texts with ten
candidates per text. The full candidate pool reached 9.35% CER / 9.51% WER;
the selected outputs reached 1.14% / 1.86%, with zero early stops, overruns,
final truncations, or speaker-boundary fallbacks. Across-text speaker
similarity was 0.511, reference-similarity P10 was 0.444, and median/P95 pace
was 4.04/4.17 chars/sec. A second ASR that did not participate in selection
measured 6.83% CER / 6.78% WER; on matched target recordings it measured
9.20% / 8.76%. The same single code-switch item was the only outlier in both.
These are test-time selection results and must not be described as a trained
checkpoint improvement.

The selected long-form contract was repeated in four independent processes,
each with three multi-sentence texts and five request seeds. It uses
80-character chunks, 4.0 target chars/sec, no retry selection, one reference
embedding per request, and no generated-audio continuation context.

| Metric | Four-run range |
| --- | ---: |
| CER | 9.40%-11.90% |
| WER | 11.13%-13.14% |
| Median / P95 chars per second | 3.87-3.89 / 3.99 |
| Duration CV across seeds | 1.31%-2.38% |
| Across-text speaker similarity | 0.759-0.771 |
| Reference similarity P10 | 0.547-0.557 |
| First-to-last sentence speaker drop P95 | 0.017-0.030 |
| Sentence-final truncation / ASR outliers | 0 / 0 |

A separate 12-text, one-seed stress slice reached 14.64%-15.19% CER with no
final truncation; one ASCII-and-numeric-heavy item was an ASR outlier. These
metrics use a corrected text-normalization contract and are not directly
comparable to older long-form reports. Feeding generated audio into the next
chunk consistently worsened CER, pace variance, and speaker drift, so
`--continuation-context-sec 0` remains the production default. Raw reference
tokens are not represented by these production metrics.

A separately frozen 12-text long-form holdout reached 4.00% selected CER and
4.91% WER under the selection ASR, but one numeric/ASCII-heavy request had no
candidate below the 0.03 speaker-boundary gate. Independent-ASR CER was 15.25%
with one cross-chunk outlier. The quality-reranked long-form path is therefore
not promoted; use the standard measured 80-character/no-context command above.

## Limitations

- Metrics are not a public benchmark and should be used only for internal model
  selection.
- Speaker similarity depends on the quality and rights-cleared status of the
  supplied reference audio or centroid.
- The windowed speaker-embedding backend is the recommended production path.
- Free-running endpoint behavior is not fully calibrated. Use the documented
  target pace, stop hysteresis, and chunking contract; inspect important output.
- Raw reference tokens remain a separately gated research backend.
- Very long passages should be chunked to avoid stop-token and prosody drift.
- Quality reranking is currently promoted only for speaker-reference
  short/medium synthesis; it is not the default centroid or long-form path.
- Generated-audio continuation context is experimental and disabled by default.
- Generated speech may be incorrect; do not use it as a real-world notification
  without human review.

## Files

- `pytorch_model.bin`: BlueMagpie model weights.
- `audiovae.pth`: AudioVAE weights.
- `config.json`: BlueMagpie architecture/runtime config.
- `tokenizer.json`, `tokenizer_config.json`: tokenizer files.
- `checkpoints/*speaker_centroids.pt`: small centroid demo artifacts retained
  on the Hub for centroid-only examples.
- `USAGE.md`: expanded usage guide.