kgrosero14 commited on
Commit
ebb4e29
·
verified ·
1 Parent(s): cc49c04

Add ChiReSSD v5 checkpoint (inference-only), config, manifest, model card

Browse files
Files changed (4) hide show
  1. README.md +153 -0
  2. config.yml +66 -0
  3. manifest.json +50 -0
  4. model.pth +3 -0
README.md CHANGED
@@ -1,3 +1,156 @@
1
  ---
 
 
 
2
  license: mit
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ library_name: chiressd
3
+ pipeline_tag: text-to-speech
4
+ language: en
5
  license: mit
6
+ base_model: yl4579/StyleTTS2
7
+ tags:
8
+ - styletts2
9
+ - speech-reconstruction
10
+ - disordered-speech
11
+ - child-speech
12
+ - clinical
13
  ---
14
+
15
+ # ChiReSSD
16
+
17
+ Speaker-preserving reconstruction of pediatric disordered speech.
18
+
19
+ Given a target transcript and a short reference recording from a child with a speech sound
20
+ disorder (SSD), ChiReSSD synthesizes the target utterance with canonical pronunciation while
21
+ keeping the child's voice and prosody. Pronunciation enters through the text pathway; identity
22
+ and prosody come from the style pathway. That separation is the point: ordinary style-preserving
23
+ TTS treats disordered articulation as part of the speaker's style and so reproduces the
24
+ mispronunciation it was meant to correct.
25
+
26
+ - Code: <https://github.com/Lab-MSP/ChiReSSD>
27
+ - Paper: *Generative Reconstruction of Pediatric Disordered Speech for Automated Clinical
28
+ Evaluation*
29
+
30
+ ## Model description
31
+
32
+ StyleTTS2 fine-tuned on child disordered speech. Two 128-dimensional style vectors are extracted
33
+ from the reference recording — acoustic (timbre) and prosodic — and interpolated with a style
34
+ sampled from the adapted diffusion prior. `alpha` weights the acoustic side and `beta` the
35
+ prosodic; 1.0 is fully diffusion-sampled, 0.0 fully reference-driven.
36
+
37
+ It works on unseen speakers from a reference as short as a few seconds. No per-child model is
38
+ trained, and no paired typical/atypical recordings are required.
39
+
40
+ ## Base model and license chain
41
+
42
+ Fine-tuned from the StyleTTS2 LibriTTS second-stage checkpoint by
43
+ [yl4579](https://github.com/yl4579/StyleTTS2) (MIT). **That base checkpoint is not redistributed
44
+ here** — obtain it from the upstream release. The frozen helper models (ASR text aligner, JDCNet
45
+ pitch extractor, PL-BERT) likewise ship inside the upstream repository and are not redistributed.
46
+
47
+ ChiReSSD's own modification to StyleTTS2 is two lines, published as patch files in the code
48
+ repository rather than as a fork.
49
+
50
+ ## Intended use
51
+
52
+ Research on speech reconstruction and on automated clinical evaluation of pediatric speech sound
53
+ disorders.
54
+
55
+ ## Out of scope
56
+
57
+ - **Not a medical device.** No diagnostic or treatment decision should rest on its output.
58
+ - Not a replacement for assessment by a licensed speech-language pathologist.
59
+ - Not validated for populations outside the training corpus — it was adapted to children aged
60
+ 5–12 with a Central Scottish accent, and the phonemizer default (`en-gb`) reflects that.
61
+ - Not a general-purpose voice cloning system.
62
+
63
+ ## Training data
64
+
65
+ Child speech from UltraSuite/CLP-derived recordings, obtained under their own data use
66
+ agreements. **The corpus is not released here and is not redistributable**: it is identifiable
67
+ child clinical speech. See `DATA.md` in the code repository for how to obtain comparable data;
68
+ <https://huggingface.co/datasets/changelinglab/ultrasuite-benchmark> is a suitable starting point
69
+ for fine-tuning your own model.
70
+
71
+ ## Training configuration
72
+
73
+ | | |
74
+ |---|---|
75
+ | Epochs | 4 |
76
+ | Batch size | 4 |
77
+ | Max length | 600 frames |
78
+ | Learning rate | 1e-5 (`lr`, `bert_lr`, `ft_lr`) |
79
+ | `lambda_F0` | 20 (upstream: 1) |
80
+ | `lambda_mel` | 5 |
81
+ | Style diffusion from epoch | 2 |
82
+ | Joint SLM-adversarial from epoch | 3 |
83
+ | Decoder | HiFi-GAN, multispeaker |
84
+ | Sample rate | 24 kHz |
85
+ | LR schedule | OneCycleLR, `pct_start=0.1` (upstream: 0) |
86
+
87
+ The heavy pitch weighting is deliberate: the pitch extractor was pretrained on adult voices, and
88
+ children's F0 is both higher and more variable, so it needs the strongest adaptation of any
89
+ component. Conversely, only four epochs — longer schedules start fitting the disordered
90
+ articulation itself.
91
+
92
+ ## Inference
93
+
94
+ | Preset | alpha | beta | steps | Purpose |
95
+ |---|---|---|---|---|
96
+ | `star_v5` | 0.8 | 0.5 | 10 | Released operating point |
97
+ | `star_v5_alt` | 1.0 | 0.4 | 10 | Variant |
98
+ | `star_oneshot_base` | 1.0 | 0.5 | 5 | One-shot baseline (no fine-tuning) |
99
+ | `star_adult_tts` | 0.3 | 0.7 | 5 | Standard adult-TTS baseline |
100
+ | `torgo_v5` | 1.0 | 0.5 | 15 | Adult dysarthric speech |
101
+
102
+ `alpha` is high on purpose. A low `alpha` leans on the reference acoustics, which is exactly
103
+ where the disordered articulation lives.
104
+
105
+ **Synthesis is stochastic.** The initial style latent is zeros rather than a Gaussian draw, but
106
+ the ADPM2 sampler is ancestral and adds fresh noise at every step, so repeated calls differ in
107
+ waveform and in duration. Pass a `seed` for reproducible output.
108
+
109
+ ## Evaluation
110
+
111
+ See the paper. **This model repository does not reproduce the paper's reported numbers**, and the
112
+ code repository does not either: the evaluation corpora are not redistributable, so the numbers
113
+ cannot be independently recomputed from what is published here.
114
+
115
+ ## Ethical considerations
116
+
117
+ This model is trained on identifiable clinical recordings of children and is designed to
118
+ reproduce a child's vocal identity accurately. High speaker similarity is the method's goal and
119
+ therefore also its misuse vector: the same property that makes feedback usable in a child's own
120
+ voice makes the model a voice-cloning tool for a minor. Use it only with appropriate consent and
121
+ ethical oversight.
122
+
123
+ ## Usage
124
+
125
+ ```python
126
+ from chiressd.model import ChiReSSD
127
+ from chiressd.config import load_preset
128
+
129
+ model = ChiReSSD.from_pretrained() # downloads this checkpoint
130
+ style = model.compute_style("child_reference.wav")
131
+ wav = model.synthesize(
132
+ "butterfly butterfly butterfly", style, seed=1234, **load_preset("star_v5").as_kwargs()
133
+ )
134
+ ```
135
+
136
+ Run `chiressd-setup` first: it clones and patches the upstream StyleTTS2 checkout that supplies
137
+ the architecture and frozen helper models.
138
+
139
+ ## Checkpoint provenance
140
+
141
+ Derived from training checkpoint `epoch_2nd_00003.pth` of the v5 run by keeping `state['net']`
142
+ only, removing the `module.` prefix that DataParallel added to ten of the thirteen submodules,
143
+ and making every tensor detached, CPU-resident and contiguous. Precision is unchanged (float32;
144
+ no fp16 cast, which would alter outputs).
145
+
146
+ - 2258 MB → 767 MB; the 1491 MB removed is optimizer state, which inference never reads
147
+ - `sha256` of `model.pth`: `6c5faf5b4967b4d26cb4c58d4517dcc477729a579039316741bd88b9bd0f16e0`
148
+ - Verified bit-exact against the original: identical 256-d style vector and all 105 000 waveform
149
+ samples under a fixed seed (`scripts/verify_checkpoint.py`)
150
+
151
+ Full details in `manifest.json`.
152
+
153
+ ## Citation
154
+
155
+ Rosero, Yeo, Mortensen, Van't Slot, Hallac, and Busso. *Generative Reconstruction of Pediatric
156
+ Disordered Speech for Automated Clinical Evaluation.*
config.yml ADDED
@@ -0,0 +1,66 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ASR_config: Utils/ASR/config.yml
2
+ ASR_path: Utils/ASR/epoch_00080.pth
3
+ F0_path: Utils/JDC/bst.t7
4
+ PLBERT_dir: Utils/PLBERT/
5
+ model_params:
6
+ decoder:
7
+ resblock_dilation_sizes:
8
+ - - 1
9
+ - 3
10
+ - 5
11
+ - - 1
12
+ - 3
13
+ - 5
14
+ - - 1
15
+ - 3
16
+ - 5
17
+ resblock_kernel_sizes:
18
+ - 3
19
+ - 7
20
+ - 11
21
+ type: hifigan
22
+ upsample_initial_channel: 512
23
+ upsample_kernel_sizes:
24
+ - 20
25
+ - 10
26
+ - 6
27
+ - 4
28
+ upsample_rates:
29
+ - 10
30
+ - 5
31
+ - 3
32
+ - 2
33
+ diffusion:
34
+ dist:
35
+ estimate_sigma_data: true
36
+ mean: -3.0
37
+ sigma_data: 0.2740232650313019
38
+ std: 1.0
39
+ embedding_mask_proba: 0.1
40
+ transformer:
41
+ head_features: 64
42
+ multiplier: 2
43
+ num_heads: 8
44
+ num_layers: 3
45
+ dim_in: 64
46
+ dropout: 0.2
47
+ hidden_dim: 512
48
+ max_conv_dim: 512
49
+ max_dur: 50
50
+ multispeaker: true
51
+ n_layer: 3
52
+ n_mels: 80
53
+ n_token: 178
54
+ slm:
55
+ hidden: 768
56
+ initial_channel: 64
57
+ model: microsoft/wavlm-base-plus
58
+ nlayers: 13
59
+ sr: 16000
60
+ style_dim: 128
61
+ preprocess_params:
62
+ spect_params:
63
+ hop_length: 300
64
+ n_fft: 2048
65
+ win_length: 1200
66
+ sr: 24000
manifest.json ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "sha256": "6c5faf5b4967b4d26cb4c58d4517dcc477729a579039316741bd88b9bd0f16e0",
3
+ "filename": "model.pth",
4
+ "created": "2026-09-14T00:48:03+00:00",
5
+ "chiressd_commit": "e1b8bbd68255229c55a0797c4a8f2cb0f02fdb84",
6
+ "upstream_styletts2_commit": "5cedc71c333f8d8b8551ca59378bdcc7af4c9529",
7
+ "source_checkpoint": "epoch_2nd_00003.pth",
8
+ "source_config": "config_ft_F0_v5.yml",
9
+ "strip_recipe": "keep state['net'] only; drop optimizer and scalar bookkeeping; remove the leading 'module.' that DataParallel added; detach/cpu/contiguous every tensor; keep float32 (no fp16 cast, which would change outputs); re-save with _use_new_zipfile_serialization=True",
10
+ "modules": [
11
+ "bert",
12
+ "bert_encoder",
13
+ "decoder",
14
+ "diffusion",
15
+ "mpd",
16
+ "msd",
17
+ "pitch_extractor",
18
+ "predictor",
19
+ "predictor_encoder",
20
+ "style_encoder",
21
+ "text_aligner",
22
+ "text_encoder",
23
+ "wd"
24
+ ],
25
+ "dataparallel_prefixed_modules": [
26
+ "bert",
27
+ "bert_encoder",
28
+ "decoder",
29
+ "diffusion",
30
+ "pitch_extractor",
31
+ "predictor",
32
+ "predictor_encoder",
33
+ "style_encoder",
34
+ "text_aligner",
35
+ "text_encoder"
36
+ ],
37
+ "dropped_keys": [
38
+ "epoch",
39
+ "iters",
40
+ "optimizer",
41
+ "val_loss"
42
+ ],
43
+ "source_metadata": {
44
+ "epoch": 3,
45
+ "iters": 1700,
46
+ "val_loss": 0.3068065345287323
47
+ },
48
+ "source_bytes": 2257595196,
49
+ "released_bytes": 766549479
50
+ }
model.pth ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6c5faf5b4967b4d26cb4c58d4517dcc477729a579039316741bd88b9bd0f16e0
3
+ size 766549479