|
Download README.md from zeechimp/whisper-decoder: direct link, hf CLI and curl.
- Browser
- Download file 10.3 kB
-
https://huggingface.co/zeechimp/whisper-decoder/resolve/main/README.md
- Command line
-
hf download hf://zeechimp/whisper-decoder/README.md
-
curl -L -o README.md https://huggingface.co/zeechimp/whisper-decoder/resolve/main/README.md
10.3 kB
| license: apache-2.0 | |
| library_name: whisper-decoder | |
| tags: | |
| - speech | |
| - whisper | |
| - formants | |
| - mfcc | |
| - keyword-spotting | |
| - numpy | |
| - cpu | |
| pipeline_tag: audio-classification | |
| language: | |
| - en | |
| # whisper-decoder | |
| **A classifier trained only on whispered speech identifies the | |
| same words when they are voiced at 83.2% accuracy — 8.3× chance.** | |
| Whispering removes the pitch source. The vocal folds stop vibrating | |
| and the only sound that leaves the mouth is turbulent noise shaped | |
| by the vocal tract. The formants survive. The word survives. Only | |
| the voicing is gone. | |
| A classifier trained on whispered digits and evaluated on voiced | |
| digits recovers the word from the formants alone. This is what the | |
| experiment measures: how much of a word's identity lives in the | |
| formant structure, independent of the pitch that normal speech | |
| provides. | |
| ## The claim in one sentence | |
| Formant structure is sufficient to identify a closed-vocabulary | |
| word even when the voicing source that produced the word is | |
| completely different. | |
| ## What it produces | |
| A closed-vocabulary classifier over ten English digits | |
| (`zero` through `nine`). Given a 1.2-second utterance, it returns | |
| the predicted word and a probability distribution over the | |
| vocabulary. | |
| Internally it computes 49 features: | |
| | feature family | count | purpose | | |
| |---|---|---| | |
| | MFCC mean + std | 26 | spectral envelope, formants | | |
| | low-band mel mean + std | 20 | F1 and F2 explicitly, 200–2500 Hz | | |
| | harmonicity mean + std | 2 | voicing detector (near 0 for whisper) | | |
| | trimmed duration | 1 | word length after silence removal | | |
| ## Install | |
| ```bash | |
| pip install numpy | |
| ``` | |
| No other dependencies for the classifier. `matplotlib` is optional | |
| for the figures. No downloads. No model weights. The classifier is | |
| 5,610 parameters and trains from scratch in 1.6 seconds. | |
| ## Usage | |
| ### Run the demo | |
| ```bash | |
| python whisper_decoder.py | |
| ``` | |
| Prints self-test results, generates training and test data, | |
| trains the classifier, and evaluates on three corpora. Writes one | |
| figure to `whisper_figures/`. | |
| ### Programmatic | |
| ```python | |
| from whisper_decoder import ( | |
| MLP, train_mlp, extract_features, | |
| synth_word, generate_corpus, FEATURE_DIM, | |
| ) | |
| # Generate a training corpus (whisper mode, 200 samples per word) | |
| train = generate_corpus('whisper', 200, 40, seed=0, | |
| reverb_train_frac=0.25) | |
| # Standardize | |
| mu = train.X_train.mean(axis=0) | |
| sigma = train.X_train.std(axis=0) + 1e-9 | |
| X_tr = (train.X_train - mu) / sigma | |
| # Train | |
| model = MLP(FEATURE_DIM, 64, 32, 10, seed=0) | |
| train_mlp(model, X_tr, train.y_train, epochs=200, batch=64) | |
| # Classify a new utterance | |
| sig = synth_word('seven', 'whisper', seed=42) | |
| feats = extract_features(sig) | |
| probs = model.predict_proba(((feats - mu) / sigma)[None, :])[0] | |
| predicted = train.word_index_inv[int(probs.argmax())] | |
| print(predicted, probs.max()) | |
| ``` | |
| ### Self-test | |
| ```bash | |
| python whisper_decoder.py --self-test | |
| ``` | |
| 11 checks covering the synthesizer, harmonicity, feature | |
| extraction, and MFCC reproducibility. | |
| ## Benchmarks | |
| Trained on 2,000 synthetic whispered utterances (200 per word, | |
| 25% with reverb augmentation). Evaluated on 400 held-out | |
| utterances. | |
| ### Primary results | |
| | corpus | accuracy | chance | ratio | | |
| |---|---|---|---| | |
| | **voiced speech (domain shift)** | **83.2%** | 10.0% | **8.3×** | | |
| | whisper (in-distribution) | 100.0% | 10.0% | 10.0× | | |
| | reverberant whisper | 100.0% | 10.0% | 10.0× | | |
| **The voiced result is the finding.** The classifier was never | |
| trained on voiced speech. It has never seen a pulse train. The | |
| same 49 features — MFCC, low-band mel, harmonicity, duration — | |
| still identify the word. | |
| ### What the confusion matrix looks like | |
| Perfect diagonal on the whispered test set. Every digit is | |
| classified correctly, 40 out of 40. | |
| This is a synthetic dataset, not real whispered speech. The | |
| 100% is an upper bound on the synthetic task, not a prediction | |
| of real-world accuracy. | |
| ### The reverb row needs a caveat | |
| The reverb augmentation during training and the reverb applied | |
| during testing both use `rt60_s=0.35`. The training and test | |
| distributions are the same. The 100% measures that the classifier | |
| handles reverb *from the family it was trained on*. It does not | |
| measure out-of-distribution robustness. | |
| An example script (`examples/reverb_ood.py`) tests the honest | |
| version: train on `rt60=0.35`, evaluate on `rt60=0.7`. Run it to | |
| get the real number. | |
| ## How it works | |
| ### Synthesis | |
| The source-filter model of speech production, in its simplest | |
| form. Each word is a sequence of phonemes drawn from a small | |
| inventory. Each phoneme has a formant target: a list of | |
| `(frequency_hz, bandwidth_hz, amplitude)` triples. The source is | |
| either white noise (whisper) or a pulse train at F0 (voiced). | |
| The formant bank is a sum of Gaussian resonances applied to the | |
| source spectrum. | |
| Same phonemes, same formants, different source. The only thing | |
| that changes between a whispered utterance and a voiced utterance | |
| of the same word is the excitation. | |
| ### Features | |
| - **MFCC**: 13 coefficients, mean and std over frames | |
| - **Low-band mel**: 10 bands over 200–2500 Hz, mean and std | |
| - **Harmonicity**: per-frame peak of the normalized autocorrelation | |
| in the pitch-lag range (80–300 Hz). Near zero for whisper, 0.5–0.9 | |
| for voiced. | |
| - **Trimmed duration**: length of the non-silent portion | |
| The low-band range is deliberately extended to 2500 Hz because F2 | |
| for front vowels (`i`, `e`) sits at 2000–2500 Hz. A narrower band | |
| loses `three`, `zero`, and `five`. | |
| ### Why it works on voiced speech | |
| The formants in the synthesized voiced speech occupy the same | |
| frequencies as the formants in the synthesized whispered speech. | |
| The classifier learned to weight MFCC and low-band mel, which are | |
| formant-driven. Harmonicity is a near-zero feature in the | |
| whisper training set and stays near zero in the voiced test set | |
| for unvoiced consonants but rises above 0.5 for vowels. The | |
| classifier is not thrown off because the other features carry | |
| the word identity. | |
| ## Applications | |
| - **Assistive speech interfaces for whispered input.** A user who | |
| has lost their voice to illness, surgery, or fatigue can whisper | |
| and still be understood by a device. | |
| - **Security and anonymity.** Whispering is used when the speaker | |
| does not want to be heard. A decoder that works on whispers is | |
| a decoder that works on material recorded from a distance with | |
| a directional microphone. | |
| - **Speech reconstruction from damaged recordings.** A recording | |
| that has been aggressively denoised, high-pass filtered, or | |
| compressed may still carry enough formant structure. The | |
| classifier works on the formants, not on the pitch that is the | |
| first thing to be lost. | |
| - **Cross-modal keyword spotting.** A wake-word detector that | |
| triggers whether the wake word is whispered or spoken. Most | |
| wake-word systems are trained on voiced speech and fail on | |
| whispers. | |
| ## When to use it | |
| - When the vocabulary is small and fixed. | |
| - When utterances are roughly the same length. | |
| - When the source-filter assumption holds — a single speaker, | |
| moderate background noise, no heavy reverb. | |
| ## When *not* to use it | |
| - **Open-vocabulary transcription.** This is not a speech | |
| recognizer. It is a 10-way classifier. | |
| - **Real whispered speech at scale.** The training data is | |
| synthetic. Real whispers have coarticulation, breath noise, | |
| and speaker-specific formant variation that the simulator | |
| does not model. | |
| - **Speaker-independent deployment.** The formant targets are | |
| averaged across speakers. An individual speaker's formants | |
| will sit near but not on these targets. A real deployment | |
| needs per-speaker calibration or a real training set. | |
| - **Continuous speech.** The classifier assumes the entire | |
| clip contains one word. It does not detect boundaries. | |
| ## Honest limitations | |
| - **The training data is synthetic.** Every number in the | |
| benchmark table comes from a simulator, not from a recording. | |
| Real whispered speech will have different MFCC means and | |
| different harmonicity values. The accuracy numbers here are | |
| upper bounds on the synthetic task. | |
| - **The reverb test is not out-of-distribution.** See the | |
| caveat above. | |
| - **The confusion matrix is perfect.** On the synthetic test | |
| set the classifier does not confuse any pair of digits. This | |
| is because the formant targets are fixed and the noise is | |
| white — a real dataset will have overlap. | |
| - **Coarticulation is not modeled.** In real speech, adjacent | |
| phonemes shift each other's formants. The synthetic words | |
| concatenate independent phonemes without cross-influence. | |
| - **Duration is a feature.** For real speech, duration varies | |
| by speaker and by utterance. A synthetic-trained classifier | |
| that has learned "seven is 0.51 seconds long" may not | |
| generalize. | |
| - **The vocabulary is ten digits.** Extending to 100 words would | |
| require more phonemes and probably more features. | |
| ## What is not novel | |
| - The source-filter model of speech (Fant, 1960). | |
| - MFCC features (Davis & Mermelstein, 1980). | |
| - Whispered speech recognition as a task (Ito et al., 2005; | |
| Zhang & Hansen, 2008). | |
| - The observation that formants survive whispering (well | |
| established in phonetics). | |
| The contribution here is the *controlled pair* — same word, | |
| same formants, two sources — and the demonstration that a | |
| classifier trained on one still works on the other. It is a | |
| clean measurement of how much information the formants carry. | |
| ## Version history | |
| | version | change | | |
| |---|---| | |
| | 0.1.0 | synthesis, features, MLP, 80% whisper accuracy | | |
| | 0.2.0 | trimmed features, extended low band, reverb augmentation; 100% whisper, 83.2% voiced | | |
| ## Reference | |
| Part of a series of small tools built in one session: | |
| | tool | reads | answers | | |
| |---|---|---| | |
| | `ir-source-localizer` | impulse responses | where is the source? | | |
| | `acoustic-lidar` | impulse responses + known source | where are the reflectors? | | |
| | `echo-lineage` | quiet frames | where was it recorded? | | |
| | `noise-color` | the noise floor | what kind of noise is this? | | |
| | **`whisper-decoder`** | **a whispered word** | **which word is it?** | | |
| The tools share one principle: a signal carries information in | |
| the parts that the standard pipeline throws away. For | |
| `whisper-decoder`, the throwaway is the source. The formants | |
| are the signal. | |
| ## License | |
| Apache-2.0 |