--- license: apache-2.0 library_name: whisper-decoder tags: - speech - whisper - formants - mfcc - keyword-spotting - numpy - cpu pipeline_tag: audio-classification language: - en --- # whisper-decoder **A classifier trained only on whispered speech identifies the same words when they are voiced at 83.2% accuracy — 8.3× chance.** Whispering removes the pitch source. The vocal folds stop vibrating and the only sound that leaves the mouth is turbulent noise shaped by the vocal tract. The formants survive. The word survives. Only the voicing is gone. A classifier trained on whispered digits and evaluated on voiced digits recovers the word from the formants alone. This is what the experiment measures: how much of a word's identity lives in the formant structure, independent of the pitch that normal speech provides. ## The claim in one sentence Formant structure is sufficient to identify a closed-vocabulary word even when the voicing source that produced the word is completely different. ## What it produces A closed-vocabulary classifier over ten English digits (`zero` through `nine`). Given a 1.2-second utterance, it returns the predicted word and a probability distribution over the vocabulary. Internally it computes 49 features: | feature family | count | purpose | |---|---|---| | MFCC mean + std | 26 | spectral envelope, formants | | low-band mel mean + std | 20 | F1 and F2 explicitly, 200–2500 Hz | | harmonicity mean + std | 2 | voicing detector (near 0 for whisper) | | trimmed duration | 1 | word length after silence removal | ## Install ```bash pip install numpy ``` No other dependencies for the classifier. `matplotlib` is optional for the figures. No downloads. No model weights. The classifier is 5,610 parameters and trains from scratch in 1.6 seconds. ## Usage ### Run the demo ```bash python whisper_decoder.py ``` Prints self-test results, generates training and test data, trains the classifier, and evaluates on three corpora. Writes one figure to `whisper_figures/`. ### Programmatic ```python from whisper_decoder import ( MLP, train_mlp, extract_features, synth_word, generate_corpus, FEATURE_DIM, ) # Generate a training corpus (whisper mode, 200 samples per word) train = generate_corpus('whisper', 200, 40, seed=0, reverb_train_frac=0.25) # Standardize mu = train.X_train.mean(axis=0) sigma = train.X_train.std(axis=0) + 1e-9 X_tr = (train.X_train - mu) / sigma # Train model = MLP(FEATURE_DIM, 64, 32, 10, seed=0) train_mlp(model, X_tr, train.y_train, epochs=200, batch=64) # Classify a new utterance sig = synth_word('seven', 'whisper', seed=42) feats = extract_features(sig) probs = model.predict_proba(((feats - mu) / sigma)[None, :])[0] predicted = train.word_index_inv[int(probs.argmax())] print(predicted, probs.max()) ``` ### Self-test ```bash python whisper_decoder.py --self-test ``` 11 checks covering the synthesizer, harmonicity, feature extraction, and MFCC reproducibility. ## Benchmarks Trained on 2,000 synthetic whispered utterances (200 per word, 25% with reverb augmentation). Evaluated on 400 held-out utterances. ### Primary results | corpus | accuracy | chance | ratio | |---|---|---|---| | **voiced speech (domain shift)** | **83.2%** | 10.0% | **8.3×** | | whisper (in-distribution) | 100.0% | 10.0% | 10.0× | | reverberant whisper | 100.0% | 10.0% | 10.0× | **The voiced result is the finding.** The classifier was never trained on voiced speech. It has never seen a pulse train. The same 49 features — MFCC, low-band mel, harmonicity, duration — still identify the word. ### What the confusion matrix looks like Perfect diagonal on the whispered test set. Every digit is classified correctly, 40 out of 40. This is a synthetic dataset, not real whispered speech. The 100% is an upper bound on the synthetic task, not a prediction of real-world accuracy. ### The reverb row needs a caveat The reverb augmentation during training and the reverb applied during testing both use `rt60_s=0.35`. The training and test distributions are the same. The 100% measures that the classifier handles reverb *from the family it was trained on*. It does not measure out-of-distribution robustness. An example script (`examples/reverb_ood.py`) tests the honest version: train on `rt60=0.35`, evaluate on `rt60=0.7`. Run it to get the real number. ## How it works ### Synthesis The source-filter model of speech production, in its simplest form. Each word is a sequence of phonemes drawn from a small inventory. Each phoneme has a formant target: a list of `(frequency_hz, bandwidth_hz, amplitude)` triples. The source is either white noise (whisper) or a pulse train at F0 (voiced). The formant bank is a sum of Gaussian resonances applied to the source spectrum. Same phonemes, same formants, different source. The only thing that changes between a whispered utterance and a voiced utterance of the same word is the excitation. ### Features - **MFCC**: 13 coefficients, mean and std over frames - **Low-band mel**: 10 bands over 200–2500 Hz, mean and std - **Harmonicity**: per-frame peak of the normalized autocorrelation in the pitch-lag range (80–300 Hz). Near zero for whisper, 0.5–0.9 for voiced. - **Trimmed duration**: length of the non-silent portion The low-band range is deliberately extended to 2500 Hz because F2 for front vowels (`i`, `e`) sits at 2000–2500 Hz. A narrower band loses `three`, `zero`, and `five`. ### Why it works on voiced speech The formants in the synthesized voiced speech occupy the same frequencies as the formants in the synthesized whispered speech. The classifier learned to weight MFCC and low-band mel, which are formant-driven. Harmonicity is a near-zero feature in the whisper training set and stays near zero in the voiced test set for unvoiced consonants but rises above 0.5 for vowels. The classifier is not thrown off because the other features carry the word identity. ## Applications - **Assistive speech interfaces for whispered input.** A user who has lost their voice to illness, surgery, or fatigue can whisper and still be understood by a device. - **Security and anonymity.** Whispering is used when the speaker does not want to be heard. A decoder that works on whispers is a decoder that works on material recorded from a distance with a directional microphone. - **Speech reconstruction from damaged recordings.** A recording that has been aggressively denoised, high-pass filtered, or compressed may still carry enough formant structure. The classifier works on the formants, not on the pitch that is the first thing to be lost. - **Cross-modal keyword spotting.** A wake-word detector that triggers whether the wake word is whispered or spoken. Most wake-word systems are trained on voiced speech and fail on whispers. ## When to use it - When the vocabulary is small and fixed. - When utterances are roughly the same length. - When the source-filter assumption holds — a single speaker, moderate background noise, no heavy reverb. ## When *not* to use it - **Open-vocabulary transcription.** This is not a speech recognizer. It is a 10-way classifier. - **Real whispered speech at scale.** The training data is synthetic. Real whispers have coarticulation, breath noise, and speaker-specific formant variation that the simulator does not model. - **Speaker-independent deployment.** The formant targets are averaged across speakers. An individual speaker's formants will sit near but not on these targets. A real deployment needs per-speaker calibration or a real training set. - **Continuous speech.** The classifier assumes the entire clip contains one word. It does not detect boundaries. ## Honest limitations - **The training data is synthetic.** Every number in the benchmark table comes from a simulator, not from a recording. Real whispered speech will have different MFCC means and different harmonicity values. The accuracy numbers here are upper bounds on the synthetic task. - **The reverb test is not out-of-distribution.** See the caveat above. - **The confusion matrix is perfect.** On the synthetic test set the classifier does not confuse any pair of digits. This is because the formant targets are fixed and the noise is white — a real dataset will have overlap. - **Coarticulation is not modeled.** In real speech, adjacent phonemes shift each other's formants. The synthetic words concatenate independent phonemes without cross-influence. - **Duration is a feature.** For real speech, duration varies by speaker and by utterance. A synthetic-trained classifier that has learned "seven is 0.51 seconds long" may not generalize. - **The vocabulary is ten digits.** Extending to 100 words would require more phonemes and probably more features. ## What is not novel - The source-filter model of speech (Fant, 1960). - MFCC features (Davis & Mermelstein, 1980). - Whispered speech recognition as a task (Ito et al., 2005; Zhang & Hansen, 2008). - The observation that formants survive whispering (well established in phonetics). The contribution here is the *controlled pair* — same word, same formants, two sources — and the demonstration that a classifier trained on one still works on the other. It is a clean measurement of how much information the formants carry. ## Version history | version | change | |---|---| | 0.1.0 | synthesis, features, MLP, 80% whisper accuracy | | 0.2.0 | trimmed features, extended low band, reverb augmentation; 100% whisper, 83.2% voiced | ## Reference Part of a series of small tools built in one session: | tool | reads | answers | |---|---|---| | `ir-source-localizer` | impulse responses | where is the source? | | `acoustic-lidar` | impulse responses + known source | where are the reflectors? | | `echo-lineage` | quiet frames | where was it recorded? | | `noise-color` | the noise floor | what kind of noise is this? | | **`whisper-decoder`** | **a whispered word** | **which word is it?** | The tools share one principle: a signal carries information in the parts that the standard pipeline throws away. For `whisper-decoder`, the throwaway is the source. The formants are the signal. ## License Apache-2.0