whisper-decoder
A classifier trained only on whispered speech identifies the same words when they are voiced at 83.2% accuracy β 8.3Γ chance.
Whispering removes the pitch source. The vocal folds stop vibrating and the only sound that leaves the mouth is turbulent noise shaped by the vocal tract. The formants survive. The word survives. Only the voicing is gone.
A classifier trained on whispered digits and evaluated on voiced digits recovers the word from the formants alone. This is what the experiment measures: how much of a word's identity lives in the formant structure, independent of the pitch that normal speech provides.
The claim in one sentence
Formant structure is sufficient to identify a closed-vocabulary word even when the voicing source that produced the word is completely different.
What it produces
A closed-vocabulary classifier over ten English digits
(zero through nine). Given a 1.2-second utterance, it returns
the predicted word and a probability distribution over the
vocabulary.
Internally it computes 49 features:
| feature family | count | purpose |
|---|---|---|
| MFCC mean + std | 26 | spectral envelope, formants |
| low-band mel mean + std | 20 | F1 and F2 explicitly, 200β2500 Hz |
| harmonicity mean + std | 2 | voicing detector (near 0 for whisper) |
| trimmed duration | 1 | word length after silence removal |
Install
pip install numpy
No other dependencies for the classifier. matplotlib is optional
for the figures. No downloads. No model weights. The classifier is
5,610 parameters and trains from scratch in 1.6 seconds.
Usage
Run the demo
python whisper_decoder.py
Prints self-test results, generates training and test data,
trains the classifier, and evaluates on three corpora. Writes one
figure to whisper_figures/.
Programmatic
from whisper_decoder import (
MLP, train_mlp, extract_features,
synth_word, generate_corpus, FEATURE_DIM,
)
# Generate a training corpus (whisper mode, 200 samples per word)
train = generate_corpus('whisper', 200, 40, seed=0,
reverb_train_frac=0.25)
# Standardize
mu = train.X_train.mean(axis=0)
sigma = train.X_train.std(axis=0) + 1e-9
X_tr = (train.X_train - mu) / sigma
# Train
model = MLP(FEATURE_DIM, 64, 32, 10, seed=0)
train_mlp(model, X_tr, train.y_train, epochs=200, batch=64)
# Classify a new utterance
sig = synth_word('seven', 'whisper', seed=42)
feats = extract_features(sig)
probs = model.predict_proba(((feats - mu) / sigma)[None, :])[0]
predicted = train.word_index_inv[int(probs.argmax())]
print(predicted, probs.max())
Self-test
python whisper_decoder.py --self-test
11 checks covering the synthesizer, harmonicity, feature extraction, and MFCC reproducibility.
Benchmarks
Trained on 2,000 synthetic whispered utterances (200 per word, 25% with reverb augmentation). Evaluated on 400 held-out utterances.
Primary results
| corpus | accuracy | chance | ratio |
|---|---|---|---|
| voiced speech (domain shift) | 83.2% | 10.0% | 8.3Γ |
| whisper (in-distribution) | 100.0% | 10.0% | 10.0Γ |
| reverberant whisper | 100.0% | 10.0% | 10.0Γ |
The voiced result is the finding. The classifier was never trained on voiced speech. It has never seen a pulse train. The same 49 features β MFCC, low-band mel, harmonicity, duration β still identify the word.
What the confusion matrix looks like
Perfect diagonal on the whispered test set. Every digit is classified correctly, 40 out of 40.
This is a synthetic dataset, not real whispered speech. The 100% is an upper bound on the synthetic task, not a prediction of real-world accuracy.
The reverb row needs a caveat
The reverb augmentation during training and the reverb applied
during testing both use rt60_s=0.35. The training and test
distributions are the same. The 100% measures that the classifier
handles reverb from the family it was trained on. It does not
measure out-of-distribution robustness.
An example script (examples/reverb_ood.py) tests the honest
version: train on rt60=0.35, evaluate on rt60=0.7. Run it to
get the real number.
How it works
Synthesis
The source-filter model of speech production, in its simplest
form. Each word is a sequence of phonemes drawn from a small
inventory. Each phoneme has a formant target: a list of
(frequency_hz, bandwidth_hz, amplitude) triples. The source is
either white noise (whisper) or a pulse train at F0 (voiced).
The formant bank is a sum of Gaussian resonances applied to the
source spectrum.
Same phonemes, same formants, different source. The only thing that changes between a whispered utterance and a voiced utterance of the same word is the excitation.
Features
- MFCC: 13 coefficients, mean and std over frames
- Low-band mel: 10 bands over 200β2500 Hz, mean and std
- Harmonicity: per-frame peak of the normalized autocorrelation in the pitch-lag range (80β300 Hz). Near zero for whisper, 0.5β0.9 for voiced.
- Trimmed duration: length of the non-silent portion
The low-band range is deliberately extended to 2500 Hz because F2
for front vowels (i, e) sits at 2000β2500 Hz. A narrower band
loses three, zero, and five.
Why it works on voiced speech
The formants in the synthesized voiced speech occupy the same frequencies as the formants in the synthesized whispered speech. The classifier learned to weight MFCC and low-band mel, which are formant-driven. Harmonicity is a near-zero feature in the whisper training set and stays near zero in the voiced test set for unvoiced consonants but rises above 0.5 for vowels. The classifier is not thrown off because the other features carry the word identity.
Applications
- Assistive speech interfaces for whispered input. A user who has lost their voice to illness, surgery, or fatigue can whisper and still be understood by a device.
- Security and anonymity. Whispering is used when the speaker does not want to be heard. A decoder that works on whispers is a decoder that works on material recorded from a distance with a directional microphone.
- Speech reconstruction from damaged recordings. A recording that has been aggressively denoised, high-pass filtered, or compressed may still carry enough formant structure. The classifier works on the formants, not on the pitch that is the first thing to be lost.
- Cross-modal keyword spotting. A wake-word detector that triggers whether the wake word is whispered or spoken. Most wake-word systems are trained on voiced speech and fail on whispers.
When to use it
- When the vocabulary is small and fixed.
- When utterances are roughly the same length.
- When the source-filter assumption holds β a single speaker, moderate background noise, no heavy reverb.
When not to use it
- Open-vocabulary transcription. This is not a speech recognizer. It is a 10-way classifier.
- Real whispered speech at scale. The training data is synthetic. Real whispers have coarticulation, breath noise, and speaker-specific formant variation that the simulator does not model.
- Speaker-independent deployment. The formant targets are averaged across speakers. An individual speaker's formants will sit near but not on these targets. A real deployment needs per-speaker calibration or a real training set.
- Continuous speech. The classifier assumes the entire clip contains one word. It does not detect boundaries.
Honest limitations
- The training data is synthetic. Every number in the benchmark table comes from a simulator, not from a recording. Real whispered speech will have different MFCC means and different harmonicity values. The accuracy numbers here are upper bounds on the synthetic task.
- The reverb test is not out-of-distribution. See the caveat above.
- The confusion matrix is perfect. On the synthetic test set the classifier does not confuse any pair of digits. This is because the formant targets are fixed and the noise is white β a real dataset will have overlap.
- Coarticulation is not modeled. In real speech, adjacent phonemes shift each other's formants. The synthetic words concatenate independent phonemes without cross-influence.
- Duration is a feature. For real speech, duration varies by speaker and by utterance. A synthetic-trained classifier that has learned "seven is 0.51 seconds long" may not generalize.
- The vocabulary is ten digits. Extending to 100 words would require more phonemes and probably more features.
What is not novel
- The source-filter model of speech (Fant, 1960).
- MFCC features (Davis & Mermelstein, 1980).
- Whispered speech recognition as a task (Ito et al., 2005; Zhang & Hansen, 2008).
- The observation that formants survive whispering (well established in phonetics).
The contribution here is the controlled pair β same word, same formants, two sources β and the demonstration that a classifier trained on one still works on the other. It is a clean measurement of how much information the formants carry.
Version history
| version | change |
|---|---|
| 0.1.0 | synthesis, features, MLP, 80% whisper accuracy |
| 0.2.0 | trimmed features, extended low band, reverb augmentation; 100% whisper, 83.2% voiced |
Reference
Part of a series of small tools built in one session:
| tool | reads | answers |
|---|---|---|
ir-source-localizer |
impulse responses | where is the source? |
acoustic-lidar |
impulse responses + known source | where are the reflectors? |
echo-lineage |
quiet frames | where was it recorded? |
noise-color |
the noise floor | what kind of noise is this? |
whisper-decoder |
a whispered word | which word is it? |
The tools share one principle: a signal carries information in
the parts that the standard pipeline throws away. For
whisper-decoder, the throwaway is the source. The formants
are the signal.
License
Apache-2.0