whisper-decoder / README.md
zeechimp's picture
Create README.md
5691579 verified
|
Raw History Blame Contribute Delete
10.3 kB
---
license: apache-2.0
library_name: whisper-decoder
tags:
- speech
- whisper
- formants
- mfcc
- keyword-spotting
- numpy
- cpu
pipeline_tag: audio-classification
language:
- en
---
# whisper-decoder
**A classifier trained only on whispered speech identifies the
same words when they are voiced at 83.2% accuracy — 8.3× chance.**
Whispering removes the pitch source. The vocal folds stop vibrating
and the only sound that leaves the mouth is turbulent noise shaped
by the vocal tract. The formants survive. The word survives. Only
the voicing is gone.
A classifier trained on whispered digits and evaluated on voiced
digits recovers the word from the formants alone. This is what the
experiment measures: how much of a word's identity lives in the
formant structure, independent of the pitch that normal speech
provides.
## The claim in one sentence
Formant structure is sufficient to identify a closed-vocabulary
word even when the voicing source that produced the word is
completely different.
## What it produces
A closed-vocabulary classifier over ten English digits
(`zero` through `nine`). Given a 1.2-second utterance, it returns
the predicted word and a probability distribution over the
vocabulary.
Internally it computes 49 features:
| feature family | count | purpose |
|---|---|---|
| MFCC mean + std | 26 | spectral envelope, formants |
| low-band mel mean + std | 20 | F1 and F2 explicitly, 200–2500 Hz |
| harmonicity mean + std | 2 | voicing detector (near 0 for whisper) |
| trimmed duration | 1 | word length after silence removal |
## Install
```bash
pip install numpy
```
No other dependencies for the classifier. `matplotlib` is optional
for the figures. No downloads. No model weights. The classifier is
5,610 parameters and trains from scratch in 1.6 seconds.
## Usage
### Run the demo
```bash
python whisper_decoder.py
```
Prints self-test results, generates training and test data,
trains the classifier, and evaluates on three corpora. Writes one
figure to `whisper_figures/`.
### Programmatic
```python
from whisper_decoder import (
MLP, train_mlp, extract_features,
synth_word, generate_corpus, FEATURE_DIM,
)
# Generate a training corpus (whisper mode, 200 samples per word)
train = generate_corpus('whisper', 200, 40, seed=0,
reverb_train_frac=0.25)
# Standardize
mu = train.X_train.mean(axis=0)
sigma = train.X_train.std(axis=0) + 1e-9
X_tr = (train.X_train - mu) / sigma
# Train
model = MLP(FEATURE_DIM, 64, 32, 10, seed=0)
train_mlp(model, X_tr, train.y_train, epochs=200, batch=64)
# Classify a new utterance
sig = synth_word('seven', 'whisper', seed=42)
feats = extract_features(sig)
probs = model.predict_proba(((feats - mu) / sigma)[None, :])[0]
predicted = train.word_index_inv[int(probs.argmax())]
print(predicted, probs.max())
```
### Self-test
```bash
python whisper_decoder.py --self-test
```
11 checks covering the synthesizer, harmonicity, feature
extraction, and MFCC reproducibility.
## Benchmarks
Trained on 2,000 synthetic whispered utterances (200 per word,
25% with reverb augmentation). Evaluated on 400 held-out
utterances.
### Primary results
| corpus | accuracy | chance | ratio |
|---|---|---|---|
| **voiced speech (domain shift)** | **83.2%** | 10.0% | **8.3×** |
| whisper (in-distribution) | 100.0% | 10.0% | 10.0× |
| reverberant whisper | 100.0% | 10.0% | 10.0× |
**The voiced result is the finding.** The classifier was never
trained on voiced speech. It has never seen a pulse train. The
same 49 features — MFCC, low-band mel, harmonicity, duration —
still identify the word.
### What the confusion matrix looks like
Perfect diagonal on the whispered test set. Every digit is
classified correctly, 40 out of 40.
This is a synthetic dataset, not real whispered speech. The
100% is an upper bound on the synthetic task, not a prediction
of real-world accuracy.
### The reverb row needs a caveat
The reverb augmentation during training and the reverb applied
during testing both use `rt60_s=0.35`. The training and test
distributions are the same. The 100% measures that the classifier
handles reverb *from the family it was trained on*. It does not
measure out-of-distribution robustness.
An example script (`examples/reverb_ood.py`) tests the honest
version: train on `rt60=0.35`, evaluate on `rt60=0.7`. Run it to
get the real number.
## How it works
### Synthesis
The source-filter model of speech production, in its simplest
form. Each word is a sequence of phonemes drawn from a small
inventory. Each phoneme has a formant target: a list of
`(frequency_hz, bandwidth_hz, amplitude)` triples. The source is
either white noise (whisper) or a pulse train at F0 (voiced).
The formant bank is a sum of Gaussian resonances applied to the
source spectrum.
Same phonemes, same formants, different source. The only thing
that changes between a whispered utterance and a voiced utterance
of the same word is the excitation.
### Features
- **MFCC**: 13 coefficients, mean and std over frames
- **Low-band mel**: 10 bands over 200–2500 Hz, mean and std
- **Harmonicity**: per-frame peak of the normalized autocorrelation
in the pitch-lag range (80–300 Hz). Near zero for whisper, 0.5–0.9
for voiced.
- **Trimmed duration**: length of the non-silent portion
The low-band range is deliberately extended to 2500 Hz because F2
for front vowels (`i`, `e`) sits at 2000–2500 Hz. A narrower band
loses `three`, `zero`, and `five`.
### Why it works on voiced speech
The formants in the synthesized voiced speech occupy the same
frequencies as the formants in the synthesized whispered speech.
The classifier learned to weight MFCC and low-band mel, which are
formant-driven. Harmonicity is a near-zero feature in the
whisper training set and stays near zero in the voiced test set
for unvoiced consonants but rises above 0.5 for vowels. The
classifier is not thrown off because the other features carry
the word identity.
## Applications
- **Assistive speech interfaces for whispered input.** A user who
has lost their voice to illness, surgery, or fatigue can whisper
and still be understood by a device.
- **Security and anonymity.** Whispering is used when the speaker
does not want to be heard. A decoder that works on whispers is
a decoder that works on material recorded from a distance with
a directional microphone.
- **Speech reconstruction from damaged recordings.** A recording
that has been aggressively denoised, high-pass filtered, or
compressed may still carry enough formant structure. The
classifier works on the formants, not on the pitch that is the
first thing to be lost.
- **Cross-modal keyword spotting.** A wake-word detector that
triggers whether the wake word is whispered or spoken. Most
wake-word systems are trained on voiced speech and fail on
whispers.
## When to use it
- When the vocabulary is small and fixed.
- When utterances are roughly the same length.
- When the source-filter assumption holds — a single speaker,
moderate background noise, no heavy reverb.
## When *not* to use it
- **Open-vocabulary transcription.** This is not a speech
recognizer. It is a 10-way classifier.
- **Real whispered speech at scale.** The training data is
synthetic. Real whispers have coarticulation, breath noise,
and speaker-specific formant variation that the simulator
does not model.
- **Speaker-independent deployment.** The formant targets are
averaged across speakers. An individual speaker's formants
will sit near but not on these targets. A real deployment
needs per-speaker calibration or a real training set.
- **Continuous speech.** The classifier assumes the entire
clip contains one word. It does not detect boundaries.
## Honest limitations
- **The training data is synthetic.** Every number in the
benchmark table comes from a simulator, not from a recording.
Real whispered speech will have different MFCC means and
different harmonicity values. The accuracy numbers here are
upper bounds on the synthetic task.
- **The reverb test is not out-of-distribution.** See the
caveat above.
- **The confusion matrix is perfect.** On the synthetic test
set the classifier does not confuse any pair of digits. This
is because the formant targets are fixed and the noise is
white — a real dataset will have overlap.
- **Coarticulation is not modeled.** In real speech, adjacent
phonemes shift each other's formants. The synthetic words
concatenate independent phonemes without cross-influence.
- **Duration is a feature.** For real speech, duration varies
by speaker and by utterance. A synthetic-trained classifier
that has learned "seven is 0.51 seconds long" may not
generalize.
- **The vocabulary is ten digits.** Extending to 100 words would
require more phonemes and probably more features.
## What is not novel
- The source-filter model of speech (Fant, 1960).
- MFCC features (Davis & Mermelstein, 1980).
- Whispered speech recognition as a task (Ito et al., 2005;
Zhang & Hansen, 2008).
- The observation that formants survive whispering (well
established in phonetics).
The contribution here is the *controlled pair* — same word,
same formants, two sources — and the demonstration that a
classifier trained on one still works on the other. It is a
clean measurement of how much information the formants carry.
## Version history
| version | change |
|---|---|
| 0.1.0 | synthesis, features, MLP, 80% whisper accuracy |
| 0.2.0 | trimmed features, extended low band, reverb augmentation; 100% whisper, 83.2% voiced |
## Reference
Part of a series of small tools built in one session:
| tool | reads | answers |
|---|---|---|
| `ir-source-localizer` | impulse responses | where is the source? |
| `acoustic-lidar` | impulse responses + known source | where are the reflectors? |
| `echo-lineage` | quiet frames | where was it recorded? |
| `noise-color` | the noise floor | what kind of noise is this? |
| **`whisper-decoder`** | **a whispered word** | **which word is it?** |
The tools share one principle: a signal carries information in
the parts that the standard pipeline throws away. For
`whisper-decoder`, the throwaway is the source. The formants
are the signal.
## License
Apache-2.0