--- license: apache-2.0 library_name: onnx language: - en - es - de - nl - it tags: - wake-word - keyword-spotting - openvoiceos - wakeforge - int8 - onnx base_model: TigreGotico/wakehubert-tiny datasets: - TigreGotico/synthetic-wakeword-jarvis - TigreGotico/synthetic-wakeword-alexa - TigreGotico/synthetic-wakeword-ok_nabu - TigreGotico/synthetic-wakeword-hey_jarvis - TigreGotico/synthetic-wakeword-hey_marvin - TigreGotico/synthetic-wakeword-home_assistant - TigreGotico/synthetic-wakeword-hello_nabu - TigreGotico/synthetic-wakeword-computer - TigreGotico/synthetic-wakeword-wake_up - TigreGotico/synthetic-wakeword-hey_chatterbox - TigreGotico/synthetic-wakeword-hey_floyd - TigreGotico/synthetic-wakeword-hey_rhasspy - TigreGotico/synthetic-wakeword-hey_robin - TigreGotico/synthetic-wakeword-marvin - TigreGotico/synthetic-wakeword-sheila - TigreGotico/synthetic-wakeword-stop - TigreGotico/synthetic-wakeword-hey_mycroft - TigreGotico/synthetic-wakeword-android - TigreGotico/synthetic-wakeword-hey_computer - TigreGotico/synthetic-wakeword-hey_k9 - TigreGotico/synthetic-wakeword-hey_scout - TigreGotico/not-wake-words-speech-en - TigreGotico/synthetic-wakeword-despierta - TigreGotico/synthetic-wakeword-aufwachen - TigreGotico/synthetic-wakeword-wakker_worden - TigreGotico/synthetic-wakeword-sveglia pipeline_tag: audio-classification --- # WakeHuBERT wake words Ready wake-word models for [OpenVoiceOS](https://openvoiceos.org), trained with [wakeforge](https://github.com/TigreGotico/wakeforge). Each model is a small GRU classifier that reads features from [WakeHuBERT tiny](https://huggingface.co/TigreGotico/wakehubert-tiny), a 0.64M-parameter speech feature extractor distilled from HuBERT-base. A model scores the last 1.5 s of audio (75 feature frames) and outputs a logit; its sigmoid is the probability that the window holds the wake word. **Try them in your browser:** [WakeHuBERT wake-words Space](https://huggingface.co/spaces/OpenVoiceOS/wakehubert-wakewords-space) runs every model here on your microphone or an uploaded file, with a live score and the calibrated threshold. Everything runs locally in the browser. The repository holds one ONNX file per word under `models/`, and `models.json`, which lists each model's word, featurizer, calibration, default threshold, SHA-256 and measured results. Every model carries the same facts in its own ONNX metadata (`wake_word`, `pretrained_featurizer`, `default_threshold`, `window_frames`, `license`, `training_data`, and on calibrated models `calibrated`, `calib_a` and `calib_b`). | model | word | featurizer | calibrated | default threshold | |---|---|---|---|---| | `wakehubert_jarvis` | jarvis | wakehubert-int8 | yes | 0.57 | | `wakehubert_alexa` | alexa | wakehubert-int8 | yes | 0.44 | | `wakehubert_hey_jarvis` | hey jarvis | wakehubert-int8 | yes | 0.16 | | `wakehubert_hey_marvin` | hey marvin | wakehubert-int8 | yes | 0.34 | | `wakehubert_home_assistant` | home assistant | wakehubert-int8 | yes | 0.19 | | `wakehubert_okay_nabu` | okay nabu | wakehubert-int8 | yes | 0.45 | | `wakehubert_hello_nabu` | hello nabu | wakehubert-int8 | yes | 0.49 | | `wakehubert_hey_chatterbox` | hey chatterbox | wakehubert-int8 | yes | 0.13 | | `wakehubert_hey_floyd` | hey floyd | wakehubert-int8 | yes | 0.43 | | `wakehubert_hey_rhasspy` | hey rhasspy | wakehubert-int8 | yes | 0.40 | | `wakehubert_hey_robin` | hey robin | wakehubert-int8 | yes | 0.05 | | `wakehubert_marvin` | marvin | wakehubert-int8 | yes | 0.06 | | `wakehubert_sheila` | sheila | wakehubert-int8 | yes | 0.47 | | `wakehubert_stop` | stop | wakehubert-int8 | yes | 0.14 | | `wakehubert_android` | android | wakehubert-int8 | yes | 0.42 | | `wakehubert_hey_computer` | hey computer | wakehubert-int8 | yes | 0.31 | | `wakehubert_hey_k9` | hey k9 | wakehubert-int8 | yes | 0.06 | | `wakehubert_hey_scout` | hey scout | wakehubert-int8 | yes | 0.36 | | `wakehubert_wake_up` | wake up | wakehubert-int8 | yes | 0.36 | | `wakehubert_hey_ziggy` | hey ziggy | wakehubert-int8 | yes | 0.27 | | `wakehubert_hey_potato` | hey potato | wakehubert-int8 | yes | 0.15 | | `wakehubert_hey_stemcom` | hey stemcom | wakehubert-int8 | yes | 0.28 | | `wakehubert_computer` | computer | wakehubert (float32) | no | 0.99 | | `wakehubert_hey_mycroft` | hey mycroft | wakehubert-int8 | yes | 0.18 | | `wakehubert_despierta` | despierta (Spanish) | wakehubert-int8 | yes | 0.51 | | `wakehubert_aufwachen` | aufwachen (German) | wakehubert-int8 | yes | 0.57 | | `wakehubert_wakker_worden` | wakker worden (Dutch) | wakehubert-int8 | yes | 0.56 | | `wakehubert_sveglia` | sveglia (Italian) | wakehubert-int8 | yes | 0.47 | ## Use with OpenVoiceOS The models run in [ovos-ww-plugin-wakeforge](https://github.com/OpenVoiceOS/ovos-ww-plugin-wakeforge), which downloads them from this repository on first use, at a pinned revision, checks each file against its `sha256` in `models.json`, and keeps them in the Hugging Face cache; the featurizer comes from [TigreGotico/wakehubert-tiny](https://huggingface.co/TigreGotico/wakehubert-tiny) the same way. Install the plugin and name the model in `mycroft.conf`: ```bash pip install --pre ovos-ww-plugin-wakeforge ``` ```json { "listener": { "wake_word": "hey_jarvis" }, "hotwords": { "hey_jarvis": { "module": "ovos-ww-plugin-wakeforge", "model": "wakehubert_hey_jarvis", "listen": true } } } ``` A model file from this repository also loads by path: set `model` to the local `.onnx` file. The plugin reads the featurizer and the default threshold from the model's metadata. ## Use in your own code The models need only `numpy`, `onnxruntime` and `huggingface_hub`; nothing from OpenVoiceOS. Each model is a small classifier that reads WakeHuBERT-tiny features, so you run two ONNX files: the featurizer, then the wake-word model. ```python import numpy as np import onnxruntime as ort from huggingface_hub import hf_hub_download word = "jarvis" head = ort.InferenceSession(hf_hub_download("OpenVoiceOS/wakehubert-wakewords", f"models/wakehubert_{word}.onnx")) meta = head.get_modelmeta().custom_metadata_map featurizer_file = "wakehubert_int8.onnx" if meta["pretrained_featurizer"] == "wakehubert-int8" else "wakehubert.onnx" feat = ort.InferenceSession(hf_hub_download("TigreGotico/wakehubert-tiny", featurizer_file)) threshold = float(meta["default_threshold"]) WINDOW = 24000 # 1.5 s of 16 kHz audio: the window the models were trained on (75 feature frames) BLOCK = 1280 # score every 80 ms DEBOUNCE_BLOCKS = 25 # ignore 2 s after a detection class WakeWordDetector: def __init__(self): self.buf = np.zeros(WINDOW, np.float32) self.cooldown = 0 def push(self, chunk): """chunk: 1280 float32 samples at 16 kHz in -1..1. Returns (score, detected).""" self.buf = np.concatenate([self.buf, chunk.astype(np.float32)])[-WINDOW:] features = feat.run(None, {"waveform": self.buf[None]})[0] # [1, 75, 128] logit = head.run(None, {"features": features})[0][0] score = float(1.0 / (1.0 + np.exp(-logit))) self.cooldown = max(0, self.cooldown - 1) detected = score >= threshold and self.cooldown == 0 if detected: self.cooldown = DEBOUNCE_BLOCKS return score, detected ``` From a microphone, for example with `sounddevice`: ```python import sounddevice as sd detector = WakeWordDetector() with sd.InputStream(samplerate=16000, channels=1, dtype="float32", blocksize=BLOCK) as stream: while True: chunk, _ = stream.read(BLOCK) score, detected = detector.push(chunk[:, 0]) if detected: print(f"{word} detected (score {score:.2f})") ``` Each block featurizes the last 1.5 s on its own, exactly as the models were trained and scored. The featurizer is causal, the per-block cost is about 1–2 ms on one CPU core for the int8 featurizer, and one featurizer run can feed any number of wake-word models: run `feat` once per block and pass the same `features` to each model. Replace `threshold` to change sensitivity (see below). Checked against the plugin on real recordings: the snippet fires on the same clips. ## Choosing a threshold Set `"threshold"` in the hotword config to trade missed wake words against false activations. A lower value fires more readily and falsely more often; a higher value misses more wake words and fires falsely less often. On a calibrated model the number means the same thing for every word. Its score is mapped so that a threshold of 0.5 gives about one false activation per hour on held-out speech and noise. The default threshold is the point that maximises F2, which weighs recall above precision. Raising the threshold toward 0.8 or 0.9 trades recall for fewer false activations, and lowering it does the opposite. The uncalibrated model (`wakehubert_computer`) outputs a probability too, but its score is not mapped to a false-activation rate. Its threshold is not a calibrated knob, and useful values sit close to 1. ## Results Measured through the plugin at the default threshold and at 0.8. Recall is the share of test clips detected; false activations are counted per hour of negative audio. | model | recall at default | false activations/h at default | recall at 0.8 | false activations/h at 0.8 | recall test set | |---|---|---|---|---|---| | `wakehubert_jarvis` | 98.4% | 0.69 | 93.8% | 0.15 | 384 Picovoice recordings of real speakers | | `wakehubert_alexa` | 98.4% | 0.60 | 91.4% | 0.28 | 315 Picovoice recordings of real speakers | | `wakehubert_hey_jarvis` | 94.5% | 0.09 | 86.2% | 0.02 | 384 clips in held-out synthetic voices | | `wakehubert_hey_marvin` | 97.4% | 0.95 | 90.7% | 0.24 | 386 clips in held-out synthetic voices | | `wakehubert_home_assistant` | 91.1% | 0.30 | 84.2% | 0.15 | 380 clips in held-out synthetic voices | | `wakehubert_okay_nabu` | 92.0% | 0.26 | 75.4% | 0.02 | 386 clips in held-out synthetic voices | | `wakehubert_hello_nabu` | 80.9% | 0.39 | 65.2% | 0.02 | 382 clips in held-out synthetic voices | | `wakehubert_hey_chatterbox` | 82.8% | 0.09 | 61.2% | 0.00 | 116 OVOS community recordings of real speakers | | `wakehubert_hey_floyd` | 91.7% | 0.47 | 82.3% | 0.02 | 96 OVOS community recordings of real speakers | | `wakehubert_hey_rhasspy` | 100.0% | 0.60 | 97.6% | 0.11 | 374 clips in held-out synthetic voices | | `wakehubert_hey_robin` | 99.5% | 0.39 | 97.9% | 0.06 | 380 clips in held-out synthetic voices | | `wakehubert_marvin` | 73.3% | 0.77 | 71.8% | 0.67 | 195 Speech Commands test recordings of real speakers | | `wakehubert_sheila` | 88.7% | 2.08 | 84.4% | 0.54 | 212 Speech Commands test recordings of real speakers | | `wakehubert_stop` | 86.6% | 2.06 | 76.9% | 0.45 | 411 Speech Commands test recordings of real speakers | | `wakehubert_android` | 98.7% | 0.95 | 96.4% | 0.11 | 390 clips in held-out synthetic voices | | `wakehubert_hey_computer` | 96.4% | 0.47 | 93.3% | 0.04 | 390 clips in held-out synthetic voices | | `wakehubert_hey_k9` | 99.4% | 0.19 | 93.5% | 0.04 | 338 clips in held-out synthetic voices | | `wakehubert_hey_scout` | 95.9% | 0.09 | 93.8% | 0.04 | 390 clips in held-out synthetic voices | | `wakehubert_wake_up` | 98.1% | 1.10 | 96.8% | 0.19 | 378 clips in held-out synthetic voices | | `wakehubert_hey_ziggy` | 93.6% | 0.64 | 89.3% | 0.24 | 374 clips of OmniVoice and held-out edge-tts voices | | `wakehubert_hey_potato` | 88.9% | 0.34 | 80.6% | 0.02 | 360 clips in held-out edge-tts voices only, without OmniVoice test clips | | `wakehubert_hey_stemcom` | 89.0% | 0.21 | 83.1% | 0.06 | 337 clips of OmniVoice and held-out edge-tts voices | | `wakehubert_hey_mycroft` | 97.4% | 1.57 | 84.8% | 0.24 | 285 clips in held-out edge-tts voices and 379 Kokoro clips converted to unseen speakers | | `wakehubert_despierta` | 98.4% | 0.58 | 94.1% | 0.03 | 370 Spanish OmniVoice clips with unseen seeds | | `wakehubert_aufwachen` | 99.5% | 0.22 | 97.2% | 0.00 | 390 German OmniVoice clips with unseen seeds | | `wakehubert_wakker_worden` | 99.7% | 0.29 | 98.5% | 0.03 | 390 Dutch OmniVoice clips with unseen seeds | | `wakehubert_sveglia` | 99.5% | 1.16 | 97.7% | 0.10 | 222 Italian OmniVoice clips with unseen seeds | The false activations are counted over 46.5 h of negative audio (speech, non-speech and household audio). The held-out synthetic voices are text-to-speech voices that no training clip uses, converted to the voices of speakers who appear in no training clip. For the localised wake-up models (the words with a language in brackets in the table above), the false activations are counted over the 31.1 h of speech and non-speech audio without the household audio, and recall is measured on clips of OmniVoice voices drawn from seeds that no training clip uses. No figures are published for the uncalibrated model. ## Training data The calibrated models were trained on synthetic speech only, with six edge-tts voices held out of training for testing. The positives are an edge-tts voice grid, further edge-tts and Google Translate TTS voices and OmniVoice clips, and for `wakehubert_jarvis`, `wakehubert_hey_chatterbox`, `wakehubert_hey_floyd`, `wakehubert_marvin`, `wakehubert_sheila` and `wakehubert_stop` also voice-converted copies of edge-tts clips. `wakehubert_alexa` trained instead on the multi-engine, edge-tts and Piper (LibriTTS-R speakers) clips of its dataset and on OmniVoice clips. Each model's `training_data` metadata names its own sources. Most of these clips are published in the `TigreGotico/synthetic-wakeword-` datasets listed above. Speech from LibriSpeech train-clean-100 is mixed into training clips as background babble. The negatives are the wakeforge negative list and half of an AudioSet noise sample; the other half is held out. The localised wake-up models were trained on synthetic speech only: OmniVoice clips with no reference speaker, one seed per clip, kept when a speech recogniser transcript matched the phrase (or, for languages no recogniser covers, when duration, level and speech-activity checks passed), together with the clips of the phrase's `TigreGotico/synthetic-wakeword-` dataset, which also holds the kept OmniVoice clips and their held-out test split. Speech from LibriSpeech train-clean-100 is mixed in as background babble, and the negatives are the same as above. `wakehubert_computer` was trained on synthetic speech only, from `TigreGotico/synthetic-wakeword-computer`, with negatives from `TigreGotico/not-wake-words-speech-en` and AudioSet-derived clips. ## Calibration A calibrated model has an affine map folded into its graph, applied to the GRU's logit before the sigmoid. The map is fitted on a stream of held-out noise and speech that the model did not train on, so that a probability of 0.5 falls at about one false activation per hour of that stream. The default threshold is then the F2-optimal point on the calibrated scale. The fitted slope and offset are in each model's `calib_a` and `calib_b` metadata. ## Limitations - Twenty of the twenty-seven calibrated models are scored on held-out synthetic voices, so recall on real speech can be lower for those words. `wakehubert_jarvis` and `wakehubert_alexa` are scored on the Picovoice recordings, `wakehubert_hey_chatterbox` and `wakehubert_hey_floyd` on OVOS community recordings, and `wakehubert_marvin`, `wakehubert_sheila` and `wakehubert_stop` on Speech Commands, all real speakers. - `wakehubert_sheila` and `wakehubert_stop` give about two false activations per hour at their default threshold; raise it toward 0.8 for about one every two hours. `wakehubert_marvin` has a steep calibration, so its recall and false-activation rate change little between its 0.06 default and 0.8. - The calibration is fitted on about 10 h of audio with few false activations in it (2 to 14 per model), so the map is extrapolated, and "0.5 is about one false activation per hour" is approximate. - On real jarvis recordings, `wakehubert_jarvis` peaks just above its 0.57 default, so a quiet or distant speaker has little margin. - Similar-sounding words trigger each other's model. `wakehubert_jarvis` fires on "hey jarvis", `wakehubert_marvin` on "hey marvin", and `wakehubert_hey_jarvis` on "hey chatterbox". `wakehubert_hello_nabu`, `wakehubert_hey_marvin` and `wakehubert_okay_nabu` can fire on each other's words, `wakehubert_hey_marvin` also on "hey rhasspy", "hey robin" and "marvin", `wakehubert_okay_nabu` on "hey rhasspy", `wakehubert_hey_robin` on "hey marvin", "hey rhasspy" and "okay nabu", `wakehubert_okay_nabu` sometimes on "hey k9", `wakehubert_android` sometimes on "hey floyd", `wakehubert_computer` and `wakehubert_hey_computer` on each other's words, and `wakehubert_sheila` on "computer", all at their default thresholds. Raise the threshold when two of these models run side by side. - The localised wake-up models fire on inflections of the same stem, measured on edge-tts near-word probe clips at their default thresholds: `wakehubert_despierta` on "despiertas", "despierto", "depierta", "desperta" and "despertar", `wakehubert_aufwachen` on "aufmachen" and "aufwachten", and `wakehubert_sveglia` on "sveglio", "sveglie" and "svegliati" and on the rhymes "meraviglia", "bottiglia" and "voglia". `wakehubert_wakker_worden` fired on none of its probes ("wakker", "worden"). Raising the threshold toward 0.8 removes about half of these fires. Each model's `notes` in `models.json` gives its probe count. - The models are English except the localised wake-up models, whose language is in brackets in the table above. The Spanish, German, Dutch and Italian models train on speech in their own language as negatives too, and their calibration stream is 20 h, so their 0.5 point is read without extrapolation. On held-out speech in their own language (Common Voice 17 test and the FLEURS development set) at the default threshold they gave: `wakehubert_despierta` 0.60 per hour over 3.3 h; `wakehubert_aufwachen` 1.85 per hour over 3.2 h; `wakehubert_wakker_worden` 0.00 per hour over 2.4 h; `wakehubert_sveglia` 1.70 per hour over 3.5 h. ## License Apache-2.0. The featurizer, [TigreGotico/wakehubert-tiny](https://huggingface.co/TigreGotico/wakehubert-tiny), is Apache-2.0. The `synthetic-wakeword-*` and `not-wake-words-speech-en` datasets are CC BY 4.0; LibriSpeech is CC BY 4.0; AudioSet labels are CC BY 4.0 and its audio comes from YouTube videos under their uploaders' terms. The voice-conversion and voice-cloning folders of the `synthetic-wakeword-*` datasets take their voices from Mozilla Common Voice contributors, through the MLCommons Multilingual Spoken Words Corpus (CC BY 4.0).