WakeHuBERT tiny (streaming speech features for wake words)

TigreGotico's 0.64M-parameter causal feature extractor, distilled from HuBERT-base for wake-word detection, exported for loom.cpp in four precisions. Family 13: audio in, one 128-dimensional feature vector per 20 ms frame out.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from TigreGotico/wakehubert-tiny. The F32 file carries the same parameters unmodified, in loom.cpp's GGUF format; the others are quantized from them (see Files).

License

apache-2.0, inherited from the base model above.

Language(s)

en, de, nl, fr, es, it, pt, pl

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub]"
import loom

# This repo holds four precisions of one model; name the file you want (see "Files" below).
model = loom.Model.from_pretrained("loom-ai-org/wakehubert-tiny-loom", "wakehubert-tiny-f32.gguf")

# Audio is a mono float list at 16 kHz. The answer is one row of features per frame, and each row
# knows when its frame starts.
features = model.speech2embeddings.infer(audio)
print(len(features), features.dim)
# one row per 20 ms of audio, 128 features each: 550 128 for an 11 s clip
for start, row in zip(features.times[:3], features.rows[:3]):
    print(f"{start:.2f}s", [round(x, 3) for x in row[:4]])

# Each row describes the speech about 100 ms before its frame ends; a wake-word classifier is trained
# on a window of these rows, and its decision rule is yours.

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Known limitations

Needs loom 1.0.0-rc15 or later. This file's embeddings are per FRAME: model.speech2embeddings.infer(audio) returns a FrameEmbeddings (rows, times, dim, frame_rate), cut by the width the file declares. rc14 and earlier refuse the file at that door with an error naming the granularity; there, model.infer(waveform=audio) returns the same frames as one flat row-major list, 128 values per frame.

One call is one whole clip, offline. The model is strictly causal -- frame t depends only on audio before sample 320 x (t + 1), within a 2.5 s receptive field -- and it was trained to describe each HuBERT frame 100 ms late, so a detector built on it reacts about that long after the word ends. To stream it, call it on a sliding buffer that keeps 40,000 samples (2.5 s) of context and take the newest frames: those then equal the offline ones. A clip shorter than 320 samples gives no frames. Audio must be mono 16 kHz.

Every precision keeps the front end in F32. The log-mel's DFT basis is the one weight this model cannot spare precision in: at F16 its rounding in the near-empty high bins of a loud frame is amplified by the log, which took the worst frame to cosine 0.976 before the basis was exempted. Upstream's own int8 file keeps its front end in float too.

Upstream evaluates it on one English wake word and one benchmark, and notes that agreement with the teacher predicts detection quality poorly: judge it, and any of the quantized files, by a detector trained on its features.

Files

One model in several precisions; every file carries the same graph and driver. The quantized ones pack the convolution and projection weights (loom-export --quantize <type>), so their numbers differ from the original model's by the amount noted.

  • wakehubert-tiny-f32.gguf -- full precision, 3.3 MB. Matches upstream's PyTorch model to float rounding: largest difference 3.4e-5 on features up to 11.5, cosine 1.000000 on every frame.
  • wakehubert-tiny-f16.gguf -- 2.0 MB. Mean cosine 1.000000 to the PyTorch model, worst frame 0.999998. A size choice, not a speed one: on an x86 CPU it runs about 1.4x SLOWER than F32 (the quantized files run at F32's speed).
  • wakehubert-tiny-q8_0.gguf -- 1.4 MB. Mean cosine 0.99992, worst frame 0.99978 -- closer than upstream's own wakehubert_int8.onnx (mean 0.9975).
  • wakehubert-tiny-q4_1.gguf -- 1.2 MB. Mean cosine 0.987, worst frame 0.972: a measurably different feature space. A detector trained on the F32 features needs re-validating on these, or training on them.
Downloads last month
125
GGUF
Model size
811k params
Architecture
loom-wakehubert
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for loom-ai-org/wakehubert-tiny-loom

Quantized
(3)
this model

Collection including loom-ai-org/wakehubert-tiny-loom