Moonshine EOU Head β End-of-Utterance Classifier
A lightweight end-of-utterance (EOU) confirmation head trained on top of Moonshine-tiny encoder features. Designed to run alongside Silero VAD as a secondary contextual confirmer β when VAD detects 50-80ms of silence, this model confirms whether the speaker is actually done talking or just pausing.
Architecture
- Encoder: Moonshine-tiny (7.7M params, frozen) β produces 288-dim frame-level features
- EOU Head: BiGRU (2 layers, 128D) + attention pooling + classifier (565K trainable params)
Linear(288 β 128) + GELUBiGRU(128, 2 layers, bidirectional)Attention pooling over GRU outputsLayerNorm(256) β Linear(256 β 128) + GELU β Linear(128 β 1)
Performance
Trained on the full SmartTurn v3.2 dataset (271K clips, 50/50 balanced):
| Metric | Value |
|---|---|
| AUC | 0.9665 |
| Accuracy | 0.908 |
| F1 | 0.910 |
Inference Latency (CPU, 4 threads)
Full pipeline is 100% ONNX Runtime β no PyTorch needed at inference.
| Component | 3s audio |
|---|---|
| Silero VAD (per 32ms chunk) | ~1ms |
| Moonshine encoder (ONNX fp32) | ~25ms |
| EOU head (ONNX) | ~1ms |
| Total EOU decision | ~26ms |
Quick Start
Clone and run β no pip install needed, just numpy and onnxruntime:
from eou_gate import create_engine, Event
# Once per process β loads all 3 ONNX models
engine = create_engine(model_dir="models")
# Per call / user β lightweight state only
stream = engine.create_stream(min_silence_ms=60)
# Feed 16kHz float32 audio chunks (any size)
for chunk in audio_chunks:
event = stream.process_chunk(chunk)
if event == Event.TURN_END:
print("User finished speaking")
elif event == Event.PAUSE:
print("User is pausing (not done yet)")
# Or async
event = await stream.aprocess_chunk(chunk)
# Direct EOU probability (no VAD, call on demand)
prob = engine.eou_infer(audio_3s) # -> float in [0, 1]
Multi-concurrent streams
The engine is a shared singleton; each stream holds only its own VAD state and ring buffer. Thread pool handles concurrent inference:
import asyncio
engine = create_engine(model_dir="models")
async def handle_call(user_audio_stream):
stream = engine.create_stream(min_silence_ms=60)
async for chunk in user_audio_stream:
event = await stream.aprocess_chunk(chunk)
if event == Event.TURN_END:
# respond to user
stream.reset()
# 8 concurrent calls sharing one engine
await asyncio.gather(*[handle_call(s) for s in streams])
Standalone EOU inference (no VAD)
import onnxruntime as ort
import numpy as np
enc = ort.InferenceSession("models/moonshine_enc_fp32.onnx")
head = ort.InferenceSession("models/eou_head.onnx")
# Raw 16kHz float32 audio β EOU probability
audio = load_audio("clip.wav") # (samples,) float32
features = enc.run(None, {"input_values": audio.reshape(1, -1)})[0]
logit = head.run(None, {
"encoder_features": features,
"frame_lengths": np.array([features.shape[1]], dtype=np.int64),
})[0]
prob = 1.0 / (1.0 + np.exp(-logit[0]))
# prob > 0.5 β turn is done
Files
models/
βββ silero_vad.onnx # Silero VAD v5 (2.2 MB)
βββ moonshine_enc_fp32.onnx # Moonshine-tiny encoder ONNX (0.9 MB + 30 MB data)
βββ moonshine_enc_fp32.onnx.data # Encoder weights (external data)
βββ eou_head.onnx # Trained EOU head ONNX (2.2 MB)
eou_gate/ # Drop-in Python package
βββ __init__.py # create_engine() factory
βββ engine.py # EOUEngine β loads & runs all models
βββ stream.py # StreamState β per-user VAD + EOU state
βββ audio.py # RingBuffer β circular audio buffer
train.py # Full training script (Colab/RunPod ready)
eou_head_best.pt # PyTorch checkpoint (for fine-tuning)
How It Works
- Silero VAD processes 32ms audio chunks, detecting speech vs silence
- When silence reaches
min_silence_ms(e.g. 60ms) after speech: - Moonshine encoder processes the last 3s of audio context
- EOU head classifies the encoder features β "turn done" or "still pausing"
- If EOU confirms β
TURN_END; if not βPAUSE(keep listening)
This lets min_silence_ms drop from ~256ms to ~60-80ms β making voice agents feel much more responsive without cutting people off mid-sentence.
Training
- Dataset: pipecat-ai/smart-turn-data-v3.2-train (271K clips) / test (31.5K clips)
- Encoder: Moonshine-tiny, frozen, GPU
- Optimizer: AdamW, lr=3e-4, weight_decay=0.01
- Scheduler: OneCycleLR, 3 epochs
- Loss: BCEWithLogitsLoss
- Audio: 16kHz, max 4s (last 4s kept if longer)
- GPU: RunPod A40
- Training time: ~3 hours
Runtime Dependencies
Only two runtime dependencies β no PyTorch or transformers needed:
numpy
onnxruntime
Intended Use
This model is a confirmation gate, not a standalone VAD. It is designed to:
- Run on-demand when Silero VAD detects brief silence (50-80ms)
- Confirm whether the silence marks a true turn end or a mid-utterance pause
- Allow
min_silence_msto drop from ~256ms to ~60-80ms for faster turn-taking in voice agents
Credits
- Encoder: Moonshine-tiny by Useful Sensors
- VAD: Silero VAD v5
- Dataset: SmartTurn v3.2 by Pipecat AI
License
MIT
Model tree for AswanthCManoj/moonshine-eou-head
Base model
moonshine-ai/moonshine-tiny