Moonshine EOU Head β€” End-of-Utterance Classifier

A lightweight end-of-utterance (EOU) confirmation head trained on top of Moonshine-tiny encoder features. Designed to run alongside Silero VAD as a secondary contextual confirmer β€” when VAD detects 50-80ms of silence, this model confirms whether the speaker is actually done talking or just pausing.

Architecture

  • Encoder: Moonshine-tiny (7.7M params, frozen) β€” produces 288-dim frame-level features
  • EOU Head: BiGRU (2 layers, 128D) + attention pooling + classifier (565K trainable params)
    • Linear(288 β†’ 128) + GELU
    • BiGRU(128, 2 layers, bidirectional)
    • Attention pooling over GRU outputs
    • LayerNorm(256) β†’ Linear(256 β†’ 128) + GELU β†’ Linear(128 β†’ 1)

Performance

Trained on the full SmartTurn v3.2 dataset (271K clips, 50/50 balanced):

Metric Value
AUC 0.9665
Accuracy 0.908
F1 0.910

Inference Latency (CPU, 4 threads)

Full pipeline is 100% ONNX Runtime β€” no PyTorch needed at inference.

Component 3s audio
Silero VAD (per 32ms chunk) ~1ms
Moonshine encoder (ONNX fp32) ~25ms
EOU head (ONNX) ~1ms
Total EOU decision ~26ms

Quick Start

Clone and run β€” no pip install needed, just numpy and onnxruntime:

from eou_gate import create_engine, Event

# Once per process β€” loads all 3 ONNX models
engine = create_engine(model_dir="models")

# Per call / user β€” lightweight state only
stream = engine.create_stream(min_silence_ms=60)

# Feed 16kHz float32 audio chunks (any size)
for chunk in audio_chunks:
    event = stream.process_chunk(chunk)
    if event == Event.TURN_END:
        print("User finished speaking")
    elif event == Event.PAUSE:
        print("User is pausing (not done yet)")

# Or async
event = await stream.aprocess_chunk(chunk)

# Direct EOU probability (no VAD, call on demand)
prob = engine.eou_infer(audio_3s)  # -> float in [0, 1]

Multi-concurrent streams

The engine is a shared singleton; each stream holds only its own VAD state and ring buffer. Thread pool handles concurrent inference:

import asyncio

engine = create_engine(model_dir="models")

async def handle_call(user_audio_stream):
    stream = engine.create_stream(min_silence_ms=60)
    async for chunk in user_audio_stream:
        event = await stream.aprocess_chunk(chunk)
        if event == Event.TURN_END:
            # respond to user
            stream.reset()

# 8 concurrent calls sharing one engine
await asyncio.gather(*[handle_call(s) for s in streams])

Standalone EOU inference (no VAD)

import onnxruntime as ort
import numpy as np

enc = ort.InferenceSession("models/moonshine_enc_fp32.onnx")
head = ort.InferenceSession("models/eou_head.onnx")

# Raw 16kHz float32 audio β†’ EOU probability
audio = load_audio("clip.wav")  # (samples,) float32
features = enc.run(None, {"input_values": audio.reshape(1, -1)})[0]
logit = head.run(None, {
    "encoder_features": features,
    "frame_lengths": np.array([features.shape[1]], dtype=np.int64),
})[0]
prob = 1.0 / (1.0 + np.exp(-logit[0]))
# prob > 0.5 β†’ turn is done

Files

models/
β”œβ”€β”€ silero_vad.onnx              # Silero VAD v5 (2.2 MB)
β”œβ”€β”€ moonshine_enc_fp32.onnx      # Moonshine-tiny encoder ONNX (0.9 MB + 30 MB data)
β”œβ”€β”€ moonshine_enc_fp32.onnx.data # Encoder weights (external data)
└── eou_head.onnx                # Trained EOU head ONNX (2.2 MB)

eou_gate/                        # Drop-in Python package
β”œβ”€β”€ __init__.py                  # create_engine() factory
β”œβ”€β”€ engine.py                    # EOUEngine β€” loads & runs all models
β”œβ”€β”€ stream.py                    # StreamState β€” per-user VAD + EOU state
└── audio.py                     # RingBuffer β€” circular audio buffer

train.py                         # Full training script (Colab/RunPod ready)
eou_head_best.pt                 # PyTorch checkpoint (for fine-tuning)

How It Works

  1. Silero VAD processes 32ms audio chunks, detecting speech vs silence
  2. When silence reaches min_silence_ms (e.g. 60ms) after speech:
  3. Moonshine encoder processes the last 3s of audio context
  4. EOU head classifies the encoder features β†’ "turn done" or "still pausing"
  5. If EOU confirms β†’ TURN_END; if not β†’ PAUSE (keep listening)

This lets min_silence_ms drop from ~256ms to ~60-80ms β€” making voice agents feel much more responsive without cutting people off mid-sentence.

Training

  • Dataset: pipecat-ai/smart-turn-data-v3.2-train (271K clips) / test (31.5K clips)
  • Encoder: Moonshine-tiny, frozen, GPU
  • Optimizer: AdamW, lr=3e-4, weight_decay=0.01
  • Scheduler: OneCycleLR, 3 epochs
  • Loss: BCEWithLogitsLoss
  • Audio: 16kHz, max 4s (last 4s kept if longer)
  • GPU: RunPod A40
  • Training time: ~3 hours

Runtime Dependencies

Only two runtime dependencies β€” no PyTorch or transformers needed:

numpy
onnxruntime

Intended Use

This model is a confirmation gate, not a standalone VAD. It is designed to:

  1. Run on-demand when Silero VAD detects brief silence (50-80ms)
  2. Confirm whether the silence marks a true turn end or a mid-utterance pause
  3. Allow min_silence_ms to drop from ~256ms to ~60-80ms for faster turn-taking in voice agents

Credits

License

MIT

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for AswanthCManoj/moonshine-eou-head

Quantized
(12)
this model

Datasets used to train AswanthCManoj/moonshine-eou-head