RootDeveloperDS's picture
Publish Visar wake-word models: ONNX, TFLite, and fine-tuning state
f2215dc verified
|
Raw History Blame Contribute Delete
2.03 kB
---
language:
- hi
- en
tags:
- openwakeword
- wake-word-detection
- voice-assistant
- edge-ai
- onnx
- tflite
license: apache-2.0
---
# Visar (वी-सार) — Custom openWakeWord Edge Model
A lightweight, high-accuracy wake-word detection model custom-trained for low-power edge SBCs, PCs, and offline home automation systems.
## Model Highlights
- **Target Phrase:** "Visar" / "वी-सार" / "Vee-saar"
- **Acoustic Tuning:** Native Hindi cadence and Indian English phonetic variations (`hi-IN-Madhur`, `hi-IN-Swara`, `en-IN-Prabhat`, `en-IN-Neerja`).
- **Ground Truth:** Real-world microphone samples convolved with MIT environmental impulse responses (RIRs).
- **Hard Negatives:** Penalized against acoustically close Indian words (*vichar*, *vishal*, *vikas*, *bazaar*).
- **Target Hardware:** Raspberry Pi, Linux Edge SBCs, Intel x86, Android, ARM64 microcontrollers.
## Audio Input Requirements
- **Sample Rate:** 16,000 Hz
- **Channels:** 1 (Mono)
- **Format:** 16-bit Signed Linear PCM
## Repository Contents
| File / Directory | Description |
| :--- | :--- |
| `visar_edge.onnx` | Production ONNX inference graph |
| `visar_edge.onnx.data` | External weight buffer for ONNX runtime |
| `visar_edge.tflite` | Quantized FlatBuffer model for mobile/embedded devices |
| `training_config.yaml` | Exact hyperparameter configuration used during training |
| `training_state/` | Step checkpoints and optimizer states for fine-tuning |
| `user_calibration_audio/` | Microphone ground-truth calibration recordings |
## Quickstart (Python Inference)
```python
import numpy as np
import openwakeword
from openwakeword.model import Model
# Initialize wake word engine with custom Visar model
oww_model = Model(wakeword_models=["visar_edge.onnx"])
# Pass raw 16kHz 16-bit PCM audio frames (1280 samples / 80ms chunk)
# audio_frame = np.frombuffer(mic_stream.read(1280), dtype=np.int16)
# prediction = oww_model.predict(audio_frame)
# if prediction["visar_edge"] > 0.5:
# print("Wake-word detected: Visar!")