File size: 2,029 Bytes
f2215dc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
---
language:
- hi
- en
tags:
- openwakeword
- wake-word-detection
- voice-assistant
- edge-ai
- onnx
- tflite
license: apache-2.0
---

# Visar (वी-सार) — Custom openWakeWord Edge Model

A lightweight, high-accuracy wake-word detection model custom-trained for low-power edge SBCs, PCs, and offline home automation systems.

## Model Highlights
- **Target Phrase:** "Visar" / "वी-सार" / "Vee-saar"
- **Acoustic Tuning:** Native Hindi cadence and Indian English phonetic variations (`hi-IN-Madhur`, `hi-IN-Swara`, `en-IN-Prabhat`, `en-IN-Neerja`).
- **Ground Truth:** Real-world microphone samples convolved with MIT environmental impulse responses (RIRs).
- **Hard Negatives:** Penalized against acoustically close Indian words (*vichar*, *vishal*, *vikas*, *bazaar*).
- **Target Hardware:** Raspberry Pi, Linux Edge SBCs, Intel x86, Android, ARM64 microcontrollers.

## Audio Input Requirements
- **Sample Rate:** 16,000 Hz
- **Channels:** 1 (Mono)
- **Format:** 16-bit Signed Linear PCM

## Repository Contents
| File / Directory | Description |
| :--- | :--- |
| `visar_edge.onnx` | Production ONNX inference graph |
| `visar_edge.onnx.data` | External weight buffer for ONNX runtime |
| `visar_edge.tflite` | Quantized FlatBuffer model for mobile/embedded devices |
| `training_config.yaml` | Exact hyperparameter configuration used during training |
| `training_state/` | Step checkpoints and optimizer states for fine-tuning |
| `user_calibration_audio/` | Microphone ground-truth calibration recordings |

## Quickstart (Python Inference)
```python
import numpy as np
import openwakeword
from openwakeword.model import Model

# Initialize wake word engine with custom Visar model
oww_model = Model(wakeword_models=["visar_edge.onnx"])

# Pass raw 16kHz 16-bit PCM audio frames (1280 samples / 80ms chunk)
# audio_frame = np.frombuffer(mic_stream.read(1280), dtype=np.int16)
# prediction = oww_model.predict(audio_frame)
# if prediction["visar_edge"] > 0.5:
#     print("Wake-word detected: Visar!")