OneVoice / Vega β on-device speech translation model bundle
The full ONNX model bundle for QUAVIO/Vega, a 100% offline (no-cloud)
speech-to-speech translator for noisy industrial environments, built for the
Saigon AI Hub Γ Qualcomm "OneVoice AI Challenge". One cascaded engine
(denoise β VAD β ASR β MT β TTS) shipped on the Qualcomm Dragonwing
QCS6490 (Rubik Pi 3 / RB3 Gen 2), in two carriers:
- Vega Hand β ruggedized push-to-talk handheld (one-to-one).
- Vega Air β drone-mounted multilingual broadcast payload (one-to-many).
Languages: Vietnamese β English / Chinese (Mandarin) / Korean (6 directions).
This repo holds the .onnx artifacts (and Piper's espeak-ng-data phoneme
tables) that the Rust runtime (src/) loads via ONNX Runtime / QNN. It is the
same tree scripts/pack_models.sh ships as the GitHub Release bundle and that
scripts/setup_models.sh installs.
Models
| Stage | Model | Size | License | Source |
|---|---|---|---|---|
| VAD | Silero VAD | 2.1 MB | MIT | onnx-community/silero-vad |
| ASR (VI) | PhoWhisper-base (int8) | 22 MB enc / 51 MB dec | BSD-3-Clause | huuquyet/PhoWhisper-base |
| ASR (EN/ZH/KO) | Whisper-small (int8) | 88 MB enc / 149 MB dec | MIT | onnx-community/whisper-small |
| MT (VIβEN) | Opus-MT (INT8) | 45 MB enc / 78 MB dec | Apache-2.0 / CC-BY | Helsinki-NLP/opus-mt-* |
| TTS (VI) | Piper vais1000-medium (VITS) |
60 MB | CC BY 4.0 | rhasspy/piper-voices |
| TTS (EN) | Piper lessac-medium (VITS) |
60 MB | CC BY 4.0 | rhasspy/piper-voices |
| TTS (ZH) | Piper huayan-medium (VITS) |
60 MB | CC BY 4.0 | rhasspy/piper-voices |
| TTS (KO) | MMS-TTS Korean (VITS) | 109 MB | MIT | facebook/mms-tts-kor |
| TTS runtime data | espeak-ng-data (Piper phonemes) |
19 MB | GPL-3.0 | piper-tts PyPI package |
Denoise (RNNoise) has no model file β its weights are compiled into the Rust
nnnoiseless crate. Encoders/decoders use the merged-decoder KV-cache ONNX
layout (decoder_model_merged.onnx) for autoregressive greedy decode.
Repository structure
βββ vad/model.onnx
βββ asr/vi/{encoder_model.onnx, decoder_model_merged.onnx, tokenizer.json, ...}
βββ asr/en-zh-ko/{...}
βββ mt/vi-en/{encoder_model.onnx, decoder_model_merged.onnx, tokenizer.json, ...}
βββ mt/en-vi/{...}
βββ tts/{vi,en,zh,ko}/*.onnx (+ Piper *.onnx.json sidecars)
βββ tts/espeak-ng-data/ # Piper phoneme data (vi/en/zh)
Each directory carries a MANIFEST.md recording source checkpoint, export
date, quantization scheme, and eval metrics.
Usage
git lfs install
git clone https://huggingface.co/LakoreAI/onevoice-vega models/
# run the full Rust pipeline (denoise β VAD β ASR β MT β TTS)
# see the project repo: https://github.com/MinLee0210/onevoice-vega
scripts/run_pipeline.sh data/eval/ptt_sample_vi_48k.wav out.wav -- --source vi --target en
Or download only what setup_models.sh needs from the project repo β it
points at this bundle as the "rebuild" source for vad/asr/tts.
Performance
Measured on the QCS6490 (Hexagon v68 HTP), full-INT8 QNN context binaries β
see the companion qualcomm-ai-hub-community/whisper-*-qcs6490-qnn-int8 repos:
| Stage | Latency |
|---|---|
| ASR encoder (PhoWhisper-base, 30 s) | ~49.5 ms (NPU) |
| ASR decoder (per token) | ~12.1 ms (NPU) |
| End-to-end pipeline (CPU dev machine) | ~741 ms button-release β audio |
License note
This is a mixed-license bundle β each stage carries its own license (table
above). In particular, espeak-ng-data/ is GPL-3.0 (sourced from the
piper-tts package); embedding it in a distributed binary raises an
unresolved GPL question, relevant only to commercial distribution, not to the
prototype/demo. See docs/decisions/piper-vi-voice-license-trap.md in the
project repo for the full analysis. The Korean TTS path (MMS-TTS, MIT) avoids
espeak entirely.
Attribution
- Whisper β OpenAI (MIT); PhoWhisper β VinAI / huuquyet.
- Opus-MT β Helsinki-NLP (Apache-2.0).
- Silero VAD β Silero Team (MIT).
- Piper voices β rhasspy/piper-voices; espeak-ng data from piper-tts.
- MMS-TTS β Meta (MIT).
Project: MinLee0210/onevoice-vega.
Model tree for LakoreAI/onevoice-vega
Base model
Helsinki-NLP/opus-mt-en-viEvaluation results
- WER (clean speech) on VIVOSself-reported2.200
- glossary term accuracy on OneVoice industrial glossary carrier setself-reported100.000