Instructions to use webmp3/Sakura-sanoTTS-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Piper
How to use webmp3/Sakura-sanoTTS-ONNX with Piper:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Sakura-sanoTTS-ONNX (CPU-Optimized)
- Sakura-sanoTTS-ONNX(针对 CPU 优化)

Part of the Sakura Audio collection.
Sakura-sanoTTS-ONNX (CPU-Optimized)
High-fidelity, CPU-optimized, and reproducible ONNX and INT8 release of sanoTTS (ampixa/sanoTTS).
Target platform: Windows / Linux x86-64 CPU & ARM64, running via the standard ONNX Runtime CPU Execution Provider. No GPU or CUDA required.
1. Which Variant to Choose?
Each of the 15 voices is available in two validated formats:
Choose FP32 if:
- Storage and download size are not critical.
- Maximum reference fidelity is the priority.
- x86-64 CPU performance is the main goal (measured as the fastest execution on x86 CPU).
Choose INT8-HQ if:
- Smaller model and download size matters (up to 60% smaller).
- Bandwidth, disk footprint, or memory is constrained.
- Compact edge or embedded deployment is targeted.
Current speed and RTF measurements are preliminary because the benchmarking system was running under parallel background CPU workloads.
2. Download a Single Voice
You do not need to download the full repository (which contains all 15 voices). Using the official Hugging Face CLI (hf), you can download only the specific voice and runtime files needed:
Example: German Thorsten INT8-HQ (~1.5 MB download)
Linux / macOS:
hf download webmp3/Sakura-sanoTTS-ONNX \
--include "config.json" \
--include "voices/de-thorsten/int8-hq/*" \
--include "runtime/*" \
--include "synthesize.py" \
--include "requirements.txt" \
--local-dir Sakura-sanoTTS
Windows PowerShell:
hf download webmp3/Sakura-sanoTTS-ONNX `
--include "config.json" `
--include "voices/de-thorsten/int8-hq/*" `
--include "runtime/*" `
--include "synthesize.py" `
--include "requirements.txt" `
--local-dir Sakura-sanoTTS
Run synthesis:
cd Sakura-sanoTTS
pip install -r requirements.txt
python synthesize.py `
--model voices/de-thorsten/int8-hq `
--text "Guten Tag, dies ist ein Test." `
--output output.wav
Example: English Amy INT8-HQ (~4.0 MB download)
Linux / macOS:
hf download webmp3/Sakura-sanoTTS-ONNX \
--include "config.json" \
--include "voices/amy/int8-hq/*" \
--include "runtime/*" \
--include "synthesize.py" \
--include "requirements.txt" \
--local-dir Sakura-sanoTTS
Windows PowerShell:
hf download webmp3/Sakura-sanoTTS-ONNX `
--include "config.json" `
--include "voices/amy/int8-hq/*" `
--include "runtime/*" `
--include "synthesize.py" `
--include "requirements.txt" `
--local-dir Sakura-sanoTTS
Tip: Replace
de-thorstenoramywith any voice name from the voice table below. To download an unquantized FP32 voice, replaceint8-hqwithfp32(e.g.--include "voices/amy/fp32/*").
Optional: Download Entire Repository (~114 MB)
To download all 15 voices across both FP32 and INT8-HQ:
hf download webmp3/Sakura-sanoTTS-ONNX --local-dir Sakura-sanoTTS-ONNX
3. Voice Matrix (All 15 Voices)
All 15 official sanoTTS voices have been successfully exported and verified:
| Voice | Language | Sample Rate | FP32 Size | INT8-HQ Size | Recommended Variant | Description |
|---|---|---|---|---|---|---|
amy |
en_US |
22,050 Hz | 5.59 MB | 3.84 MB | FP32 (speed) / INT8-HQ (compact) |
Default English female voice |
amy-1p1m |
en_US |
22,050 Hz | 4.18 MB | 2.87 MB | INT8-HQ (compact) |
Ultra-compact English model (1.1M params) |
amy-1p8m |
en_US |
22,050 Hz | 7.04 MB | 5.29 MB | FP32 (speed) / INT8-HQ (compact) |
High-capacity English voice (1.8M params) |
kristin |
en_US |
22,050 Hz | 5.36 MB | 3.62 MB | FP32 (speed) / INT8-HQ (compact) |
Expressive American English voice |
hfc |
en_US |
22,050 Hz | 7.04 MB | 5.29 MB | FP32 (speed) / INT8-HQ (compact) |
Studio-recorded English voice |
vi |
vi_VN |
22,050 Hz | 6.01 MB | 4.70 MB | FP32 (speed) / INT8-HQ (compact) |
Native Vietnamese tonal voice |
id |
id_ID |
22,050 Hz | 6.00 MB | 4.69 MB | FP32 (speed) / INT8-HQ (compact) |
Native Indonesian voice |
heart-nano |
en_US |
24,000 Hz | 0.91 MB | 0.36 MB | INT8-HQ (ultra-compact) |
Ultra-tiny MCU nano model (294k params) |
heart |
en_US |
24,000 Hz | 8.72 MB | 4.60 MB | FP32 (speed) / INT8-HQ (compact) |
ConvNeXt neural vocoder model (2.27M params) |
de-thorsten |
de_DE |
22,050 Hz | 1.99 MB | 1.39 MB | FP32 (speed) / INT8-HQ (compact) |
German voice (Thorsten corpus) |
es-davefx |
es_ES |
22,050 Hz | 1.98 MB | 1.39 MB | FP32 (speed) / INT8-HQ (compact) |
Castilian Spanish voice |
fr-siwis |
fr_FR |
22,050 Hz | 6.01 MB | 4.70 MB | FP32 (speed) / INT8-HQ (compact) |
French female voice (Siwis corpus) |
it-serena |
it_IT |
22,050 Hz | 1.99 MB | 1.39 MB | FP32 (speed) / INT8-HQ (compact) |
Italian voice |
pt-cadu |
pt_BR |
22,050 Hz | 1.99 MB | 1.39 MB | FP32 (speed) / INT8-HQ (compact) |
Brazilian Portuguese voice |
ru-irina |
ru_RU |
22,050 Hz | 1.99 MB | 1.39 MB | FP32 (speed) / INT8-HQ (compact) |
Russian voice |
4. Python API Usage
from runtime.sanotts_onnx import SanoTTSOnnx
# Load any voice directory (supports FP32 and INT8-HQ)
tts = SanoTTSOnnx("voices/amy/int8-hq")
# Synthesize audio waveform
audio = tts.synthesize("Hello world! This is sanoTTS running locally on ONNX Runtime.")
# Save as standard WAV file
tts.save_wav("output.wav", audio)
print(f"Synthesized {len(audio)} samples ({len(audio)/tts.sample_rate:.2f} s) at {tts.sample_rate} Hz.")
5. Architecture & Quantization
The sanoTTS neural synthesis pipeline is modularly partitioned:
[Text Input]
│
▼ (Python Text Frontend / G2P)
[Phoneme IDs]
│
├───► duration.onnx (Phoneme duration prediction)
│ │
│ ▼
│ [durations]
└───┬────────┘
▼
acoustic_token.onnx (Token-level context embedding)
│
▼
Context Expansion (Repeat per token duration + positional features)
│
▼
acoustic_frame.onnx (Frame-level convolutional blocks)
│
▼
decoder.onnx (Convolutional upsampling + Residual Banks)
│
▼
[Waveform PCM Audio] (22,050 Hz / 24,000 Hz float32)
Quantization Methodology
- Selective Mixed-Precision:
INT8-HQuses selective mixed-precision quantization to preserve sensitive components while reducing overall model size. - Quality-Preserving Precision: Sensitive operations are retained at higher precision where required by quality validation.
- Deterministic Temporal Alignment: The duration modeling preserves exact timing alignment relative to the unquantized reference.
6. Audio Quality Evaluation
Evaluated against the official unquantized sanoTTS reference across standardized test sentences:
| Model Variant | Format | Cosine Similarity | RMS Error | Mel L1 Spectral Diff | Status |
|---|---|---|---|---|---|
| Original Baseline | NumPy Reference | 1.000000 | 0.000000 | 0.0000 | Ground Truth Reference |
| ONNX FP32 | ONNX FP32 | 1.000000 | 0.000018 | 0.0057 | Reference Parity |
| ONNX INT8-HQ | ONNX INT8 | 0.995135 | 0.006614 | 0.6730 | High Fidelity |
- Listening Evaluation: No obvious audible degradation was observed in the current listening tests.
- Timing Parity: Identical sample count and zero timing drift relative to the unquantized baseline.
7. Preliminary CPU Performance
Measurements were performed on standard x86-64 CPU under parallel background system load. Absolute values are preliminary.
| Variant | Model Size | Typical RTF (Preliminary) | Compute Factor | Peak RAM |
|---|---|---|---|---|
| ONNX FP32 | 5.59 MB | ~0.030 | ~33x Real-Time | 222.8 MB |
| ONNX INT8-HQ | 3.84 MB | ~0.260 | ~3.8x Real-Time | 222.2 MB |
Both FP32 and INT8 comfortably achieve real-time speech synthesis on CPU without requiring GPU acceleration.
8. Directory Structure
Sakura-sanoTTS-ONNX/
├── README.md
├── LICENSE
├── NOTICE.md
├── SHA256SUMS
├── requirements.txt
├── synthesize.py
├── runtime/
│ ├── __init__.py
│ └── sanotts_onnx.py
├── docs/
│ ├── VOICE_RELEASE_MATRIX.md
│ ├── QUALITY.md
│ └── SPEED.md
└── voices/
├── amy/ (fp32/ & int8-hq/)
├── amy-1p1m/ (fp32/ & int8-hq/)
├── amy-1p8m/ (fp32/ & int8-hq/)
├── kristin/ (fp32/ & int8-hq/)
├── hfc/ (fp32/ & int8-hq/)
├── vi/ (fp32/ & int8-hq/)
├── id/ (fp32/ & int8-hq/)
├── heart-nano/ (fp32/ & int8-hq/)
├── heart/ (fp32/ & int8-hq/)
├── de-thorsten/ (fp32/ & int8-hq/)
├── es-davefx/ (fp32/ & int8-hq/)
├── fr-siwis/ (fp32/ & int8-hq/)
├── it-serena/ (fp32/ & int8-hq/)
├── pt-cadu/ (fp32/ & int8-hq/)
└── ru-irina/ (fp32/ & int8-hq/)
9. License & Attribution
- Base Model: ampixa/sanoTTS
- Original Authors: Ampixa
- License: GNU General Public License v3.0 (GPL-3.0) (see LICENSE)
- Notice & Attribution: Derivative conversion by Sakura / webmp3 (see NOTICE.md). This project is not an official Ampixa release.
中文说明 · 樱花 (Simplified Chinese)
English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。

Sakura-sanoTTS-ONNX(针对 CPU 优化)
sanoTTS(ampixa/sanoTTS)的高保真、针对 CPU 优化、可复现的 ONNX 和 INT8 发布版本。
目标平台:Windows / Linux x86-64 CPU 和 ARM64,通过标准的 ONNX Runtime CPU 执行提供程序 运行。不需要 GPU 或 CUDA。
1. 该选哪个变体?
15 个声音中的每一个都提供两种经过验证的格式:
在以下情况选择 FP32:
- 存储空间和下载大小不是关键因素。
- 最高的参照保真度是首要目标。
- 主要目标是 x86-64 CPU 上的性能(测得在 x86 CPU 上执行速度最快)。
在以下情况选择 INT8-HQ:
- 较小的模型和下载大小很重要(最多小 60%)。
- 带宽、磁盘占用或内存受限。
- 面向紧凑的边缘或嵌入式部署。
目前的速度和 RTF 测量是初步的,因为基准测试系统当时正在并行的后台 CPU 负载下运行。
2. 下载单个声音
你 不需要下载整个仓库(其中包含全部 15 个声音)。使用官方的 Hugging Face CLI(hf),你可以只下载所需的特定声音和运行时文件:
示例:德语 Thorsten INT8-HQ(约 1.5 MB 的下载)
Linux / macOS:
hf download webmp3/Sakura-sanoTTS-ONNX \
--include "config.json" \
--include "voices/de-thorsten/int8-hq/*" \
--include "runtime/*" \
--include "synthesize.py" \
--include "requirements.txt" \
--local-dir Sakura-sanoTTS
Windows PowerShell:
hf download webmp3/Sakura-sanoTTS-ONNX `
--include "config.json" `
--include "voices/de-thorsten/int8-hq/*" `
--include "runtime/*" `
--include "synthesize.py" `
--include "requirements.txt" `
--local-dir Sakura-sanoTTS
运行合成:
cd Sakura-sanoTTS
pip install -r requirements.txt
python synthesize.py `
--model voices/de-thorsten/int8-hq `
--text "Guten Tag, dies ist ein Test." `
--output output.wav
示例:英语 Amy INT8-HQ(约 4.0 MB 的下载)
Linux / macOS:
hf download webmp3/Sakura-sanoTTS-ONNX \
--include "config.json" \
--include "voices/amy/int8-hq/*" \
--include "runtime/*" \
--include "synthesize.py" \
--include "requirements.txt" \
--local-dir Sakura-sanoTTS
Windows PowerShell:
hf download webmp3/Sakura-sanoTTS-ONNX `
--include "config.json" `
--include "voices/amy/int8-hq/*" `
--include "runtime/*" `
--include "synthesize.py" `
--include "requirements.txt" `
--local-dir Sakura-sanoTTS
提示: 将
de-thorsten或amy替换为下面声音表中的任意声音名称。若要下载未量化的 FP32 声音,请把int8-hq替换为fp32(例如--include "voices/amy/fp32/*")。
可选:下载整个仓库(约 114 MB)
下载全部 15 个声音的 FP32 和 INT8-HQ 两种格式:
hf download webmp3/Sakura-sanoTTS-ONNX --local-dir Sakura-sanoTTS-ONNX
3. 声音一览(全部 15 个声音)
sanoTTS 官方的全部 15 个声音都已成功导出并验证:
| 声音 | 语言 | 采样率 | FP32 大小 | INT8-HQ 大小 | 推荐变体 | 说明 |
|---|---|---|---|---|---|---|
amy |
en_US |
22,050 Hz | 5.59 MB | 3.84 MB | FP32 (速度) / INT8-HQ (紧凑) |
默认的英语女声 |
amy-1p1m |
en_US |
22,050 Hz | 4.18 MB | 2.87 MB | INT8-HQ (紧凑) |
极其紧凑的英语模型(1.1M 参数) |
amy-1p8m |
en_US |
22,050 Hz | 7.04 MB | 5.29 MB | FP32 (速度) / INT8-HQ (紧凑) |
高容量的英语声音(1.8M 参数) |
kristin |
en_US |
22,050 Hz | 5.36 MB | 3.62 MB | FP32 (速度) / INT8-HQ (紧凑) |
富有表现力的美式英语声音 |
hfc |
en_US |
22,050 Hz | 7.04 MB | 5.29 MB | FP32 (速度) / INT8-HQ (紧凑) |
录音棚录制的英语声音 |
vi |
vi_VN |
22,050 Hz | 6.01 MB | 4.70 MB | FP32 (速度) / INT8-HQ (紧凑) |
原生越南语声调声音 |
id |
id_ID |
22,050 Hz | 6.00 MB | 4.69 MB | FP32 (速度) / INT8-HQ (紧凑) |
原生印度尼西亚语声音 |
heart-nano |
en_US |
24,000 Hz | 0.91 MB | 0.36 MB | INT8-HQ (超紧凑) |
超小的 MCU 纳米模型(294k 参数) |
heart |
en_US |
24,000 Hz | 8.72 MB | 4.60 MB | FP32 (速度) / INT8-HQ (紧凑) |
ConvNeXt 神经声码器模型(2.27M 参数) |
de-thorsten |
de_DE |
22,050 Hz | 1.99 MB | 1.39 MB | FP32 (速度) / INT8-HQ (紧凑) |
德语声音(Thorsten 语料库) |
es-davefx |
es_ES |
22,050 Hz | 1.98 MB | 1.39 MB | FP32 (速度) / INT8-HQ (紧凑) |
卡斯蒂利亚西班牙语声音 |
fr-siwis |
fr_FR |
22,050 Hz | 6.01 MB | 4.70 MB | FP32 (速度) / INT8-HQ (紧凑) |
法语女声(Siwis 语料库) |
it-serena |
it_IT |
22,050 Hz | 1.99 MB | 1.39 MB | FP32 (速度) / INT8-HQ (紧凑) |
意大利语声音 |
pt-cadu |
pt_BR |
22,050 Hz | 1.99 MB | 1.39 MB | FP32 (速度) / INT8-HQ (紧凑) |
巴西葡萄牙语声音 |
ru-irina |
ru_RU |
22,050 Hz | 1.99 MB | 1.39 MB | FP32 (速度) / INT8-HQ (紧凑) |
俄语声音 |
4. Python API 用法
from runtime.sanotts_onnx import SanoTTSOnnx
# Load any voice directory (supports FP32 and INT8-HQ)
tts = SanoTTSOnnx("voices/amy/int8-hq")
# Synthesize audio waveform
audio = tts.synthesize("Hello world! This is sanoTTS running locally on ONNX Runtime.")
# Save as standard WAV file
tts.save_wav("output.wav", audio)
print(f"Synthesized {len(audio)} samples ({len(audio)/tts.sample_rate:.2f} s) at {tts.sample_rate} Hz.")
5. 架构与量化
sanoTTS 的神经合成流水线被模块化地划分:
[Text Input]
│
▼ (Python Text Frontend / G2P)
[Phoneme IDs]
│
├───► duration.onnx (Phoneme duration prediction)
│ │
│ ▼
│ [durations]
└───┬────────┘
▼
acoustic_token.onnx (Token-level context embedding)
│
▼
Context Expansion (Repeat per token duration + positional features)
│
▼
acoustic_frame.onnx (Frame-level convolutional blocks)
│
▼
decoder.onnx (Convolutional upsampling + Residual Banks)
│
▼
[Waveform PCM Audio] (22,050 Hz / 24,000 Hz float32)
量化方法
- 选择性混合精度:
INT8-HQ使用选择性混合精度量化,在减小整体模型大小的同时保护敏感组件。 - 保持质量的精度:在质量验证要求的地方,敏感运算保持较高精度。
- 确定性的时间对齐:时长建模相对于未量化的参照保持精确的时间对齐。
6. 音频质量评估
针对标准化测试句子,与官方未量化的 sanoTTS 参照进行评估:
| 模型变体 | 格式 | 余弦相似度 | RMS 误差 | Mel L1 频谱差异 | 状态 |
|---|---|---|---|---|---|
| 原始基线 | NumPy 参照 | 1.000000 | 0.000000 | 0.0000 | 标准答案参照 |
| ONNX FP32 | ONNX FP32 | 1.000000 | 0.000018 | 0.0057 | 与参照一致 |
| ONNX INT8-HQ | ONNX INT8 | 0.995135 | 0.006614 | 0.6730 | 高保真 |
- 听感评估:在目前的听感测试中,没有观察到明显可听见的劣化。
- 时间一致性:相对于未量化的基线,采样点数量相同,且零时间漂移。
7. 初步的 CPU 性能
测量是在标准 x86-64 CPU 上、并行的后台系统负载下进行的。绝对数值是初步的。
| 变体 | 模型大小 | 典型 RTF(初步) | 计算系数 | 峰值 RAM |
|---|---|---|---|---|
| ONNX FP32 | 5.59 MB | ~0.030 | ~33x 实时 | 222.8 MB |
| ONNX INT8-HQ | 3.84 MB | ~0.260 | ~3.8x 实时 | 222.2 MB |
FP32 和 INT8 都可以在 CPU 上轻松实现实时语音合成,无需 GPU 加速。
8. 目录结构
Sakura-sanoTTS-ONNX/
├── README.md
├── LICENSE
├── NOTICE.md
├── SHA256SUMS
├── requirements.txt
├── synthesize.py
├── runtime/
│ ├── __init__.py
│ └── sanotts_onnx.py
├── docs/
│ ├── VOICE_RELEASE_MATRIX.md
│ ├── QUALITY.md
│ └── SPEED.md
└── voices/
├── amy/ (fp32/ & int8-hq/)
├── amy-1p1m/ (fp32/ & int8-hq/)
├── amy-1p8m/ (fp32/ & int8-hq/)
├── kristin/ (fp32/ & int8-hq/)
├── hfc/ (fp32/ & int8-hq/)
├── vi/ (fp32/ & int8-hq/)
├── id/ (fp32/ & int8-hq/)
├── heart-nano/ (fp32/ & int8-hq/)
├── heart/ (fp32/ & int8-hq/)
├── de-thorsten/ (fp32/ & int8-hq/)
├── es-davefx/ (fp32/ & int8-hq/)
├── fr-siwis/ (fp32/ & int8-hq/)
├── it-serena/ (fp32/ & int8-hq/)
├── pt-cadu/ (fp32/ & int8-hq/)
└── ru-irina/ (fp32/ & int8-hq/)
9. 许可证与归属
- 基础模型:ampixa/sanoTTS
- 原作者:Ampixa
- 许可证:GNU General Public License v3.0 (GPL-3.0)(见 LICENSE)
- 声明与归属:由 Sakura / webmp3 完成的衍生转换(见 NOTICE.md)。本项目不是 Ampixa 的官方发布。
- Downloads last month
- 29
Model tree for webmp3/Sakura-sanoTTS-ONNX
Base model
ampixa/sanoTTS