Sakura Audio

Part of the Sakura Audio collection.

Sakura-sanoTTS-ONNX (CPU-Optimized)

Hugging Face Model License: GPL v3 ONNX Runtime

High-fidelity, CPU-optimized, and reproducible ONNX and INT8 release of sanoTTS (ampixa/sanoTTS).

Target platform: Windows / Linux x86-64 CPU & ARM64, running via the standard ONNX Runtime CPU Execution Provider. No GPU or CUDA required.


1. Which Variant to Choose?

Each of the 15 voices is available in two validated formats:

Choose FP32 if:

  • Storage and download size are not critical.
  • Maximum reference fidelity is the priority.
  • x86-64 CPU performance is the main goal (measured as the fastest execution on x86 CPU).

Choose INT8-HQ if:

  • Smaller model and download size matters (up to 60% smaller).
  • Bandwidth, disk footprint, or memory is constrained.
  • Compact edge or embedded deployment is targeted.

Current speed and RTF measurements are preliminary because the benchmarking system was running under parallel background CPU workloads.


2. Download a Single Voice

You do not need to download the full repository (which contains all 15 voices). Using the official Hugging Face CLI (hf), you can download only the specific voice and runtime files needed:

Example: German Thorsten INT8-HQ (~1.5 MB download)

Linux / macOS:

hf download webmp3/Sakura-sanoTTS-ONNX \
  --include "config.json" \
  --include "voices/de-thorsten/int8-hq/*" \
  --include "runtime/*" \
  --include "synthesize.py" \
  --include "requirements.txt" \
  --local-dir Sakura-sanoTTS

Windows PowerShell:

hf download webmp3/Sakura-sanoTTS-ONNX `
  --include "config.json" `
  --include "voices/de-thorsten/int8-hq/*" `
  --include "runtime/*" `
  --include "synthesize.py" `
  --include "requirements.txt" `
  --local-dir Sakura-sanoTTS

Run synthesis:

cd Sakura-sanoTTS
pip install -r requirements.txt

python synthesize.py `
  --model voices/de-thorsten/int8-hq `
  --text "Guten Tag, dies ist ein Test." `
  --output output.wav

Example: English Amy INT8-HQ (~4.0 MB download)

Linux / macOS:

hf download webmp3/Sakura-sanoTTS-ONNX \
  --include "config.json" \
  --include "voices/amy/int8-hq/*" \
  --include "runtime/*" \
  --include "synthesize.py" \
  --include "requirements.txt" \
  --local-dir Sakura-sanoTTS

Windows PowerShell:

hf download webmp3/Sakura-sanoTTS-ONNX `
  --include "config.json" `
  --include "voices/amy/int8-hq/*" `
  --include "runtime/*" `
  --include "synthesize.py" `
  --include "requirements.txt" `
  --local-dir Sakura-sanoTTS

Tip: Replace de-thorsten or amy with any voice name from the voice table below. To download an unquantized FP32 voice, replace int8-hq with fp32 (e.g. --include "voices/amy/fp32/*").


Optional: Download Entire Repository (~114 MB)

To download all 15 voices across both FP32 and INT8-HQ:

hf download webmp3/Sakura-sanoTTS-ONNX --local-dir Sakura-sanoTTS-ONNX

3. Voice Matrix (All 15 Voices)

All 15 official sanoTTS voices have been successfully exported and verified:

Voice Language Sample Rate FP32 Size INT8-HQ Size Recommended Variant Description
amy en_US 22,050 Hz 5.59 MB 3.84 MB FP32 (speed) / INT8-HQ (compact) Default English female voice
amy-1p1m en_US 22,050 Hz 4.18 MB 2.87 MB INT8-HQ (compact) Ultra-compact English model (1.1M params)
amy-1p8m en_US 22,050 Hz 7.04 MB 5.29 MB FP32 (speed) / INT8-HQ (compact) High-capacity English voice (1.8M params)
kristin en_US 22,050 Hz 5.36 MB 3.62 MB FP32 (speed) / INT8-HQ (compact) Expressive American English voice
hfc en_US 22,050 Hz 7.04 MB 5.29 MB FP32 (speed) / INT8-HQ (compact) Studio-recorded English voice
vi vi_VN 22,050 Hz 6.01 MB 4.70 MB FP32 (speed) / INT8-HQ (compact) Native Vietnamese tonal voice
id id_ID 22,050 Hz 6.00 MB 4.69 MB FP32 (speed) / INT8-HQ (compact) Native Indonesian voice
heart-nano en_US 24,000 Hz 0.91 MB 0.36 MB INT8-HQ (ultra-compact) Ultra-tiny MCU nano model (294k params)
heart en_US 24,000 Hz 8.72 MB 4.60 MB FP32 (speed) / INT8-HQ (compact) ConvNeXt neural vocoder model (2.27M params)
de-thorsten de_DE 22,050 Hz 1.99 MB 1.39 MB FP32 (speed) / INT8-HQ (compact) German voice (Thorsten corpus)
es-davefx es_ES 22,050 Hz 1.98 MB 1.39 MB FP32 (speed) / INT8-HQ (compact) Castilian Spanish voice
fr-siwis fr_FR 22,050 Hz 6.01 MB 4.70 MB FP32 (speed) / INT8-HQ (compact) French female voice (Siwis corpus)
it-serena it_IT 22,050 Hz 1.99 MB 1.39 MB FP32 (speed) / INT8-HQ (compact) Italian voice
pt-cadu pt_BR 22,050 Hz 1.99 MB 1.39 MB FP32 (speed) / INT8-HQ (compact) Brazilian Portuguese voice
ru-irina ru_RU 22,050 Hz 1.99 MB 1.39 MB FP32 (speed) / INT8-HQ (compact) Russian voice

4. Python API Usage

from runtime.sanotts_onnx import SanoTTSOnnx

# Load any voice directory (supports FP32 and INT8-HQ)
tts = SanoTTSOnnx("voices/amy/int8-hq")

# Synthesize audio waveform
audio = tts.synthesize("Hello world! This is sanoTTS running locally on ONNX Runtime.")

# Save as standard WAV file
tts.save_wav("output.wav", audio)
print(f"Synthesized {len(audio)} samples ({len(audio)/tts.sample_rate:.2f} s) at {tts.sample_rate} Hz.")

5. Architecture & Quantization

The sanoTTS neural synthesis pipeline is modularly partitioned:

[Text Input]
     │
     ▼ (Python Text Frontend / G2P)
[Phoneme IDs]
     │
     ├───► duration.onnx       (Phoneme duration prediction)
     │            │
     │            ▼
     │       [durations]
     └───┬────────┘
         ▼
      acoustic_token.onnx      (Token-level context embedding)
         │
         ▼
      Context Expansion        (Repeat per token duration + positional features)
         │
         ▼
      acoustic_frame.onnx      (Frame-level convolutional blocks)
         │
         ▼
      decoder.onnx             (Convolutional upsampling + Residual Banks)
         │
         ▼
      [Waveform PCM Audio]     (22,050 Hz / 24,000 Hz float32)

Quantization Methodology

  • Selective Mixed-Precision: INT8-HQ uses selective mixed-precision quantization to preserve sensitive components while reducing overall model size.
  • Quality-Preserving Precision: Sensitive operations are retained at higher precision where required by quality validation.
  • Deterministic Temporal Alignment: The duration modeling preserves exact timing alignment relative to the unquantized reference.

6. Audio Quality Evaluation

Evaluated against the official unquantized sanoTTS reference across standardized test sentences:

Model Variant Format Cosine Similarity RMS Error Mel L1 Spectral Diff Status
Original Baseline NumPy Reference 1.000000 0.000000 0.0000 Ground Truth Reference
ONNX FP32 ONNX FP32 1.000000 0.000018 0.0057 Reference Parity
ONNX INT8-HQ ONNX INT8 0.995135 0.006614 0.6730 High Fidelity
  • Listening Evaluation: No obvious audible degradation was observed in the current listening tests.
  • Timing Parity: Identical sample count and zero timing drift relative to the unquantized baseline.

7. Preliminary CPU Performance

Measurements were performed on standard x86-64 CPU under parallel background system load. Absolute values are preliminary.

Variant Model Size Typical RTF (Preliminary) Compute Factor Peak RAM
ONNX FP32 5.59 MB ~0.030 ~33x Real-Time 222.8 MB
ONNX INT8-HQ 3.84 MB ~0.260 ~3.8x Real-Time 222.2 MB

Both FP32 and INT8 comfortably achieve real-time speech synthesis on CPU without requiring GPU acceleration.


8. Directory Structure

Sakura-sanoTTS-ONNX/
├── README.md
├── LICENSE
├── NOTICE.md
├── SHA256SUMS
├── requirements.txt
├── synthesize.py
├── runtime/
│   ├── __init__.py
│   └── sanotts_onnx.py
├── docs/
│   ├── VOICE_RELEASE_MATRIX.md
│   ├── QUALITY.md
│   └── SPEED.md
└── voices/
    ├── amy/            (fp32/ & int8-hq/)
    ├── amy-1p1m/       (fp32/ & int8-hq/)
    ├── amy-1p8m/       (fp32/ & int8-hq/)
    ├── kristin/        (fp32/ & int8-hq/)
    ├── hfc/            (fp32/ & int8-hq/)
    ├── vi/             (fp32/ & int8-hq/)
    ├── id/             (fp32/ & int8-hq/)
    ├── heart-nano/     (fp32/ & int8-hq/)
    ├── heart/          (fp32/ & int8-hq/)
    ├── de-thorsten/    (fp32/ & int8-hq/)
    ├── es-davefx/      (fp32/ & int8-hq/)
    ├── fr-siwis/       (fp32/ & int8-hq/)
    ├── it-serena/      (fp32/ & int8-hq/)
    ├── pt-cadu/        (fp32/ & int8-hq/)
    └── ru-irina/       (fp32/ & int8-hq/)

9. License & Attribution

  • Base Model: ampixa/sanoTTS
  • Original Authors: Ampixa
  • License: GNU General Public License v3.0 (GPL-3.0) (see LICENSE)
  • Notice & Attribution: Derivative conversion by Sakura / webmp3 (see NOTICE.md). This project is not an official Ampixa release.

中文说明 · 樱花 (Simplified Chinese)

English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。

Sakura Audio

属于 Sakura Audio 合集。

Sakura-sanoTTS-ONNX(针对 CPU 优化)

Hugging Face Model License: GPL v3 ONNX Runtime

sanoTTS(ampixa/sanoTTS)的高保真、针对 CPU 优化、可复现的 ONNX 和 INT8 发布版本。

目标平台:Windows / Linux x86-64 CPU 和 ARM64,通过标准的 ONNX Runtime CPU 执行提供程序 运行。不需要 GPU 或 CUDA。


1. 该选哪个变体?

15 个声音中的每一个都提供两种经过验证的格式:

在以下情况选择 FP32:

  • 存储空间和下载大小不是关键因素。
  • 最高的参照保真度是首要目标。
  • 主要目标是 x86-64 CPU 上的性能(测得在 x86 CPU 上执行速度最快)。

在以下情况选择 INT8-HQ:

  • 较小的模型和下载大小很重要(最多小 60%)。
  • 带宽、磁盘占用或内存受限。
  • 面向紧凑的边缘或嵌入式部署。

目前的速度和 RTF 测量是初步的,因为基准测试系统当时正在并行的后台 CPU 负载下运行。


2. 下载单个声音

你 不需要下载整个仓库(其中包含全部 15 个声音)。使用官方的 Hugging Face CLI(hf),你可以只下载所需的特定声音和运行时文件:

示例:德语 Thorsten INT8-HQ(约 1.5 MB 的下载)

Linux / macOS:

hf download webmp3/Sakura-sanoTTS-ONNX \
  --include "config.json" \
  --include "voices/de-thorsten/int8-hq/*" \
  --include "runtime/*" \
  --include "synthesize.py" \
  --include "requirements.txt" \
  --local-dir Sakura-sanoTTS

Windows PowerShell:

hf download webmp3/Sakura-sanoTTS-ONNX `
  --include "config.json" `
  --include "voices/de-thorsten/int8-hq/*" `
  --include "runtime/*" `
  --include "synthesize.py" `
  --include "requirements.txt" `
  --local-dir Sakura-sanoTTS

运行合成:

cd Sakura-sanoTTS
pip install -r requirements.txt

python synthesize.py `
  --model voices/de-thorsten/int8-hq `
  --text "Guten Tag, dies ist ein Test." `
  --output output.wav

示例:英语 Amy INT8-HQ(约 4.0 MB 的下载)

Linux / macOS:

hf download webmp3/Sakura-sanoTTS-ONNX \
  --include "config.json" \
  --include "voices/amy/int8-hq/*" \
  --include "runtime/*" \
  --include "synthesize.py" \
  --include "requirements.txt" \
  --local-dir Sakura-sanoTTS

Windows PowerShell:

hf download webmp3/Sakura-sanoTTS-ONNX `
  --include "config.json" `
  --include "voices/amy/int8-hq/*" `
  --include "runtime/*" `
  --include "synthesize.py" `
  --include "requirements.txt" `
  --local-dir Sakura-sanoTTS

提示: 将 de-thorsten 或 amy 替换为下面声音表中的任意声音名称。若要下载未量化的 FP32 声音,请把 int8-hq 替换为 fp32(例如 --include "voices/amy/fp32/*")。


可选:下载整个仓库(约 114 MB)

下载全部 15 个声音的 FP32 和 INT8-HQ 两种格式:

hf download webmp3/Sakura-sanoTTS-ONNX --local-dir Sakura-sanoTTS-ONNX

3. 声音一览(全部 15 个声音)

sanoTTS 官方的全部 15 个声音都已成功导出并验证:

声音 语言 采样率 FP32 大小 INT8-HQ 大小 推荐变体 说明
amy en_US 22,050 Hz 5.59 MB 3.84 MB FP32 (速度) / INT8-HQ (紧凑) 默认的英语女声
amy-1p1m en_US 22,050 Hz 4.18 MB 2.87 MB INT8-HQ (紧凑) 极其紧凑的英语模型(1.1M 参数)
amy-1p8m en_US 22,050 Hz 7.04 MB 5.29 MB FP32 (速度) / INT8-HQ (紧凑) 高容量的英语声音(1.8M 参数)
kristin en_US 22,050 Hz 5.36 MB 3.62 MB FP32 (速度) / INT8-HQ (紧凑) 富有表现力的美式英语声音
hfc en_US 22,050 Hz 7.04 MB 5.29 MB FP32 (速度) / INT8-HQ (紧凑) 录音棚录制的英语声音
vi vi_VN 22,050 Hz 6.01 MB 4.70 MB FP32 (速度) / INT8-HQ (紧凑) 原生越南语声调声音
id id_ID 22,050 Hz 6.00 MB 4.69 MB FP32 (速度) / INT8-HQ (紧凑) 原生印度尼西亚语声音
heart-nano en_US 24,000 Hz 0.91 MB 0.36 MB INT8-HQ (超紧凑) 超小的 MCU 纳米模型(294k 参数)
heart en_US 24,000 Hz 8.72 MB 4.60 MB FP32 (速度) / INT8-HQ (紧凑) ConvNeXt 神经声码器模型(2.27M 参数)
de-thorsten de_DE 22,050 Hz 1.99 MB 1.39 MB FP32 (速度) / INT8-HQ (紧凑) 德语声音(Thorsten 语料库)
es-davefx es_ES 22,050 Hz 1.98 MB 1.39 MB FP32 (速度) / INT8-HQ (紧凑) 卡斯蒂利亚西班牙语声音
fr-siwis fr_FR 22,050 Hz 6.01 MB 4.70 MB FP32 (速度) / INT8-HQ (紧凑) 法语女声(Siwis 语料库)
it-serena it_IT 22,050 Hz 1.99 MB 1.39 MB FP32 (速度) / INT8-HQ (紧凑) 意大利语声音
pt-cadu pt_BR 22,050 Hz 1.99 MB 1.39 MB FP32 (速度) / INT8-HQ (紧凑) 巴西葡萄牙语声音
ru-irina ru_RU 22,050 Hz 1.99 MB 1.39 MB FP32 (速度) / INT8-HQ (紧凑) 俄语声音

4. Python API 用法

from runtime.sanotts_onnx import SanoTTSOnnx

# Load any voice directory (supports FP32 and INT8-HQ)
tts = SanoTTSOnnx("voices/amy/int8-hq")

# Synthesize audio waveform
audio = tts.synthesize("Hello world! This is sanoTTS running locally on ONNX Runtime.")

# Save as standard WAV file
tts.save_wav("output.wav", audio)
print(f"Synthesized {len(audio)} samples ({len(audio)/tts.sample_rate:.2f} s) at {tts.sample_rate} Hz.")

5. 架构与量化

sanoTTS 的神经合成流水线被模块化地划分:

[Text Input]
     │
     ▼ (Python Text Frontend / G2P)
[Phoneme IDs]
     │
     ├───► duration.onnx       (Phoneme duration prediction)
     │            │
     │            ▼
     │       [durations]
     └───┬────────┘
         ▼
      acoustic_token.onnx      (Token-level context embedding)
         │
         ▼
      Context Expansion        (Repeat per token duration + positional features)
         │
         ▼
      acoustic_frame.onnx      (Frame-level convolutional blocks)
         │
         ▼
      decoder.onnx             (Convolutional upsampling + Residual Banks)
         │
         ▼
      [Waveform PCM Audio]     (22,050 Hz / 24,000 Hz float32)

量化方法

  • 选择性混合精度:INT8-HQ 使用选择性混合精度量化,在减小整体模型大小的同时保护敏感组件。
  • 保持质量的精度:在质量验证要求的地方,敏感运算保持较高精度。
  • 确定性的时间对齐:时长建模相对于未量化的参照保持精确的时间对齐。

6. 音频质量评估

针对标准化测试句子,与官方未量化的 sanoTTS 参照进行评估:

模型变体 格式 余弦相似度 RMS 误差 Mel L1 频谱差异 状态
原始基线 NumPy 参照 1.000000 0.000000 0.0000 标准答案参照
ONNX FP32 ONNX FP32 1.000000 0.000018 0.0057 与参照一致
ONNX INT8-HQ ONNX INT8 0.995135 0.006614 0.6730 高保真
  • 听感评估:在目前的听感测试中,没有观察到明显可听见的劣化。
  • 时间一致性:相对于未量化的基线,采样点数量相同,且零时间漂移。

7. 初步的 CPU 性能

测量是在标准 x86-64 CPU 上、并行的后台系统负载下进行的。绝对数值是初步的。

变体 模型大小 典型 RTF(初步) 计算系数 峰值 RAM
ONNX FP32 5.59 MB ~0.030 ~33x 实时 222.8 MB
ONNX INT8-HQ 3.84 MB ~0.260 ~3.8x 实时 222.2 MB

FP32 和 INT8 都可以在 CPU 上轻松实现实时语音合成,无需 GPU 加速。


8. 目录结构

Sakura-sanoTTS-ONNX/
├── README.md
├── LICENSE
├── NOTICE.md
├── SHA256SUMS
├── requirements.txt
├── synthesize.py
├── runtime/
│   ├── __init__.py
│   └── sanotts_onnx.py
├── docs/
│   ├── VOICE_RELEASE_MATRIX.md
│   ├── QUALITY.md
│   └── SPEED.md
└── voices/
    ├── amy/            (fp32/ & int8-hq/)
    ├── amy-1p1m/       (fp32/ & int8-hq/)
    ├── amy-1p8m/       (fp32/ & int8-hq/)
    ├── kristin/        (fp32/ & int8-hq/)
    ├── hfc/            (fp32/ & int8-hq/)
    ├── vi/             (fp32/ & int8-hq/)
    ├── id/             (fp32/ & int8-hq/)
    ├── heart-nano/     (fp32/ & int8-hq/)
    ├── heart/          (fp32/ & int8-hq/)
    ├── de-thorsten/    (fp32/ & int8-hq/)
    ├── es-davefx/      (fp32/ & int8-hq/)
    ├── fr-siwis/       (fp32/ & int8-hq/)
    ├── it-serena/      (fp32/ & int8-hq/)
    ├── pt-cadu/        (fp32/ & int8-hq/)
    └── ru-irina/       (fp32/ & int8-hq/)

9. 许可证与归属

  • 基础模型:ampixa/sanoTTS
  • 原作者:Ampixa
  • 许可证:GNU General Public License v3.0 (GPL-3.0)(见 LICENSE)
  • 声明与归属:由 Sakura / webmp3 完成的衍生转换(见 NOTICE.md)。本项目不是 Ampixa 的官方发布。
Downloads last month
29
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webmp3/Sakura-sanoTTS-ONNX

Base model

ampixa/sanoTTS
Quantized
(3)
this model

Collection including webmp3/Sakura-sanoTTS-ONNX