Sakura Audio

Part of the Sakura Audio collection.

Sakura Whistle ONNX

High-parity CPU-first ONNX port of Cactus Whistle for local multilingual speech recognition.

Sakura Whistle ONNX provides a complete, CPU-optimized ONNX Runtime build and INT8 quantization of Cactus Compute's Whistle speech foundation model.

This release represents a high-fidelity full-architecture ONNX CPU port (55.15M parameters), retaining the complete conformer/transformer architecture, positional encodings, residual routing, causal convolutions, and n-gram memory tables for portable, standalone CPU inference.


Why this port?

  • Standard ONNX Runtime: Runs out of the box with standard onnxruntime on CPU without custom C++ compilers or proprietary binary engines.
  • CPU-First INT8 Release: Primary INT8 build (125.4 MB) provides an efficient, portable footprint on standard CPUs.
  • Python-Friendly Runtime: Clean, pure-Python runtime without any PyTorch or GPU dependency for deployment.
  • Multilingual Speech Recognition: Supports 7 European languages (English, German, French, Spanish, Italian, Dutch, Polish).
  • INT8 Main & FP32 Reference: INT8 is provided as the default recommended build; full-precision FP32 is available as a reference build.
  • Language Detection: Zero-shot speech language identification across supported languages.
  • Heuristic Word Timestamps: Token-aligned word boundaries extracted from cross-attention weights.
  • Keyword Biasing: Logit biasing can increase the probability of selected terminology during transcription.
  • Exposed Encoder Embeddings: Continuous 512-dimensional acoustic representations are exposed for experimentation.

Model Variants & Specifications

Format Directory Path Total Size Parameters Steady RAM Description
ONNX INT8 (Primary) models/full-int8/ 125.42 MB 55.15M ~180 MB Recommended CPU release: per-channel dynamic INT8.
ONNX FP32 (Reference) models/full-fp32/ 227.46 MB 55.15M ~350 MB Full-precision uncompressed reference build.
Upstream .cact Cactus Compute 16.91 MB 55.15M ~45 MB Upstream proprietary binary container.

Comparison to Upstream: The upstream Cactus .cact binary container is substantially more compact (16.91 MB vs. 125.4 MB) and faster in specialized native C++ execution. Sakura Whistle ONNX trades peak native kernel size for broad ecosystem portability, standard ONNX deployment, and pure Python integration across diverse CPU platforms.


Download

Download the recommended standalone INT8 release package using the Hugging Face CLI (hf):

Linux / macOS (Bash)

hf download webmp3/Sakura-Whistle-ONNX \
  --include "config.json" \
  --include "models/full-int8/*" \
  --include "runtime/*" \
  --include "transcribe.py" \
  --include "requirements.txt" \
  --local-dir Sakura-Whistle-ONNX

Windows (PowerShell)

hf download webmp3/Sakura-Whistle-ONNX `
  --include "config.json" `
  --include "models/full-int8/*" `
  --include "runtime/*" `
  --include "transcribe.py" `
  --include "requirements.txt" `
  --local-dir Sakura-Whistle-ONNX

Installation

Install the minimal runtime dependencies (pure CPU, zero PyTorch dependency):

pip install -r requirements.txt

(Requirements: onnxruntime>=1.19.0, numpy>=1.24.0, soundfile>=0.12.1, sentencepiece>=0.2.0, scipy>=1.10.0)


CLI Usage

Use the included transcribe.py command-line interface:

# German speech transcription with word timestamps
python transcribe.py --audio speech_de.wav --language de --timestamps

# English speech transcription with auto-detection and JSON output
python transcribe.py --audio speech_en.wav --json

# Transcription with optional keyword biasing
python transcribe.py --audio speech.wav --keywords "Sakura,ONNX,Hugging Face"

CLI Output Example:

Language:   de
Transcript: Guten Morgen, wie wird das Wetter heute?

Word Timestamps:
    0.32 -  0.63 s  Guten                     (p=0.980)
    0.72 -  1.43 s  Morgen,                   (p=0.980)
    1.43 -  1.53 s  wie                       (p=0.980)
    1.53 -  1.59 s  wird                      (p=0.980)
    1.68 -  1.83 s  das                       (p=0.980)
    1.83 -  2.23 s  Wetter                    (p=0.980)
    2.24 -  3.44 s  heute?                    (p=0.980)

Stats: TTFT=93.9ms | Decode TPS=54.9 | Total Time=3.540s

Python API

from runtime.whistle_full_onnx import WhistleFullONNX

# Initialize ONNX Runtime model on CPU (default: models/full-int8)
model = WhistleFullONNX("models/full-int8")

# Transcribe 16 kHz audio (file path or numpy array)
result = model.transcribe(
    "speech.wav",
    language="de",         # "de", "en", "fr", "es", "it", "nl", "pl", or None for auto-detect
    word_timestamps=True,  # Extract heuristic word timestamps
    keywords=["Sakura"]    # Optional keyword biasing
)

print(f"Language:   {result.language}")
print(f"Transcript: {result.text}")

# Word timestamps
for word_info in result.words:
    print(f"  {word_info['start']:5.2f} - {word_info['end']:5.2f} s: {word_info['word']}")

# Extract 512-dim continuous speech representations for experimentation
embeddings = model.embed("speech.wav")
print("Embeddings shape:", embeddings.shape)  # [num_frames, 512]

Supported Languages

Code Language Native Example
en English "Good morning! How is the weather today?"
de Deutsch (German) "Guten Morgen, wie wird das Wetter heute?"
fr Français (French) "Bonjour, quel temps fait-il aujourd'hui?"
es Español (Spanish) "Buenos días, ¿cómo está el clima hoy?"
it Italiano (Italian) "Buongiorno, che tempo fa oggi?"
nl Nederlands (Dutch) "Goedemorgen, hoe is het weer vandaag?"
pl Polski (Polish) "Dzień dobry, jaka jest dzisiaj pogoda?"

Quality Benchmarks

Word Error Rate (WER) measured on the included local validation set (20 German, 20 English, 5 Multilingual):

Backend Engine German WER (de) English WER (en) Model Size Status
Cactus Official (.cact) 23.59% 19.36% 16.9 MB Upstream Baseline
Full FP32 ONNX CPU 23.98% 21.32% 227.5 MB Reference Build
Full INT8 ONNX CPU 24.70% 21.20% 125.4 MB Primary Release

Note: These values are from the included local validation set and should not be interpreted as broad benchmark results.

INT8 quantization achieves a 44.9% reduction in file size compared to FP32 while remaining close to the FP32 and Cactus baselines on the included validation set.

Representative Parity Examples

  • German (de_01_wetter.wav):

    • Ground Truth: "Guten Morgen, wie wird das Wetter heute?"
    • Cactus: "Guten Morgen, wie wird das Wetter heute?"
    • ONNX INT8: "Guten Morgen, wie wird das Wetter heute?"
    • Status: Exact transcript match on this sample (0.00% WER).
  • English (en_01_weather.wav):

    • Ground Truth: "Good morning, how is the weather today?"
    • Cactus: "Good morning! How is the weather to day?\""
    • ONNX INT8: "Good morning! How is the weather today?"
    • Status: Exact transcript match to Ground Truth (0.00% WER vs. Target). Measured WER against Cactus is 25.00% (2 out of 8 words) due to Cactus segmenting "today" as "to day".

Limitations

  • Heuristic Timestamps: Word timestamps are heuristic cross-attention alignments (80 ms resolution), not a forced-alignment system.
  • Keyword Biasing: Keyword logit biasing can increase the probability of selected terms, but does not guarantee recognition.
  • Streaming: This model is non-streaming; chunked or VAD-based processing can be used for near-real-time applications.
  • Encoder Embeddings: Exposed encoder embeddings are provided for experimentation and have not been benchmarked for diarization, speaker recognition, or semantic search.
  • Short Audio Language Classification: Very short audio clips (< 2 s) between closely related languages (such as Dutch and German) can occasionally be confused. Specifying explicit language tags is recommended for brief audio cues.

Architecture Details

Full architecture parity required reconstruction of the original audio frontend, positional encoding, residual routing, causal convolution, and memory components:

  • Audio Frontend: Utterance AGC, Hann windowing, 512 RFFT with 80-mel filterbank, log compression, and per-utterance CMVN.
  • AudioStem & Conformer Encoder: 8-layer Conformer stack with depthwise convolutions, RoPE self-attention, and 4-lane Sinkhorn manifold routing.
  • Decoder: 8-layer GQA Decoder with RoPE positional encodings, 3-tap causal convolutions on Q/K/V, cross-attention, Engram n-gram memory lookups, and 4-lane routing.

License & Attribution


中文说明 · 樱花 (Simplified Chinese)

English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。

Sakura Audio

属于 Sakura Audio 合集。

Sakura Whistle ONNX

Cactus Whistle 的高一致性、CPU 优先的 ONNX 移植版,用于本地多语言语音识别。

Sakura Whistle ONNX 提供了 Cactus Compute 的 Whistle 语音基础模型的完整的、针对 CPU 优化的 ONNX Runtime 构建以及 INT8 量化版本。

本次发布是一个高保真的、完整架构的 ONNX CPU 移植(55.15M 参数),保留了完整的 conformer/transformer 架构、位置编码、残差路由、因果卷积和 n-gram 记忆表,用于可移植的、独立的 CPU 推理。


为什么要做这个移植?

  • 标准 ONNX Runtime: 在 CPU 上使用标准 onnxruntime 即可开箱即用,无需自定义 C++ 编译器或专有二进制引擎。
  • CPU 优先的 INT8 发布: 主要的 INT8 构建(125.4 MB)在标准 CPU 上提供高效、可移植的体积。
  • 对 Python 友好的运行时: 干净的纯 Python 运行时,部署时没有任何 PyTorch 或 GPU 依赖。
  • 多语言语音识别: 支持 7 种欧洲语言(英语、德语、法语、西班牙语、意大利语、荷兰语、波兰语)。
  • INT8 为主,FP32 为参照: 默认推荐 INT8 构建;全精度 FP32 作为参照构建提供。
  • 语言检测: 在受支持的语言之间进行零样本语音语种识别。
  • 启发式词级时间戳: 从交叉注意力权重中提取与 token 对齐的词边界。
  • 关键词偏置: 对 logit 进行偏置可提高转写过程中所选术语的概率。
  • 公开的编码器嵌入: 公开 512 维连续声学表示,供实验使用。

模型变体与规格

格式 目录路径 总大小 参数量 稳定 RAM 说明
ONNX INT8(主要) models/full-int8/ 125.42 MB 55.15M ~180 MB 推荐的 CPU 发布版本:逐通道动态 INT8。
ONNX FP32(参照) models/full-fp32/ 227.46 MB 55.15M ~350 MB 全精度、未压缩的参照构建。
上游 .cact Cactus Compute 16.91 MB 55.15M ~45 MB 上游专有的二进制容器。

与上游的比较: 上游的 Cactus .cact 二进制容器要紧凑得多(16.91 MB 对 125.4 MB),并且在专门的原生 C++ 执行中更快。Sakura Whistle ONNX 放弃了原生内核的极致体积,换取广泛的生态系统可移植性、标准的 ONNX 部署,以及在各种 CPU 平台上的纯 Python 集成。


下载

使用 Hugging Face CLI(hf)下载推荐的独立 INT8 发布包:

Linux / macOS (Bash)

hf download webmp3/Sakura-Whistle-ONNX \
  --include "config.json" \
  --include "models/full-int8/*" \
  --include "runtime/*" \
  --include "transcribe.py" \
  --include "requirements.txt" \
  --local-dir Sakura-Whistle-ONNX

Windows (PowerShell)

hf download webmp3/Sakura-Whistle-ONNX `
  --include "config.json" `
  --include "models/full-int8/*" `
  --include "runtime/*" `
  --include "transcribe.py" `
  --include "requirements.txt" `
  --local-dir Sakura-Whistle-ONNX

安装

安装最少的运行时依赖(纯 CPU,零 PyTorch 依赖):

pip install -r requirements.txt

(依赖:onnxruntime>=1.19.0、numpy>=1.24.0、soundfile>=0.12.1、sentencepiece>=0.2.0、scipy>=1.10.0)


命令行用法

使用随附的 transcribe.py 命令行界面:

# German speech transcription with word timestamps
python transcribe.py --audio speech_de.wav --language de --timestamps

# English speech transcription with auto-detection and JSON output
python transcribe.py --audio speech_en.wav --json

# Transcription with optional keyword biasing
python transcribe.py --audio speech.wav --keywords "Sakura,ONNX,Hugging Face"

命令行输出示例:

Language:   de
Transcript: Guten Morgen, wie wird das Wetter heute?

Word Timestamps:
    0.32 -  0.63 s  Guten                     (p=0.980)
    0.72 -  1.43 s  Morgen,                   (p=0.980)
    1.43 -  1.53 s  wie                       (p=0.980)
    1.53 -  1.59 s  wird                      (p=0.980)
    1.68 -  1.83 s  das                       (p=0.980)
    1.83 -  2.23 s  Wetter                    (p=0.980)
    2.24 -  3.44 s  heute?                    (p=0.980)

Stats: TTFT=93.9ms | Decode TPS=54.9 | Total Time=3.540s

Python API

from runtime.whistle_full_onnx import WhistleFullONNX

# Initialize ONNX Runtime model on CPU (default: models/full-int8)
model = WhistleFullONNX("models/full-int8")

# Transcribe 16 kHz audio (file path or numpy array)
result = model.transcribe(
    "speech.wav",
    language="de",         # "de", "en", "fr", "es", "it", "nl", "pl", or None for auto-detect
    word_timestamps=True,  # Extract heuristic word timestamps
    keywords=["Sakura"]    # Optional keyword biasing
)

print(f"Language:   {result.language}")
print(f"Transcript: {result.text}")

# Word timestamps
for word_info in result.words:
    print(f"  {word_info['start']:5.2f} - {word_info['end']:5.2f} s: {word_info['word']}")

# Extract 512-dim continuous speech representations for experimentation
embeddings = model.embed("speech.wav")
print("Embeddings shape:", embeddings.shape)  # [num_frames, 512]

支持的语言

代码 语言 本语言示例
en English "Good morning! How is the weather today?"
de Deutsch (German) "Guten Morgen, wie wird das Wetter heute?"
fr Français (French) "Bonjour, quel temps fait-il aujourd'hui?"
es Español (Spanish) "Buenos días, ¿cómo está el clima hoy?"
it Italiano (Italian) "Buongiorno, che tempo fa oggi?"
nl Nederlands (Dutch) "Goedemorgen, hoe is het weer vandaag?"
pl Polski (Polish) "Dzień dobry, jaka jest dzisiaj pogoda?"

质量基准

在随附的本地验证集(20 条德语、20 条英语、5 条多语言)上测得的词错误率(WER):

后端引擎 德语 WER (de) 英语 WER (en) 模型大小 状态
Cactus 官方(.cact) 23.59% 19.36% 16.9 MB 上游基线
完整 FP32 ONNX CPU 23.98% 21.32% 227.5 MB 参照构建
完整 INT8 ONNX CPU 24.70% 21.20% 125.4 MB 主要发布版本

说明:这些数值来自随附的本地验证集,不应被解读为广泛的基准结果。

与 FP32 相比,INT8 量化使文件大小减少了 44.9%,同时在随附的验证集上与 FP32 和 Cactus 基线保持接近。

有代表性的一致性示例

  • 德语(de_01_wetter.wav):

    • 标准答案:"Guten Morgen, wie wird das Wetter heute?"
    • Cactus:"Guten Morgen, wie wird das Wetter heute?"
    • ONNX INT8:"Guten Morgen, wie wird das Wetter heute?"
    • 状态: 在该样本上转写完全匹配(0.00% WER)。
  • 英语(en_01_weather.wav):

    • 标准答案:"Good morning, how is the weather today?"
    • Cactus:"Good morning! How is the weather to day?\""
    • ONNX INT8:"Good morning! How is the weather today?"
    • 状态: 与标准答案完全匹配(相对目标的 WER 为 0.00%)。相对于 Cactus 测得的 WER 为 25.00%(8 个词中有 2 个),原因是 Cactus 把 "today" 切分成了 "to day"。

局限

  • 启发式时间戳: 词级时间戳是启发式的交叉注意力对齐(80 ms 分辨率),不是强制对齐系统。
  • 关键词偏置: 关键词 logit 偏置可以提高所选词语的概率,但不保证一定能识别出来。
  • 流式处理: 该模型不是流式的;可以使用分块或基于 VAD 的处理来满足近实时应用。
  • 编码器嵌入: 公开的编码器嵌入仅供实验使用,尚未针对说话人分离、说话人识别或语义搜索进行基准测试。
  • 短音频语种分类: 非常短的音频片段(< 2 s)在相近语言之间(例如荷兰语和德语)偶尔会被混淆。对于简短的音频提示,建议明确指定语言标签。

架构详情

要达到完整的架构一致性,需要重建原始的音频前端、位置编码、残差路由、因果卷积和记忆组件:

  • 音频前端: 语句级 AGC、Hann 加窗、512 点 RFFT 配 80 维梅尔滤波器组、对数压缩和逐语句 CMVN。
  • AudioStem 与 Conformer 编码器: 8 层 Conformer 堆叠,带有深度卷积、RoPE 自注意力和 4 通道 Sinkhorn 流形路由。
  • 解码器: 8 层 GQA 解码器,带有 RoPE 位置编码、作用于 Q/K/V 的 3 抽头因果卷积、交叉注意力、Engram n-gram 记忆查找和 4 通道路由。

许可证与归属

Downloads last month
48
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webmp3/Sakura-Whistle-ONNX

Quantized
(3)
this model

Collection including webmp3/Sakura-Whistle-ONNX