Supertonic 3 β€” AX650 / AX8850 (LLM-8850) NPU

Supertone/supertonic-3 converted with Axera Pulsar2 7.0-patch1 to run on the AX650 / AX8850 NPU, tested on the M5Stack LLM-8850 PCIe card (AXCL). The flow-matching estimator (8 steps) and the vocoder run on the card. Text processing, the duration predictor, the text encoder and the Euler update stay on the host CPU.

On the card it reads Korean, Japanese and English almost as accurately as the original (table below). There are two static sizes ("buckets"): a message of up to about 6.7 s takes about 0.18 s, and one of up to about 13 s takes about 0.35 s. Card memory is about 220 MiB with both buckets loaded, or 99 MiB with the short one only. The host runtime is the HTTP server in JonPark0/TTSBot (server/), written for a Discord TTS bot.

ν•œκ΅­μ–΄ μ•ˆλ‚΄λŠ” μ•„λž˜μ— μžˆμŠ΅λ‹ˆλ‹€.

Accuracy and speed

60 sentences (Korean standard 20, Korean chat 20, Japanese 10, English 10) read by whisper-large-v3-turbo. Korean and Japanese are scored by CER (Japanese in kana), English by WER. UTMOS is utmos22_strong, an English-trained proxy. On human Korean speech the same ASR gives 1.7% CER.

Korean standard CER Korean chat CER Japanese CER English WER UTMOS Time per sentence
Original ONNX (CPU) 0.0% 3.9% 1.7% 1.4% 4.00 –
Static graph + Erf approximation (CPU float, same noise) 0.3% 5.5% 1.7% 1.4% 3.99 –
Card estimator + CPU vocoder 0.3% 5.5% 2.2% 1.4% 4.00 0.53 s
Card estimator + card vocoder (this repo) 0.3% 6.5% 2.2% 1.4% 3.95 0.18 s

Only two sentences are read differently by the card and by the float model with the same noise.

  • Per call on the card: estimator about 18 ms per step (Γ— 8), vocoder about 17 ms.
  • Through the TTSBot HTTP server: 0.175–0.178 s for a short message.

Long messages (192 bucket)

30 messages of 7–14 s (20 Korean, 5 Japanese, 5 English) on the card, voice F1. With the short bucket only, the server has to split each message into 2–4 pieces. With both buckets, 29 of 30 messages are read in one piece.

Card Korean CER Japanese CER English WER Korean UTMOS Median time per message Pieces
96 bucket only 0.1% 1.5% 0% 3.85 0.39 s 70
96 + 192 buckets 0.0% 0.8% 0% 3.91 0.35 s 31
  • One 192 call: estimator about 34 ms per step (Γ— 8), vocoder about 46 ms.
  • Compiled vs float (pulsar2 run): cosine 0.9999 and 0.9995 at estimator steps 0 and 7, 0.994 for the vocoder.
  • The gain is fewer breaks inside a sentence rather than accuracy. In CPU float, Korean UTMOS was 4.06 for the original, 3.85 with the 96 bucket only, and 4.01 with both buckets.
  • These runs are before the server's edge-quiet trimming (below). With trimming, the kept speech samples are identical, but Korean UTMOS reads about 0.05 lower (3.91 β†’ 3.86). The ASR error is the same.

Voices

All 10 preset voices on the card, each reading the same 20 Korean sentences (10 standard, 10 chat). With 10 sentences per group, one or two misread sentences move the numbers a lot. All voices are usable.

Voice Standard CER Chat CER UTMOS
F1 0.6% 3.7% 3.85
F2 0.0% 8.1% 3.66
F3 0.0% 2.0% 4.03
F4 0.0% 13.0% 3.74
F5 0.0% 3.3% 3.72
M1 0.0% 2.7% 3.67
M2 0.0% 0.0% 3.55
M3 0.6% 7.4% 3.54
M4 0.0% 2.5% 3.84
M5 0.0% 2.0% 3.68

Contents

File Description
st_est.axmodel Flow-matching estimator, static text length T = 96 and latent length L = 96 (about 6.7 s of audio)
st_voc.axmodel Vocoder, L = 96 β†’ 44.1 kHz audio
st_est_T192_L192.axmodel Estimator, T = L = 192 (about 13 s)
st_voc_L192.axmodel Vocoder, L = 192
st_rope_T96_L96.onnx, st_rope_T192_L192.onnx Rotary sin/cos tables for the actual text and latent lengths, one per bucket (host, CPU)
time_table.npy Time embeddings for the 8 flow steps (host)
LICENSE, NOTICE Original license (BigScience OpenRAIL-M) and the list of changes
File sha256
st_est.axmodel 5fec2e6d05f608cecfcdadf4de198916ba810473c73034508788ad450226a561
st_voc.axmodel 124620f4104678623b112f936306cf0644592315c0b422767ed3240a616696da
st_est_T192_L192.axmodel 2005389fe06fa131fbbf3fa66777856ea82718a6aaf0f7db89ed049c2e815fdb
st_voc_L192.axmodel 2f7a3635dc676d6402400ca291f34be9e41223e702a15c46d519718b9d29d4af

The duration predictor, the text encoder, the text processor and the voice styles are not in this repo. The host runtime downloads them from Supertone/supertonic-3 through the supertonic Python package on first run.

Quick start

Requirements: an AX650 / AX8850 AXCL card with the AXCL driver and runtime installed (libaxcl_rt.dll on Windows, libaxcl_rt.so on Linux), Python 3.10+.

git lfs install
git clone https://huggingface.co/jonpark0/supertonic-3-AX650
git clone https://github.com/JonPark0/TTSBot && cd TTSBot
pip install -r server/requirements.txt          # numpy, onnxruntime, supertonic

python -m server.say --models ../supertonic-3-AX650 --out out/say "μ•ˆλ…•ν•˜μ„Έμš”" "γ…‡γ…‹ 10λΆ„ 뒀에 λ“€μ–΄κ°ˆκ²Œ"
python -m server.tts_server --models ../supertonic-3-AX650 --port 8850
curl -s localhost:8850/tts -H 'Content-Type: application/json' \
  -d '{"text": "였늘 μ˜€ν›„ 3μ‹œ 30뢄에 νšŒμ˜κ°€ μžˆμŠ΅λ‹ˆλ‹€.", "lang": "ko", "voice": "F1"}' -o out.wav
  • lang: ko, ja or en. Voices: F1–F5, M1–M5.
  • The server loads every bucket in the folder (--buckets 96 keeps only the short one). Messages are split with the duration predictor to fit the largest bucket: by sentence first, then by clause, then into balanced word groups that prefer to end after a connective ending (after particles for Japanese). Each piece runs on the smallest bucket it fits.
  • The model pads each piece with about 0.5 s of quiet before speech and 0.6 s after. The server trims this to 0.06 s and 0.1 s and joins pieces with 0.2–0.3 s pauses.
  • See server/README.md in TTSBot for the API, error codes, text normalization for chat messages, and a systemd example.

Limitations

  • Static shapes: at most 192 text tokens and 192 latent frames (about 13 s) per card call, or 96 / 6.7 s with the short bucket. Longer text has to be split, as the TTSBot server does.
  • Tested with Korean, Japanese and English and the 10 preset voices. The other languages of Supertonic 3 were not measured. Custom voice styles were not tried.
  • One request at a time per card.
  • Tested on Windows 11 with the AXCL Windows driver. Linux hosts use the same code with libaxcl_rt.so but have not been tested on hardware.

How it was converted

  • Pulsar2 7.0-patch1, target AX650 (NPU3), U16 quantization, MinMax calibration on 256 estimator inputs and 32 vocoder inputs. Calibration used voices F1 and M1, with short sentences for the 96 bucket and 7–13 s sentences for the 192 bucket.
  • Graph changes for Pulsar2 and the NPU:
    • 3D Pad(mode=edge) replaced by Slice + Tile + Concat. Padded frames are first filled with the last valid frame so the static graph matches the dynamic original.
    • Rotary sin/cos and the time embedding depend on the actual lengths and the step, so they are computed on the host and fed as inputs.
    • Vocoder 1D convolutions rewritten as 1Γ—K 2D convolutions (the 1D form failed to compile).
    • The compiled Erf (in GELU) gives wrong results on the NPU in this Pulsar2 version, so it is replaced by erf(u) β‰ˆ tanh(1.1283792Β·uΒ·(1 + 0.08943Β·uΒ²)) with u clipped to Β±4 (maximum error 3.6e-4).
  • pulsar2 run on the compiled models matched the card output bit for bit.
  • Scripts: server/convert/ in JonPark0/TTSBot.

License

These files are Derivatives of the Model under the BigScience OpenRAIL-M license of Supertonic 3. See LICENSE and NOTICE. The use restrictions in Attachment A of the license apply to you and to anyone you pass the model to. For example, restriction (e) requires saying that content is machine generated when you publish it, so a bot using this model should make clear its voice is synthetic.


ν•œκ΅­μ–΄

Supertone의 Supertonic 3λ₯Ό Axera Pulsar2 7.0-patch1둜 λ³€ν™˜ν•΄ AX650 / AX8850 NPUμ—μ„œ λŒλ¦¬λŠ” λͺ¨λΈμž…λ‹ˆλ‹€. M5Stack LLM-8850 PCIe μΉ΄λ“œ(AXCL)μ—μ„œ μ‹œν—˜ν–ˆμŠ΅λ‹ˆλ‹€. μΉ΄λ“œμ—μ„œλŠ” μΆ”μ •κΈ°(8단계)와 보코더가 돌고, ν…μŠ€νŠΈ 처리·길이 μ˜ˆμΈ‘κΈ°Β·ν…μŠ€νŠΈ μΈμ½”λ”Β·μ˜€μΌλŸ¬ 갱신은 호슀트 CPUμ—μ„œ λ•λ‹ˆλ‹€.

  • 정확도: 60λ¬Έμž₯ λ°›μ•„μ“°κΈ° 였λ₯˜μœ¨μ΄ ν•œκ΅­μ–΄ ν‘œμ€€ 0.3%, μ±„νŒ… 6.5%, 일본어 2.2%, μ˜μ–΄ 1.4%μž…λ‹ˆλ‹€. 원본은 0.0%, 3.9%, 1.7%, 1.4%μž…λ‹ˆλ‹€. 같은 작음의 float 결과와 λ‹€λ₯΄κ²Œ 읽힌 λ¬Έμž₯은 두 κ°œλΏμž…λ‹ˆλ‹€.
  • 속도: μ•½ 6.7μ΄ˆκΉŒμ§€μ˜ λ©”μ‹œμ§€λŠ” μ•½ 0.18초, μ•½ 13μ΄ˆκΉŒμ§€λŠ” κΈ΄ 버킷(T=L=192)으둜 ν•œ λ²ˆμ— μ•½ 0.35μ΄ˆμž…λ‹ˆλ‹€. μΉ΄λ“œ λ©”λͺ¨λ¦¬λŠ” 두 버킷을 λ‹€ 올리면 μ•½ 220 MiB, 짧은 λ²„ν‚·λ§Œμ΄λ©΄ μ•½ 99 MiBμž…λ‹ˆλ‹€.
  • κΈ΄ λ©”μ‹œμ§€: 7–14초 λ©”μ‹œμ§€ 30개 쀑 29개λ₯Ό ν•œ λ²ˆμ— 읽어, λ¬Έμž₯ μ€‘κ°„μ—μ„œ λŠκΈ°λŠ” 곳이 μ‚¬λΌμ‘ŒμŠ΅λ‹ˆλ‹€. λ°›μ•„μ“°κΈ° μ •ν™•λ„λŠ” 두 방식 λͺ¨λ‘ 거의 μ™„λ²½ν•©λ‹ˆλ‹€.
  • λͺ©μ†Œλ¦¬: F1–F5, M1–M5 10μ’… λͺ¨λ‘ μΉ΄λ“œμ—μ„œ μ‹œν—˜ν–ˆκ³  λͺ¨λ‘ μ“Έ λ§Œν•©λ‹ˆλ‹€. λͺ©μ†Œλ¦¬λ³„ κ²°κ³ΌλŠ” μœ„ ν‘œμ— μžˆμŠ΅λ‹ˆλ‹€.
  • μ‚¬μš©: JonPark0/TTSBot의 server/κ°€ 이 μ €μž₯μ†Œ 폴더λ₯Ό --models둜 λ°›μŠ΅λ‹ˆλ‹€.
    git lfs install
    git clone https://huggingface.co/jonpark0/supertonic-3-AX650
    git clone https://github.com/JonPark0/TTSBot && cd TTSBot
    pip install -r server/requirements.txt
    python -m server.tts_server --models ../supertonic-3-AX650
    
    길이 예츑기, ν…μŠ€νŠΈ 인코더, λͺ©μ†Œλ¦¬λŠ” 처음 μ‹€ν–‰ν•  λ•Œ supertonic νŒ¨ν‚€μ§€κ°€ 원본 μ €μž₯μ†Œμ—μ„œ λ‚΄λ €λ°›μŠ΅λ‹ˆλ‹€.
  • μ œν•œ: μΉ΄λ“œ 호좜 ν•œ λ²ˆμ— ν…μŠ€νŠΈ 192토큰, μŒμ„± μ•½ 13μ΄ˆκΉŒμ§€μž…λ‹ˆλ‹€(짧은 버킷은 96, 6.7초). 더 κΈ΄ λ©”μ‹œμ§€λŠ” TTSBot μ„œλ²„κ°€ λ¬Έμž₯, μ‰Όν‘œ, 띄어쓰기 μˆœμ„œλ‘œ λ‚˜λˆ  μ½μŠ΅λ‹ˆλ‹€. ν•œκ΅­μ–΄Β·μΌλ³Έμ–΄Β·μ˜μ–΄μ™€ κΈ°λ³Έ λͺ©μ†Œλ¦¬ 10μ’…λ§Œ μ‹œν—˜ν–ˆμŠ΅λ‹ˆλ‹€. μš”μ²­μ€ μΉ΄λ“œλ‹Ή ν•œ λ²ˆμ— ν•˜λ‚˜μž…λ‹ˆλ‹€. Windows 11μ—μ„œ μ‹œν—˜ν–ˆκ³ , LinuxλŠ” 같은 μ½”λ“œμ§€λ§Œ μ‹€μ œ μΉ΄λ“œλ‘œλŠ” 아직 μ‹œν—˜ν•˜μ§€ μ•Šμ•˜μŠ΅λ‹ˆλ‹€.
  • λΌμ΄μ„ μŠ€: 원본과 같은 BigScience OpenRAIL-Mμž…λ‹ˆλ‹€(LICENSE, NOTICE). 뢀둝 A의 μ‚¬μš© μ œν•œμ΄ κ·ΈλŒ€λ‘œ μ μš©λ©λ‹ˆλ‹€. 예λ₯Ό λ“€μ–΄ (e)항에 따라 μƒμ„±ν•œ μŒμ„±μ„ 내보일 λ•Œ 기계가 λ§Œλ“  κ²ƒμž„μ„ λ°ν˜€μ•Ό ν•˜λ‹ˆ, 봇에 μ“Έ λ•ŒλŠ” ν•©μ„± μŒμ„±μ΄λΌλŠ” 점을 μ•Œλ € μ£Όμ„Έμš”.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for jonpark0/supertonic-3-AX650

Quantized
(27)
this model