Instructions to use jonpark0/supertonic-3-AX650 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Supertonic
How to use jonpark0/supertonic-3-AX650 with Supertonic:
from supertonic import TTS tts = TTS(auto_download=True) style = tts.get_voice_style(voice_name="M1") text = "The train delay was announced at 4:45 PM on Wed, Apr 3, 2024 due to track maintenance." wav, duration = tts.synthesize(text, voice_style=style) tts.save_audio(wav, "output.wav")
- Notebooks
- Google Colab
- Kaggle
Supertonic 3 β AX650 / AX8850 (LLM-8850) NPU
Supertone/supertonic-3 converted with Axera Pulsar2 7.0-patch1 to run on the AX650 / AX8850 NPU, tested on the M5Stack LLM-8850 PCIe card (AXCL). The flow-matching estimator (8 steps) and the vocoder run on the card. Text processing, the duration predictor, the text encoder and the Euler update stay on the host CPU.
On the card it reads Korean, Japanese and English almost as accurately as the original (table below). There are two static sizes ("buckets"): a message of up to about 6.7 s takes about 0.18 s, and one of up to about 13 s takes about 0.35 s. Card memory is about 220 MiB with both buckets loaded, or 99 MiB with the short one only. The host runtime is the HTTP server in JonPark0/TTSBot (server/), written for a Discord TTS bot.
νκ΅μ΄ μλ΄λ μλμ μμ΅λλ€.
Accuracy and speed
60 sentences (Korean standard 20, Korean chat 20, Japanese 10, English 10) read by whisper-large-v3-turbo. Korean and Japanese are scored by CER (Japanese in kana), English by WER. UTMOS is utmos22_strong, an English-trained proxy. On human Korean speech the same ASR gives 1.7% CER.
| Korean standard CER | Korean chat CER | Japanese CER | English WER | UTMOS | Time per sentence | |
|---|---|---|---|---|---|---|
| Original ONNX (CPU) | 0.0% | 3.9% | 1.7% | 1.4% | 4.00 | β |
| Static graph + Erf approximation (CPU float, same noise) | 0.3% | 5.5% | 1.7% | 1.4% | 3.99 | β |
| Card estimator + CPU vocoder | 0.3% | 5.5% | 2.2% | 1.4% | 4.00 | 0.53 s |
| Card estimator + card vocoder (this repo) | 0.3% | 6.5% | 2.2% | 1.4% | 3.95 | 0.18 s |
Only two sentences are read differently by the card and by the float model with the same noise.
- Per call on the card: estimator about 18 ms per step (Γ 8), vocoder about 17 ms.
- Through the TTSBot HTTP server: 0.175β0.178 s for a short message.
Long messages (192 bucket)
30 messages of 7β14 s (20 Korean, 5 Japanese, 5 English) on the card, voice F1. With the short bucket only, the server has to split each message into 2β4 pieces. With both buckets, 29 of 30 messages are read in one piece.
| Card | Korean CER | Japanese CER | English WER | Korean UTMOS | Median time per message | Pieces |
|---|---|---|---|---|---|---|
| 96 bucket only | 0.1% | 1.5% | 0% | 3.85 | 0.39 s | 70 |
| 96 + 192 buckets | 0.0% | 0.8% | 0% | 3.91 | 0.35 s | 31 |
- One 192 call: estimator about 34 ms per step (Γ 8), vocoder about 46 ms.
- Compiled vs float (
pulsar2 run): cosine 0.9999 and 0.9995 at estimator steps 0 and 7, 0.994 for the vocoder. - The gain is fewer breaks inside a sentence rather than accuracy. In CPU float, Korean UTMOS was 4.06 for the original, 3.85 with the 96 bucket only, and 4.01 with both buckets.
- These runs are before the server's edge-quiet trimming (below). With trimming, the kept speech samples are identical, but Korean UTMOS reads about 0.05 lower (3.91 β 3.86). The ASR error is the same.
Voices
All 10 preset voices on the card, each reading the same 20 Korean sentences (10 standard, 10 chat). With 10 sentences per group, one or two misread sentences move the numbers a lot. All voices are usable.
| Voice | Standard CER | Chat CER | UTMOS |
|---|---|---|---|
| F1 | 0.6% | 3.7% | 3.85 |
| F2 | 0.0% | 8.1% | 3.66 |
| F3 | 0.0% | 2.0% | 4.03 |
| F4 | 0.0% | 13.0% | 3.74 |
| F5 | 0.0% | 3.3% | 3.72 |
| M1 | 0.0% | 2.7% | 3.67 |
| M2 | 0.0% | 0.0% | 3.55 |
| M3 | 0.6% | 7.4% | 3.54 |
| M4 | 0.0% | 2.5% | 3.84 |
| M5 | 0.0% | 2.0% | 3.68 |
Contents
| File | Description |
|---|---|
st_est.axmodel |
Flow-matching estimator, static text length T = 96 and latent length L = 96 (about 6.7 s of audio) |
st_voc.axmodel |
Vocoder, L = 96 β 44.1 kHz audio |
st_est_T192_L192.axmodel |
Estimator, T = L = 192 (about 13 s) |
st_voc_L192.axmodel |
Vocoder, L = 192 |
st_rope_T96_L96.onnx, st_rope_T192_L192.onnx |
Rotary sin/cos tables for the actual text and latent lengths, one per bucket (host, CPU) |
time_table.npy |
Time embeddings for the 8 flow steps (host) |
LICENSE, NOTICE |
Original license (BigScience OpenRAIL-M) and the list of changes |
| File | sha256 |
|---|---|
st_est.axmodel |
5fec2e6d05f608cecfcdadf4de198916ba810473c73034508788ad450226a561 |
st_voc.axmodel |
124620f4104678623b112f936306cf0644592315c0b422767ed3240a616696da |
st_est_T192_L192.axmodel |
2005389fe06fa131fbbf3fa66777856ea82718a6aaf0f7db89ed049c2e815fdb |
st_voc_L192.axmodel |
2f7a3635dc676d6402400ca291f34be9e41223e702a15c46d519718b9d29d4af |
The duration predictor, the text encoder, the text processor and the voice styles are not in this repo. The host runtime downloads them from Supertone/supertonic-3 through the supertonic Python package on first run.
Quick start
Requirements: an AX650 / AX8850 AXCL card with the AXCL driver and runtime installed (libaxcl_rt.dll on Windows, libaxcl_rt.so on Linux), Python 3.10+.
git lfs install
git clone https://huggingface.co/jonpark0/supertonic-3-AX650
git clone https://github.com/JonPark0/TTSBot && cd TTSBot
pip install -r server/requirements.txt # numpy, onnxruntime, supertonic
python -m server.say --models ../supertonic-3-AX650 --out out/say "μλ
νμΈμ" "γ
γ
10λΆ λ€μ λ€μ΄κ°κ²"
python -m server.tts_server --models ../supertonic-3-AX650 --port 8850
curl -s localhost:8850/tts -H 'Content-Type: application/json' \
-d '{"text": "μ€λ μ€ν 3μ 30λΆμ νμκ° μμ΅λλ€.", "lang": "ko", "voice": "F1"}' -o out.wav
lang:ko,jaoren. Voices:F1βF5,M1βM5.- The server loads every bucket in the folder (
--buckets 96keeps only the short one). Messages are split with the duration predictor to fit the largest bucket: by sentence first, then by clause, then into balanced word groups that prefer to end after a connective ending (after particles for Japanese). Each piece runs on the smallest bucket it fits. - The model pads each piece with about 0.5 s of quiet before speech and 0.6 s after. The server trims this to 0.06 s and 0.1 s and joins pieces with 0.2β0.3 s pauses.
- See
server/README.mdin TTSBot for the API, error codes, text normalization for chat messages, and a systemd example.
Limitations
- Static shapes: at most 192 text tokens and 192 latent frames (about 13 s) per card call, or 96 / 6.7 s with the short bucket. Longer text has to be split, as the TTSBot server does.
- Tested with Korean, Japanese and English and the 10 preset voices. The other languages of Supertonic 3 were not measured. Custom voice styles were not tried.
- One request at a time per card.
- Tested on Windows 11 with the AXCL Windows driver. Linux hosts use the same code with
libaxcl_rt.sobut have not been tested on hardware.
How it was converted
- Pulsar2 7.0-patch1, target AX650 (NPU3), U16 quantization, MinMax calibration on 256 estimator inputs and 32 vocoder inputs. Calibration used voices F1 and M1, with short sentences for the 96 bucket and 7β13 s sentences for the 192 bucket.
- Graph changes for Pulsar2 and the NPU:
- 3D
Pad(mode=edge)replaced by Slice + Tile + Concat. Padded frames are first filled with the last valid frame so the static graph matches the dynamic original. - Rotary sin/cos and the time embedding depend on the actual lengths and the step, so they are computed on the host and fed as inputs.
- Vocoder 1D convolutions rewritten as 1ΓK 2D convolutions (the 1D form failed to compile).
- The compiled
Erf(in GELU) gives wrong results on the NPU in this Pulsar2 version, so it is replaced byerf(u) β tanh(1.1283792Β·uΒ·(1 + 0.08943Β·uΒ²))with u clipped to Β±4 (maximum error 3.6e-4).
- 3D
pulsar2 runon the compiled models matched the card output bit for bit.- Scripts:
server/convert/in JonPark0/TTSBot.
License
These files are Derivatives of the Model under the BigScience OpenRAIL-M license of Supertonic 3. See LICENSE and NOTICE. The use restrictions in Attachment A of the license apply to you and to anyone you pass the model to. For example, restriction (e) requires saying that content is machine generated when you publish it, so a bot using this model should make clear its voice is synthetic.
νκ΅μ΄
Supertoneμ Supertonic 3λ₯Ό Axera Pulsar2 7.0-patch1λ‘ λ³νν΄ AX650 / AX8850 NPUμμ λ리λ λͺ¨λΈμ
λλ€. M5Stack LLM-8850 PCIe μΉ΄λ(AXCL)μμ μννμ΅λλ€. μΉ΄λμμλ μΆμ κΈ°(8λ¨κ³)μ 보μ½λκ° λκ³ , ν
μ€νΈ μ²λ¦¬Β·κΈΈμ΄ μμΈ‘κΈ°Β·ν
μ€νΈ μΈμ½λΒ·μ€μΌλ¬ κ°±μ μ νΈμ€νΈ CPUμμ λλλ€.
- μ νλ: 60λ¬Έμ₯ λ°μμ°κΈ° μ€λ₯μ¨μ΄ νκ΅μ΄ νμ€ 0.3%, μ±ν 6.5%, μΌλ³Έμ΄ 2.2%, μμ΄ 1.4%μ λλ€. μλ³Έμ 0.0%, 3.9%, 1.7%, 1.4%μ λλ€. κ°μ μ‘μμ float κ²°κ³Όμ λ€λ₯΄κ² μ½ν λ¬Έμ₯μ λ κ°λΏμ λλ€.
- μλ: μ½ 6.7μ΄κΉμ§μ λ©μμ§λ μ½ 0.18μ΄, μ½ 13μ΄κΉμ§λ κΈ΄ λ²ν·(T=L=192)μΌλ‘ ν λ²μ μ½ 0.35μ΄μ λλ€. μΉ΄λ λ©λͺ¨λ¦¬λ λ λ²ν·μ λ€ μ¬λ¦¬λ©΄ μ½ 220 MiB, μ§§μ λ²ν·λ§μ΄λ©΄ μ½ 99 MiBμ λλ€.
- κΈ΄ λ©μμ§: 7β14μ΄ λ©μμ§ 30κ° μ€ 29κ°λ₯Ό ν λ²μ μ½μ΄, λ¬Έμ₯ μ€κ°μμ λκΈ°λ κ³³μ΄ μ¬λΌμ‘μ΅λλ€. λ°μμ°κΈ° μ νλλ λ λ°©μ λͺ¨λ κ±°μ μλ²½ν©λλ€.
- λͺ©μ리: F1βF5, M1βM5 10μ’ λͺ¨λ μΉ΄λμμ μννκ³ λͺ¨λ μΈ λ§ν©λλ€. λͺ©μλ¦¬λ³ κ²°κ³Όλ μ νμ μμ΅λλ€.
- μ¬μ©: JonPark0/TTSBotμ
server/κ° μ΄ μ μ₯μ ν΄λλ₯Ό--modelsλ‘ λ°μ΅λλ€.κΈΈμ΄ μμΈ‘κΈ°, ν μ€νΈ μΈμ½λ, λͺ©μ리λ μ²μ μ€νν λgit lfs install git clone https://huggingface.co/jonpark0/supertonic-3-AX650 git clone https://github.com/JonPark0/TTSBot && cd TTSBot pip install -r server/requirements.txt python -m server.tts_server --models ../supertonic-3-AX650supertonicν¨ν€μ§κ° μλ³Έ μ μ₯μμμ λ΄λ €λ°μ΅λλ€. - μ ν: μΉ΄λ νΈμΆ ν λ²μ ν μ€νΈ 192ν ν°, μμ± μ½ 13μ΄κΉμ§μ λλ€(μ§§μ λ²ν·μ 96, 6.7μ΄). λ κΈ΄ λ©μμ§λ TTSBot μλ²κ° λ¬Έμ₯, μΌν, λμ΄μ°κΈ° μμλ‘ λλ μ½μ΅λλ€. νκ΅μ΄Β·μΌλ³Έμ΄Β·μμ΄μ κΈ°λ³Έ λͺ©μ리 10μ’ λ§ μννμ΅λλ€. μμ²μ μΉ΄λλΉ ν λ²μ νλμ λλ€. Windows 11μμ μννκ³ , Linuxλ κ°μ μ½λμ§λ§ μ€μ μΉ΄λλ‘λ μμ§ μννμ§ μμμ΅λλ€.
- λΌμ΄μ μ€: μλ³Έκ³Ό κ°μ BigScience OpenRAIL-Mμ
λλ€(
LICENSE,NOTICE). λΆλ‘ Aμ μ¬μ© μ νμ΄ κ·Έλλ‘ μ μ©λ©λλ€. μλ₯Ό λ€μ΄ (e)νμ λ°λΌ μμ±ν μμ±μ λ΄λ³΄μΌ λ κΈ°κ³κ° λ§λ κ²μμ λ°νμΌ νλ, λ΄μ μΈ λλ ν©μ± μμ±μ΄λΌλ μ μ μλ € μ£ΌμΈμ.
Model tree for jonpark0/supertonic-3-AX650
Base model
Supertone/supertonic-3