VoxCPM2 GGUF
GGUF conversion of openbmb/VoxCPM2 for use with CrispASR.
Model Details
- Architecture: TSLM (28L MiniCPM-4) + RALM (8L) + LocEnc (12L) + LocDiT (12L) + AudioVAE
- Parameters: ~2B
- Output: 48 kHz mono PCM
- Languages: 30 languages (Arabic, Burmese, Chinese, Danish, Dutch, English, Finnish, French, German, Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Khmer, Korean, Lao, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, Tagalog, Thai, Turkish, Vietnamese) plus 9 Chinese dialects
- License: Apache 2.0
- Voice Cloning: Supported via reference audio
Features
- Tokenizer-free: Directly generates audio patches via diffusion (no discrete audio tokens)
- High quality: CFM-based generation with classifier-free guidance
- Multilingual: 30-language support via CJK-split BPE tokenizer (73k vocab)
- Voice cloning: Zero-shot voice cloning from reference audio
Usage with CrispASR
# Zero-shot TTS
crispasr -m voxcpm2-f16.gguf \
--tts "Hello, this is VoxCPM2 speaking." \
--tts-output output.wav
# Quantized (smaller, faster)
crispasr -m voxcpm2-q4_k.gguf \
--tts "Hello world" --tts-output output.wav
Files
| File | Size | Description |
|---|---|---|
voxcpm2-f16.gguf |
4.63 GB | F16 weights (full precision) |
voxcpm2-q4_k.gguf |
~1.5 GB | Q4_K quantized (faster, slightly lower quality) |
voxcpm2-q8_0.gguf |
2.83 GB | Q8_0 quantized (near-F16 quality) |
voxcpm2-q8_0-locdit-f16.gguf |
3.03 GB | Q8_0 with the LocDiT diffusion head kept in F16 β fastest on Vulkan GPUs (see below) |
voxcpm2-ref.gguf |
371 KB | Reference activation dump for numerical validation |
Which file on a Vulkan GPU?
On Vulkan, the diffusion head (CFM, ~80% of synthesis time at the default 10 steps) is
compute-bound in small matrix multiplications, and GPU matrix engines run those fastest
from F16 weights. The 2B text model, on the other hand, is bandwidth-bound and prefers
q8_0. voxcpm2-q8_0-locdit-f16.gguf combines the two. Measured on a T4 under Vulkan:
CFM per audio step 70.2 β 59.4 ms at 10 steps, and 46.7 ms with
CRISPASR_VOXCPM2_INFERENCE_STEPS=8. Full F16 was slower overall. Built with
crispasr-quantize voxcpm2-f16.gguf out.gguf q8_0 --tensor-type "^locdit\.=f16"; a CPU
synthesis round-trips through ASR exactly
(CrispASR#461).
Numerical Validation
The voxcpm2-ref.gguf file contains intermediate activation tensors captured from the PyTorch reference implementation. Used with crispasr-diff to validate the C++ inference path:
crispasr-diff voxcpm2-tts voxcpm2-f16.gguf voxcpm2-ref.gguf samples/jfk.wav
Current status: 12 pass, 0 fail across all transformer stages (TSLM, RALM, LocEnc, LocDiT, projections).
Captured stages: text_input_ids, locenc_in, locenc_out, enc_to_lm, tslm_layer_0_out, tslm_layer_27_out, tslm_prefill_out, ralm_prefill_out, lm_to_dit_hidden, res_to_dit_hidden, cfm_step0_z, cfm_step0_result, stop_logits_step0.
Architecture
Text β BPE tokenize β TSLM (28L causal MiniCPM-4, GQA 16h/2kv, LongRoPE)
β
FSQ bottleneck (tanhβroundβlinear)
β
RALM (8L causal, no RoPE, GQA 16h/2kv)
β
Projections: lm_to_dit + res_to_dit β mu [2048]
β
LocDiT (12L bidirectional, CFM Euler solver, 10 steps, cfg=2.0)
β
Predicted latent patch [4 frames Γ 64 dims]
β
LocEnc (12L bidirectional) β next TSLM input
β (AR loop until stop)
AudioVAE decoder β 48 kHz PCM
Conversion
Converted using models/convert-voxcpm2-to-gguf.py from the CrispASR repository:
python models/convert-voxcpm2-to-gguf.py \
--input openbmb/VoxCPM2 \
--output voxcpm2-f16.gguf
Acknowledgments
Original model by OpenBMB. Apache 2.0 license.
Provenance and EU AI Act Art. 53 note
- Upstream model: openbmb/VoxCPM2 β published by
openbmb. - Upstream licence:
apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - What was done here: format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
- Training data: documented β where it is documented at all β by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
- Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.
- Downloads last month
- 4,709
Model tree for cstr/voxcpm2-GGUF
Base model
openbmb/VoxCPM2