Instructions to use openbmb/VoxCPM2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- VoxCPM
How to use openbmb/VoxCPM2 with VoxCPM:
import soundfile as sf from voxcpm import VoxCPM model = VoxCPM.from_pretrained("openbmb/VoxCPM2") wav = model.generate( text="VoxCPM is an innovative end-to-end TTS model from ModelBest, designed to generate highly expressive speech.", prompt_wav_path=None, # optional: path to a prompt speech for voice cloning prompt_text=None, # optional: reference text cfg_value=2.0, # LM guidance on LocDiT, higher for better adherence to the prompt, but maybe worse inference_timesteps=10, # LocDiT inference timesteps, higher for better result, lower for fast speed normalize=True, # enable external TN tool denoise=True, # enable external Denoise tool retry_badcase=True, # enable retrying mode for some bad cases (unstoppable) retry_badcase_max_times=3, # maximum retrying times retry_badcase_ratio_threshold=6.0, # maximum length restriction for bad case detection (simple but effective), it could be adjusted for slow pace speech ) sf.write("output.wav", wav, 16000) print("saved: output.wav") - Notebooks
- Google Colab
- Kaggle
Thank you to everyone who developed this model.
#7
by abbas78 - opened
Thank you for sharing this impressive model. I’ve tested it extensively in Turkish, and it works flawlessly. So far, this model has delivered the best performance among open-source options. It follows English instructions very well. I’ve narrated YouTube tutorial videos and dubbed movie scenes. The results are excellent.
Dubbing is a much harder test than narration, so that's a useful data point. Narration can run long without anyone noticing, but a dubbed line has to fit the shot, and most TTS drifts on duration once the sentence gets long.
Curious how it handled Turkish suffix chains. Long agglutinative words are usually where stress placement falls apart.