Generative Modeling via Drifting
Paper • 2602.04770 • Published • 10
Experimental. For the best quality, use the released mel model Vyvo/drifting-tts-tr (v3.1).
This is the one-step Turkish TTS of drifting-tts in the latent space of a pretrained audio VAE, the paper's latent setting (Deng et al., 2026).
Compare it with v3.1 and the VoxCPM2-latent model: Vyvo/drifting-tts-tr-compare.
| file | what |
|---|---|
drifting_tts_dacvae_10k.pt |
the latent TTS model, 10k steps |
dacvae_decoder_ft.pt |
the DAC-VAE decoder fine-tuned for 40k steps on the model's generated latents, with the watermark frozen (below) |
Freya-TR-Eval, first 100 sentences, speaker 722, T = 0.3, α = 2. Whisper large-v3 on 8 kHz band-matched audio; UTMOSv2 and DNSMOS on the full band.
| model | decoder / vocoder | WER | CER | UTMOSv2 | DNSMOS OVRL |
|---|---|---|---|---|---|
| v3.1 (mels, released) | BigVGAN-v2-ft | 0.66% | 0.14% | 2.934 | 3.33 |
| DAC-VAE latents (this) | fine-tuned decoder | 1.32% | 0.30% | 2.710 | 3.26 |
| DAC-VAE latents | released decoder | 9.55% | 2.91% | 1.837 | 1.48 |
| VoxCPM2 latents | fine-tuned decoder | 1.87% | 0.43% | 2.530 | 3.25 |
Details: docs/LATENTS.md (#32).
The DAC-VAE decoder embeds an AudioSeal-style watermark in every output. The fine-tune keeps it.
pip install "drifting-tts[dacvae] @ git+https://github.com/kadirnar/drifting-tts"
from huggingface_hub import hf_hub_download
from drifting_tts.synthesize import Synthesizer
repo = "Vyvo/drifting-tts-tr-dacvae"
synth = Synthesizer(hf_hub_download(repo, "drifting_tts_dacvae_10k.pt"), "cuda",
vocoder=hf_hub_download(repo, "dacvae_decoder_ft.pt"))
wav, info = synth("Merhaba! Bu ses tek bir adımda üretildi.", speaker="studio", seed=0) # 24 kHz
facebook/dacvae-watermarked on first use; vocoder= replaces its decoder with the fine-tuned one.studio, male, female, or any integer speaker ID.Base model
facebook/dacvae-watermarked