Drifting TTS Turkish: DAC-VAE latents (experimental)

Experimental. For the best quality, use the released mel model Vyvo/drifting-tts-tr (v3.1).

This is the one-step Turkish TTS of drifting-tts in the latent space of a pretrained audio VAE, the paper's latent setting (Deng et al., 2026).

  • Output. The DriftDiT generator writes DAC-VAE latents in one network evaluation (128 channels, 25 Hz).
  • Decoding. A DAC-VAE decoder fine-tuned on these latents turns them into audio.
  • Training. The model uses the same recipe and the same training data as v3.1. Only the target space differs.

Compare it with v3.1 and the VoxCPM2-latent model: Vyvo/drifting-tts-tr-compare.

Files

file what
drifting_tts_dacvae_10k.pt the latent TTS model, 10k steps
dacvae_decoder_ft.pt the DAC-VAE decoder fine-tuned for 40k steps on the model's generated latents, with the watermark frozen (below)

Results

Freya-TR-Eval, first 100 sentences, speaker 722, T = 0.3, α = 2. Whisper large-v3 on 8 kHz band-matched audio; UTMOSv2 and DNSMOS on the full band.

model decoder / vocoder WER CER UTMOSv2 DNSMOS OVRL
v3.1 (mels, released) BigVGAN-v2-ft 0.66% 0.14% 2.934 3.33
DAC-VAE latents (this) fine-tuned decoder 1.32% 0.30% 2.710 3.26
DAC-VAE latents released decoder 9.55% 2.91% 1.837 1.48
VoxCPM2 latents fine-tuned decoder 1.87% 0.43% 2.530 3.25
  • The released decoder sounds noisy on generated latents. The generator's latents are smoother than real ones, and the released DAC-VAE decoder renders that as noise.
  • Fine-tuning the decoder fixes it. Trained on the model's own latents, it brings WER from 9.55% to 1.32% and UTMOSv2 from 1.84 to 2.71.
  • Clean decoding is kept. On real latents of held-out recordings, UTMOSv2 goes from 2.72 to 2.70.
  • Ranking. This is the most natural of the latent models, but still behind v3.1.

Details: docs/LATENTS.md (#32).

Watermark

The DAC-VAE decoder embeds an AudioSeal-style watermark in every output. The fine-tune keeps it.

  • Frozen. Only the decoder's audio path was trained. All watermark weights are frozen, bit-identical to the release, and the watermark is added to every output exactly as by the released decoder.
  • Same watermark. The watermark of the fine-tuned decoder is the released watermark generator applied to the new audio, at the same level.
  • Not confirmed by a detector. Meta has not released a detector for this watermark yet, so detection could not be tested.

Usage

pip install "drifting-tts[dacvae] @ git+https://github.com/kadirnar/drifting-tts"
from huggingface_hub import hf_hub_download
from drifting_tts.synthesize import Synthesizer

repo = "Vyvo/drifting-tts-tr-dacvae"
synth = Synthesizer(hf_hub_download(repo, "drifting_tts_dacvae_10k.pt"), "cuda",
                    vocoder=hf_hub_download(repo, "dacvae_decoder_ft.pt"))
wav, info = synth("Merhaba! Bu ses tek bir adımda üretildi.", speaker="studio", seed=0)  # 24 kHz
  • Weights download. The released DAC-VAE weights come from facebook/dacvae-watermarked on first use; vocoder= replaces its decoder with the fine-tuned one.
  • Output rate. Use 24 kHz, the default: the decoder was fine-tuned against 24 kHz recordings.
  • Voices. The same as v3.1's: studio, male, female, or any integer speaker ID.

License

  • Weights in this repository: CC BY 4.0, as v3.1.
  • The fine-tuned decoder: a derivative of DAC-VAE (Meta, Apache-2.0).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vyvo/drifting-tts-tr-dacvae

Finetuned
(3)
this model

Space using Vyvo/drifting-tts-tr-dacvae 1

Paper for Vyvo/drifting-tts-tr-dacvae