DACVAE-TTS Turkish v2

Turkish reference-conditioned text-to-speech in DACVAE latent space. This is the unchanged full-v2 checkpoint at 60,000 updates from VoiceHub/dacvae-tts-tr-combined, used by Vyvo/dacvae-tts-tr-v2-demo.

Files

  • model.pt: original PyTorch checkpoint with model and EMA weights, config and codec statistics.
  • duration/duration_trc.json: Turkish v2 duration predictor.
  • config.json: architecture and codec metadata for this release.
  • inference.json: settings used in the successful local inference checks.
  • provenance.json: source revisions and SHA-256 checksums.
  • verification.json and samples/: two generated examples and their ASR verification report.

The network has 67,442,052 parameters, width 512, 12 transformer layers, 8 attention heads and 128-dimensional codec latents. It uses character text units, RoPE, QK normalization, EDM prediction and quality conditioning. Output is 48 kHz mono audio. The facebook/dacvae-watermarked codec is downloaded separately.

Run inference

Code is in kadirnar/dacvae-tts-tr-v2. The code repository requires authorized GitHub access while it is private. Use Python 3.10–3.12 and a compatible PyTorch installation.

git clone https://github.com/kadirnar/dacvae-tts-tr-v2.git
cd dacvae-tts-tr-v2
uv venv --python 3.12
uv pip install -e .
.venv/bin/python inference.py \
  --text "Merhaba, bugün Türkçe konuşma sentezini test ediyoruz." \
  --output outputs/merhaba.wav

The wrapper downloads model.pt and the duration predictor from this repository, uses the bundled reference voice and writes PCM16 WAV plus inference metadata. For your own voice, pass --reference-audio reference.wav and --reference-text "The exact Turkish transcript of that reference recording.". See the code README for local model loading, device selection and reference examples.

This is a custom PyTorch checkpoint, not a Transformers AutoModel checkpoint. Use the DACVAE-TTS Synthesizer implementation included in the code repository. The dacvae-tr-opt/FastDiT implementation rejects its value_residual=False architecture.

Verified settings and results

The local checks used CUDA/BF16 on an NVIDIA GeForce RTX 5070 Ti, 32 sampling steps, guidance 5, sway -1, predictor duration, APG eta 0.5/momentum -0.3, quality target [4.6, 4.9, 4.4], seed 42. Turkish v2 normalization expands numbers, time expressions and currency. The wrapper also completes sentence punctuation and splits long text into chunks.

Input Duration Sample
Merhaba, bugün Türkçe konuşma sentezini test ediyoruz. 3.96 s Audio
Yarın saat 14:30'da buluşalım, toplam ücret 250 TL olacak. 5.28 s Audio

Both samples produced finite, non-silent audio and zero WER/CER after Turkish normalization when transcribed with openai/whisper-large-v3. These are two smoke-test examples, not a general quality benchmark. The full report is in verification.json.

Provenance and license

The source demo declares CC-BY-NC-4.0, retained for this model release. The original checkpoint is redistributed without changing its tensors. The codec retains its upstream license. Source revisions and file hashes are in provenance.json; code attribution and third-party notices are retained in the code repository.

Downloads last month
28
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train TTS-OPT/dacvae-tts-tr-v2

Space using TTS-OPT/dacvae-tts-tr-v2 1