DACVAE-TTS Turkish v2
Turkish reference-conditioned text-to-speech in DACVAE latent space. This is
the unchanged full-v2 checkpoint at 60,000 updates from
VoiceHub/dacvae-tts-tr-combined,
used by Vyvo/dacvae-tts-tr-v2-demo.
Files
model.pt: original PyTorch checkpoint with model and EMA weights, config and codec statistics.duration/duration_trc.json: Turkish v2 duration predictor.config.json: architecture and codec metadata for this release.inference.json: settings used in the successful local inference checks.provenance.json: source revisions and SHA-256 checksums.verification.jsonandsamples/: two generated examples and their ASR verification report.
The network has 67,442,052 parameters, width 512, 12 transformer layers,
8 attention heads and 128-dimensional codec latents. It uses character text units,
RoPE, QK normalization, EDM prediction and quality conditioning.
Output is 48 kHz mono audio. The facebook/dacvae-watermarked codec is downloaded separately.
Run inference
Code is in kadirnar/dacvae-tts-tr-v2. The code repository requires authorized GitHub access while it is private. Use Python 3.10–3.12 and a compatible PyTorch installation.
git clone https://github.com/kadirnar/dacvae-tts-tr-v2.git
cd dacvae-tts-tr-v2
uv venv --python 3.12
uv pip install -e .
.venv/bin/python inference.py \
--text "Merhaba, bugün Türkçe konuşma sentezini test ediyoruz." \
--output outputs/merhaba.wav
The wrapper downloads model.pt and the duration predictor from this repository,
uses the bundled reference voice and writes PCM16 WAV plus inference metadata.
For your own voice, pass --reference-audio reference.wav and
--reference-text "The exact Turkish transcript of that reference recording.".
See the code README for local model loading, device selection and reference examples.
This is a custom PyTorch checkpoint, not a Transformers AutoModel checkpoint.
Use the DACVAE-TTS Synthesizer implementation included in the code repository.
The dacvae-tr-opt/FastDiT implementation rejects its value_residual=False architecture.
Verified settings and results
The local checks used CUDA/BF16 on an NVIDIA GeForce RTX 5070 Ti,
32 sampling steps, guidance 5, sway -1, predictor duration,
APG eta 0.5/momentum -0.3, quality target [4.6, 4.9, 4.4], seed 42.
Turkish v2 normalization expands numbers, time expressions and currency.
The wrapper also completes sentence punctuation and splits long text into chunks.
| Input | Duration | Sample |
|---|---|---|
| Merhaba, bugün Türkçe konuşma sentezini test ediyoruz. | 3.96 s | Audio |
| Yarın saat 14:30'da buluşalım, toplam ücret 250 TL olacak. | 5.28 s | Audio |
Both samples produced finite, non-silent audio and zero WER/CER after Turkish
normalization when transcribed with openai/whisper-large-v3.
These are two smoke-test examples, not a general quality benchmark.
The full report is in verification.json.
Provenance and license
The source demo declares CC-BY-NC-4.0, retained for this model release.
The original checkpoint is redistributed without changing its tensors.
The codec retains its upstream license. Source revisions and file hashes are in
provenance.json; code attribution and third-party notices are retained in the code repository.
- Downloads last month
- 28