File size: 8,982 Bytes
574b391 9bde6c7 574b391 9bde6c7 574b391 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 | ---
language:
- en
license: apache-2.0
tags:
- text-to-speech
- tts
- audio-language-model
- voice-cloning
- zero-shot-tts
- neurovoice
- qwen2.5
- mimi-codec
- pytorch
- deep-learning
datasets:
- parler-tts/libritts_r_filtered
pipeline_tag: text-to-speech
inference: false
---
# 🎙️ NeuroVoice-0.5B (v0.1 Alpha)
<p align="center">
<img src="https://img.shields.io/badge/Model-NeuroVoice--0.5B-7928CA?style=for-the-badge&logo=openai" alt="Model">
<img src="https://img.shields.io/badge/Architecture-Two--Stage%20Audio%20LM-blue?style=for-the-badge" alt="Architecture">
<img src="https://img.shields.io/badge/Backbone-Qwen2.5--0.5B-green?style=for-the-badge" alt="Backbone">
<img src="https://img.shields.io/badge/Audio%20Codec-Kyutai%20Mimi%2024kHz-orange?style=for-the-badge" alt="Codec">
<img src="https://img.shields.io/badge/Hardware-4x%20NVIDIA%20H100-red?style=for-the-badge" alt="Hardware">
<img src="https://img.shields.io/badge/License-Apache%202.0-yellow?style=for-the-badge" alt="License">
</p>
**NeuroVoice-0.5B (v0.1 Alpha)** is an original, autoregressive **Neural Audio Language Model** built from the ground up for expressive text-to-speech and zero-shot voice cloning. It couples the deep linguistic reasoning of **Qwen2.5-0.5B** with the high-compression **Kyutai Mimi** 24 kHz neural audio codec (12.5 Hz frame rate, 8 codebooks).
The model was developed, optimized, and trained on **4x NVIDIA H100 GPUs** (DistributedDataParallel) using PyTorch FlashAttention-2 / SDPA, delivering over **99,800 tokens/sec** training throughput.
---
## 🏗️ NeuroVoice Architecture
NeuroVoice operates through a **Two-Stage Autoregressive Backbone + Depth Decoder** hierarchy:
```
[Text Prompt + Reference Audio]
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Stage 1: Main Audio Transformer (Qwen2.5-0.5B Backbone) │
│ - 24 Transformer Layers (d_model=896, 14 Heads, 2 KV Heads)│
│ - SwiGLU MLP (d_ff=4864), RoPE, RMSNorm │
│ - Autoregressively predicts Codebook 0 (Semantic & Pitch) │
└─────────────────────────────────────────────────────────────┘
│
▼ (Audio Hidden States + CB0 Tokens)
┌─────────────────────────────────────────────────────────────┐
│ Stage 2: 4-Layer Depth Decoder (Acoustic Hierarchy) │
│ - Hierarchically generates Codebooks 1..7 │
│ - Reconstructs high-frequency acoustics & vocal timbre │
└─────────────────────────────────────────────────────────────┘
│
▼ (Full 8-Codebook Matrix @ 12.5 Hz)
┌─────────────────────────────────────────────────────────────┐
│ Kyutai Mimi Neural Audio Codec Decoder │
│ - Synthesizes 24 kHz Hi-Fi Audio Waveform │
└─────────────────────────────────────────────────────────────┘
```
* **Codec Details:** Kyutai Mimi compresses 24,000 samples/sec audio into **12.5 frames per second** with **8 codebooks** (vocabulary size 2048 per codebook), achieving an extreme 1920x temporal compression.
* **Stage 1 (Main Backbone):** Autoregressively predicts the speech backbone (Codebook 0).
* **Stage 2 (Depth Decoder):** Takes backbone hidden representations and CB0 to hierarchically generate the remaining 7 acoustic codebooks (CB1 to CB7) with soft acoustic sampling.
---
## 🍳 Training Recipe & Hyperparameters
Trained on MareNostrum 5 using PyTorch DistributedDataParallel (DDP):
| Hyperparameter | Value |
| :--- | :--- |
| **Compute Hardware** | 1 Node / 4x NVIDIA H100 (80GB HBM3) |
| **Dataset** | `parler-tts/libritts_r_filtered` (Clean subset, 21,000 samples, ~35.2 hours) |
| **Global Batch Size** | 48 (12 per GPU x 4 GPUs) |
| **Optimizer** | AdamW ($\beta_1=0.9, \beta_2=0.95, \text{weight\_decay}=0.1$) |
| **Learning Rate** | $2.0 \times 10^{-4}$ (Cosine Annealing schedule) |
| **Warmup Steps** | 200 steps (Linear warmup from $1.0 \times 10^{-6}$) |
| **Precision** | `bfloat16` Mixed Precision (`torch.autocast`) |
| **Attention Kernel** | PyTorch `F.scaled_dot_product_attention` (FlashAttention-2) |
| **Total Epochs / Steps** | 15 Epochs / 6,555 Steps |
| **Training Duration** | **23 minutes 52 seconds** |
| **Peak Throughput** | **~99,800 tokens / second** |
### 📉 Loss Trajectory Across Training
```
Step 0000: Total Loss: 25.1857 | CB0 Loss: 17.3926 | Depth Loss: 7.7931
Step 1000: Total Loss: 3.1959 | CB0 Loss: 0.6863 | Depth Loss: 2.5096
Step 3000: Total Loss: 2.2011 | CB0 Loss: 0.1136 | Depth Loss: 2.0875
Step 5000: Total Loss: 1.4281 | CB0 Loss: 0.0029 | Depth Loss: 1.4252
Step 6555: Total Loss: 1.0407 | CB0 Loss: 0.0034 | Depth Loss: 1.0372 (Best Epoch Avg: 1.1138)
```
---
## 🚀 Quickstart & Inference
### 1. Installation
```bash
pip install torch transformers soundfile torchaudio
```
### 2. Loading with Hugging Face `AutoModel` (Recommended)
Thanks to `AutoConfig` and `AutoModel` integration (`trust_remote_code=True`), you can load NeuroVoice directly in Python without cloning the repository:
```python
import torch
from transformers import AutoConfig, AutoModel
repo_id = "TurkishCodeMan/NeuroVoice-0.5B"
# 1. Konfigürasyonu ve Modeli Yükle (Safetensors / bfloat16)
config = AutoConfig.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModel.from_pretrained(repo_id, trust_remote_code=True)
model = model.to("cuda" if torch.cuda.is_available() else "cpu")
model.eval()
print(f"✅ {config.model_name} başarıyla yüklendi!")
print(f"Parametre Sayısı: {sum(p.numel() for p in model.parameters()) / 1e6:.2f}M")
```
---
### 3. Standalone CLI Inference (Cloned Repository)
You can also clone the repository and use the built-in `inference.py` script:
```bash
git clone https://huggingface.co/TurkishCodeMan/NeuroVoice-0.5B
cd NeuroVoice-0.5B
```
#### A. Standard Text-to-Speech
```bash
python inference.py \
--use_qwen_backbone \
--checkpoint model.safetensors \
--text "Artificial intelligence is creating the future of speech synthesis." \
--output output_tts.wav \
--temperature 0.6 \
--depth_temperature 0.6 \
--max_tokens 80
```
#### B. Zero-Shot Voice Cloning
Use a 3 to 4-second clean audio sample (`.wav`) as reference:
```bash
python inference.py \
--use_qwen_backbone \
--checkpoint model.safetensors \
--ref_audio reference_voice.wav \
--ref_max_sec 3.5 \
--text "Hello! This speech is synthesized using the NeuroVoice neural audio model." \
--output output_cloned.wav \
--temperature 0.5 \
--depth_temperature 0.5 \
--max_tokens 80
```
---
## 🔬 Dataset Analysis & Limitations
The v0.1 model was trained on **35.17 hours** of speech from the LibriTTS filtered clean subset:
* **Total Samples:** 21,000 audio-text pairs
* **Average Sentence Duration:** 6.03 seconds (18 words, ~4.2 codec frames per word)
* **Reference Audio Range in Training:** 2.0s to 4.0s (Average 3.71s)
### ⚠️ Known Limitations in v0.1 Alpha:
1. **Audiobook Monotony:** LibriTTS is narrated audiobook data; speech has a formal, studio-reader cadence rather than conversational prosody.
2. **Optimal Reference Audio:** The model was trained with 2-4 second reference clips. Use `--ref_max_sec 3.5` for optimal voice cloning quality.
3. **Acoustic Timbre:** Depth Decoder loss converged to ~1.0, which produces intelligible speech with a slightly synthetic timbre. Scaling up to larger multi-speaker datasets in v0.2 will address this.
---
## 🗺️ Roadmap (v0.2 & Beyond)
- [ ] **Scale Training Data:** Expand from 35 hours to 500+ hours with diverse conversational audio.
- [ ] **Enhanced Depth Decoder:** Increase Depth Decoder capacity from 4 to 8 layers for warmer human timbre.
- [ ] **Multilingual NeuroVoice:** Extend tokenization to Turkish and multilingual speech.
---
## 📜 Citation & License
This project is licensed under the **Apache 2.0 License**.
```bibtex
@misc{neurovoice-0.5b-2026,
author = {TurkishCodeMan},
title = {NeuroVoice-0.5B: An Autoregressive Neural Audio Language Model},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/TurkishCodeMan/NeuroVoice-0.5B}}
}
```
|