File size: 8,982 Bytes
574b391
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9bde6c7
574b391
 
 
 
 
 
 
 
 
 
 
 
 
9bde6c7
574b391
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
---
language:
- en
license: apache-2.0
tags:
- text-to-speech
- tts
- audio-language-model
- voice-cloning
- zero-shot-tts
- neurovoice
- qwen2.5
- mimi-codec
- pytorch
- deep-learning
datasets:
- parler-tts/libritts_r_filtered
pipeline_tag: text-to-speech
inference: false
---

# 🎙️ NeuroVoice-0.5B (v0.1 Alpha)

<p align="center">
  <img src="https://img.shields.io/badge/Model-NeuroVoice--0.5B-7928CA?style=for-the-badge&logo=openai" alt="Model">
  <img src="https://img.shields.io/badge/Architecture-Two--Stage%20Audio%20LM-blue?style=for-the-badge" alt="Architecture">
  <img src="https://img.shields.io/badge/Backbone-Qwen2.5--0.5B-green?style=for-the-badge" alt="Backbone">
  <img src="https://img.shields.io/badge/Audio%20Codec-Kyutai%20Mimi%2024kHz-orange?style=for-the-badge" alt="Codec">
  <img src="https://img.shields.io/badge/Hardware-4x%20NVIDIA%20H100-red?style=for-the-badge" alt="Hardware">
  <img src="https://img.shields.io/badge/License-Apache%202.0-yellow?style=for-the-badge" alt="License">
</p>

**NeuroVoice-0.5B (v0.1 Alpha)** is an original, autoregressive **Neural Audio Language Model** built from the ground up for expressive text-to-speech and zero-shot voice cloning. It couples the deep linguistic reasoning of **Qwen2.5-0.5B** with the high-compression **Kyutai Mimi** 24 kHz neural audio codec (12.5 Hz frame rate, 8 codebooks).

The model was developed, optimized, and trained on **4x NVIDIA H100 GPUs** (DistributedDataParallel) using PyTorch FlashAttention-2 / SDPA, delivering over **99,800 tokens/sec** training throughput.

---

## 🏗️ NeuroVoice Architecture

NeuroVoice operates through a **Two-Stage Autoregressive Backbone + Depth Decoder** hierarchy:

```
[Text Prompt + Reference Audio]
             │
             ▼
┌─────────────────────────────────────────────────────────────┐
│ Stage 1: Main Audio Transformer (Qwen2.5-0.5B Backbone)    │
│  - 24 Transformer Layers (d_model=896, 14 Heads, 2 KV Heads)│
│  - SwiGLU MLP (d_ff=4864), RoPE, RMSNorm                    │
│  - Autoregressively predicts Codebook 0 (Semantic & Pitch)  │
└─────────────────────────────────────────────────────────────┘
             │
             ▼  (Audio Hidden States + CB0 Tokens)
┌─────────────────────────────────────────────────────────────┐
│ Stage 2: 4-Layer Depth Decoder (Acoustic Hierarchy)         │
│  - Hierarchically generates Codebooks 1..7                  │
│  - Reconstructs high-frequency acoustics & vocal timbre     │
└─────────────────────────────────────────────────────────────┘
             │
             ▼  (Full 8-Codebook Matrix @ 12.5 Hz)
┌─────────────────────────────────────────────────────────────┐
│ Kyutai Mimi Neural Audio Codec Decoder                      │
│  - Synthesizes 24 kHz Hi-Fi Audio Waveform                  │
└─────────────────────────────────────────────────────────────┘
```

* **Codec Details:** Kyutai Mimi compresses 24,000 samples/sec audio into **12.5 frames per second** with **8 codebooks** (vocabulary size 2048 per codebook), achieving an extreme 1920x temporal compression.
* **Stage 1 (Main Backbone):** Autoregressively predicts the speech backbone (Codebook 0).
* **Stage 2 (Depth Decoder):** Takes backbone hidden representations and CB0 to hierarchically generate the remaining 7 acoustic codebooks (CB1 to CB7) with soft acoustic sampling.

---

## 🍳 Training Recipe & Hyperparameters

Trained on MareNostrum 5 using PyTorch DistributedDataParallel (DDP):

| Hyperparameter | Value |
| :--- | :--- |
| **Compute Hardware** | 1 Node / 4x NVIDIA H100 (80GB HBM3) |
| **Dataset** | `parler-tts/libritts_r_filtered` (Clean subset, 21,000 samples, ~35.2 hours) |
| **Global Batch Size** | 48 (12 per GPU x 4 GPUs) |
| **Optimizer** | AdamW ($\beta_1=0.9, \beta_2=0.95, \text{weight\_decay}=0.1$) |
| **Learning Rate** | $2.0 \times 10^{-4}$ (Cosine Annealing schedule) |
| **Warmup Steps** | 200 steps (Linear warmup from $1.0 \times 10^{-6}$) |
| **Precision** | `bfloat16` Mixed Precision (`torch.autocast`) |
| **Attention Kernel** | PyTorch `F.scaled_dot_product_attention` (FlashAttention-2) |
| **Total Epochs / Steps** | 15 Epochs / 6,555 Steps |
| **Training Duration** | **23 minutes 52 seconds** |
| **Peak Throughput** | **~99,800 tokens / second** |

### 📉 Loss Trajectory Across Training

```
Step 0000: Total Loss: 25.1857 | CB0 Loss: 17.3926 | Depth Loss: 7.7931
Step 1000: Total Loss:  3.1959 | CB0 Loss:  0.6863 | Depth Loss: 2.5096
Step 3000: Total Loss:  2.2011 | CB0 Loss:  0.1136 | Depth Loss: 2.0875
Step 5000: Total Loss:  1.4281 | CB0 Loss:  0.0029 | Depth Loss: 1.4252
Step 6555: Total Loss:  1.0407 | CB0 Loss:  0.0034 | Depth Loss: 1.0372 (Best Epoch Avg: 1.1138)
```

---

## 🚀 Quickstart & Inference

### 1. Installation

```bash
pip install torch transformers soundfile torchaudio
```

### 2. Loading with Hugging Face `AutoModel` (Recommended)

Thanks to `AutoConfig` and `AutoModel` integration (`trust_remote_code=True`), you can load NeuroVoice directly in Python without cloning the repository:

```python
import torch
from transformers import AutoConfig, AutoModel

repo_id = "TurkishCodeMan/NeuroVoice-0.5B"

# 1. Konfigürasyonu ve Modeli Yükle (Safetensors / bfloat16)
config = AutoConfig.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModel.from_pretrained(repo_id, trust_remote_code=True)
model = model.to("cuda" if torch.cuda.is_available() else "cpu")
model.eval()

print(f"✅ {config.model_name} başarıyla yüklendi!")
print(f"Parametre Sayısı: {sum(p.numel() for p in model.parameters()) / 1e6:.2f}M")
```

---

### 3. Standalone CLI Inference (Cloned Repository)

You can also clone the repository and use the built-in `inference.py` script:

```bash
git clone https://huggingface.co/TurkishCodeMan/NeuroVoice-0.5B
cd NeuroVoice-0.5B
```

#### A. Standard Text-to-Speech
```bash
python inference.py \
    --use_qwen_backbone \
    --checkpoint model.safetensors \
    --text "Artificial intelligence is creating the future of speech synthesis." \
    --output output_tts.wav \
    --temperature 0.6 \
    --depth_temperature 0.6 \
    --max_tokens 80
```

#### B. Zero-Shot Voice Cloning
Use a 3 to 4-second clean audio sample (`.wav`) as reference:

```bash
python inference.py \
    --use_qwen_backbone \
    --checkpoint model.safetensors \
    --ref_audio reference_voice.wav \
    --ref_max_sec 3.5 \
    --text "Hello! This speech is synthesized using the NeuroVoice neural audio model." \
    --output output_cloned.wav \
    --temperature 0.5 \
    --depth_temperature 0.5 \
    --max_tokens 80
```

---

## 🔬 Dataset Analysis & Limitations

The v0.1 model was trained on **35.17 hours** of speech from the LibriTTS filtered clean subset:
* **Total Samples:** 21,000 audio-text pairs
* **Average Sentence Duration:** 6.03 seconds (18 words, ~4.2 codec frames per word)
* **Reference Audio Range in Training:** 2.0s to 4.0s (Average 3.71s)

### ⚠️ Known Limitations in v0.1 Alpha:
1. **Audiobook Monotony:** LibriTTS is narrated audiobook data; speech has a formal, studio-reader cadence rather than conversational prosody.
2. **Optimal Reference Audio:** The model was trained with 2-4 second reference clips. Use `--ref_max_sec 3.5` for optimal voice cloning quality.
3. **Acoustic Timbre:** Depth Decoder loss converged to ~1.0, which produces intelligible speech with a slightly synthetic timbre. Scaling up to larger multi-speaker datasets in v0.2 will address this.

---

## 🗺️ Roadmap (v0.2 & Beyond)

- [ ] **Scale Training Data:** Expand from 35 hours to 500+ hours with diverse conversational audio.
- [ ] **Enhanced Depth Decoder:** Increase Depth Decoder capacity from 4 to 8 layers for warmer human timbre.
- [ ] **Multilingual NeuroVoice:** Extend tokenization to Turkish and multilingual speech.

---

## 📜 Citation & License

This project is licensed under the **Apache 2.0 License**.

```bibtex
@misc{neurovoice-0.5b-2026,
  author = {TurkishCodeMan},
  title = {NeuroVoice-0.5B: An Autoregressive Neural Audio Language Model},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/TurkishCodeMan/NeuroVoice-0.5B}}
}
```