GigaAM-Uzbek v1
Uzbek (Latin script) automatic speech recognition model — a full fine-tune of
GigaAM multilingual_ctc (220M,
char-level CTC) on 646 hours of supervised Uzbek speech.
Results (WER, greedy CTC)
| Benchmark | zero-shot 220M | zero-shot 600M | this model |
|---|---|---|---|
| UzbekVoice dev (read) | 5.08% | 4.15% | 2.18% |
| Common Voice uz dev | 8.50% | 6.42% | 8.05% |
| Common Voice uz test | 7.67% | 5.77% | 7.51% |
| FLEURS uz test* | 10.47% | 7.52% | 11.55% |
Apostrophe placement (oʻ/gʻ/tutuq — the Uzbek-specific failure mode): word-level precision/recall 0.967/0.967 on UzbekVoice dev; 0.894/0.890 on CV test.
*FLEURS caveat: 17% of reference rows (digit-containing) excluded — no number expansion in v1; treat FLEURS numbers as approximate.
Training data
| Corpus | Hours (post-filter) | License |
|---|---|---|
| DavronSherbaev/uzbekvoice-filtered | 591.6 | Apache-2.0 |
| yakhyo/mozilla-common-voice-uzbek (CV 17.0) | 52.6 | CC-0 |
Frozen eval sets: Common Voice uz test (18.0h), FLEURS uz test. Train rows whose normalized sentence appears in any dev/test set were removed (2,320 rows); Common Voice speaker (client_id) disjointness across splits verified.
Normalization
Single normalizer applied to all training and reference text
(source):
Unicode NFC → apostrophe unification (ʻ ʼ ’ ‘ ´ … → ASCII ') **before** any punctuation stripping → lowercase → GigaAMnormalize_raw_text` → charset
{space, ', a–z}. Rows with digits, Cyrillic, or duration outside 0.3–20s dropped.
Hyperparameters
Continue-finetune of multilingual_ctc (fixed 70-label vocab): lr 5e-5, 6 epochs,
batch 64, bf16-mixed, SpecAugment defaults, seed 42, checkpoint selection by WER on
a 3k-utterance mixed dev set (5 val checks/epoch). Best: epoch 4. Hardware:
1× NVIDIA B200. GigaAM commit 559d88d6.
A domain-balance ablation (Common Voice ×5 upsampling, +2 epochs at lr 2e-5) did not improve CV WER (7.97% vs 8.05% dev) — the CV gap appears capacity-bound, not sampling-bound (see 600M zero-shot).
Usage
import gigaam
model = gigaam.load_model("gigaam-uzbek-v1.ckpt")
print(model.transcribe("audio.wav")) # 16 kHz mono WAV
Limitations
- Read/prompted speech domains; no conversational training data in v1.
- No number expansion: digits are absent from output vocabulary scope.
- Latin script only.
- Out-of-domain (CV/FLEURS) performance near zero-shot level; v2 (in progress) adds USC (105h), FeruzaSpeech (52h), FLEURS-train and revisits capacity.
Training code, logs, experiment registry: https://github.com/SamariddinS/uzbek-voice