GigaAM-Uzbek v1

Uzbek (Latin script) automatic speech recognition model — a full fine-tune of GigaAM multilingual_ctc (220M, char-level CTC) on 646 hours of supervised Uzbek speech.

Results (WER, greedy CTC)

Benchmark zero-shot 220M zero-shot 600M this model
UzbekVoice dev (read) 5.08% 4.15% 2.18%
Common Voice uz dev 8.50% 6.42% 8.05%
Common Voice uz test 7.67% 5.77% 7.51%
FLEURS uz test* 10.47% 7.52% 11.55%

Apostrophe placement (oʻ/gʻ/tutuq — the Uzbek-specific failure mode): word-level precision/recall 0.967/0.967 on UzbekVoice dev; 0.894/0.890 on CV test.

*FLEURS caveat: 17% of reference rows (digit-containing) excluded — no number expansion in v1; treat FLEURS numbers as approximate.

Training data

Corpus Hours (post-filter) License
DavronSherbaev/uzbekvoice-filtered 591.6 Apache-2.0
yakhyo/mozilla-common-voice-uzbek (CV 17.0) 52.6 CC-0

Frozen eval sets: Common Voice uz test (18.0h), FLEURS uz test. Train rows whose normalized sentence appears in any dev/test set were removed (2,320 rows); Common Voice speaker (client_id) disjointness across splits verified.

Normalization

Single normalizer applied to all training and reference text (source): Unicode NFC → apostrophe unification (ʻ ʼ ’ ‘ ´ … → ASCII ') **before** any punctuation stripping → lowercase → GigaAMnormalize_raw_text` → charset {space, ', a–z}. Rows with digits, Cyrillic, or duration outside 0.3–20s dropped.

Hyperparameters

Continue-finetune of multilingual_ctc (fixed 70-label vocab): lr 5e-5, 6 epochs, batch 64, bf16-mixed, SpecAugment defaults, seed 42, checkpoint selection by WER on a 3k-utterance mixed dev set (5 val checks/epoch). Best: epoch 4. Hardware: 1× NVIDIA B200. GigaAM commit 559d88d6.

A domain-balance ablation (Common Voice ×5 upsampling, +2 epochs at lr 2e-5) did not improve CV WER (7.97% vs 8.05% dev) — the CV gap appears capacity-bound, not sampling-bound (see 600M zero-shot).

Usage

import gigaam
model = gigaam.load_model("gigaam-uzbek-v1.ckpt")
print(model.transcribe("audio.wav"))  # 16 kHz mono WAV

Limitations

  • Read/prompted speech domains; no conversational training data in v1.
  • No number expansion: digits are absent from output vocabulary scope.
  • Latin script only.
  • Out-of-domain (CV/FLEURS) performance near zero-shot level; v2 (in progress) adds USC (105h), FeruzaSpeech (52h), FLEURS-train and revisits capacity.

Training code, logs, experiment registry: https://github.com/SamariddinS/uzbek-voice

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support