You need to agree to share your contact information to access this model

Download access is reviewed by RegaLabs. Use and redistribution are governed by the Apache License 2.0.

Please provide your full name, organization/affiliation, and intended use case for accessing the Govtûgô Kurdish ASR model.

Log in or Sign Up to review the conditions and access this model content.

Govtûgô: Automated Multi-Speaker Kurdish ASR Model

سیستەمی خۆکاری دەنگ-بۆ-دەق بە زمانی کوردی (سۆرانی)

Model Summary

Govtûgô-ASR-Sorani is a state-of-the-art Automatic Speech Recognition (ASR) model fine-tuned for Central Kurdish (Sorani / سۆرانی), developed by RegaLabs. Built on the Qwen3-ASR architecture, the model is trained on more than 500+ hours of Kurdish speech datasets to deliver high-accuracy speech transcription across diverse acoustic environments, including multi-speaker meetings, spoken dialogue, clean studio recordings, and telephony audio.


Model Details

  • Developed by: RegaLabs
  • Architecture: Qwen3ASRForConditionalGeneration (24 encoder layers, 16 attention heads, 1024 d_model)
  • Processor Class: Qwen3ASRProcessor (128 mel bins, 16 kHz sampling rate, 50 window size)
  • Precision / Data Type: bfloat16
  • Training Dataset: 500+ hours of Central Kurdish speech
  • Format: SafeTensors standalone model weights
  • Language: Central Kurdish (Sorani / کوردیی ناوەندی / ckb) only

Intended Uses & Limitations

Primary Intended Uses

  • Multi-speaker meeting and dialogue transcription in Sorani Kurdish.
  • Media captioning and automated subtitle generation (SRT / VTT).
  • Kurdish voice dictation, conversational AI, and audio indexing.

Out-of-Scope & Limitations

  • Language Scope: This model is designed and trained strictly for Central Kurdish (Sorani / ckb) only. It is not intended for or trained on other languages or non-Sorani dialects (such as Kurmanji, Arabic, Persian, or English).
  • Extreme Overlapping Speech: Heavy multi-speaker overlap requires upstream speaker diarization (e.g., Pyannote) to partition speech turns prior to inference.

Quickstart / How to Use

The model is natively supported in Hugging Face transformers ($\ge$ 5.16.0).

import torch
import librosa
from transformers import AutoProcessor, Qwen3ASRForConditionalGeneration

model_id = "RegaLabs/Govtugo-ASR-Sorani"

# 1. Load processor and model
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3ASRForConditionalGeneration.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

# 2. Load audio and resample to 16 kHz
audio, sr = librosa.load("kurdish_sample.wav", sr=16000)

# 3. Apply conversation chat template
messages = [
    {"role": "system", "content": ""},
    {"role": "user", "content": [{"type": "audio", "audio": audio}]}
]
text_prompt = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)

# 4. Generate transcription
inputs = processor(text=text_prompt, audios=audio, return_tensors="pt", sampling_rate=16000)
inputs = {k: v.to(model.device) for k, v in inputs.items()}

with torch.no_grad():
    generated_ids = model.generate(**inputs, max_new_tokens=256)

# 5. Decode output
transcription = processor.batch_decode(
    generated_ids[:, inputs["input_ids"].shape[1]:],
    skip_special_tokens=True
)[0]

print("Transcription:", transcription)

Benchmarks

Evaluated 2026-10-06 (greedy decoding, max_new_tokens=256, the language Kurdish<asr_text> prefix stripped). Text normalization on reference and hypothesis: NFC, lowercase, diacritics/tatweel/bidi marks removed, Arabic→Kurdish letter forms unified, digits unified, punctuation removed. Corpus-level WER / CER.

Read and studio speech (50 random clips per set, 16 kHz)

Test set Clips WER % CER %
FLEURS ckb (female speakers) 50 25.53 5.11
TTS4All – fatih 50 11.99 1.11
TTS4All – giganet 50 12.62 0.90
TTS4All – shahen 50 6.70 1.01
KurFemTTS 50 10.63 1.49
Aran 50 14.87 1.88

Evaluation note: some read-speech sets may overlap the training data; FLEURS is the most reliable read-speech figure.


Access & Gating Policy

For a Hugging Face release, configure manual approval in the repository settings; the model-card fields alone do not enable gating. The intended access process is:

  1. Click Request Access above.
  2. Provide your name, organization/affiliation, and intended use case.
  3. Access requests are reviewed and approved by the RegaLabs team.

License & Attribution

Copyright 2026 RegaLabs, for RegaLabs contributions.

This model release is licensed under the Apache License 2.0. Preserve the license and applicable attribution notices when redistributing; identify modifications to files you change. See NOTICE, THIRD_PARTY_NOTICES.md, and CHANGES.md. Download approval does not add research-only or noncommercial restrictions to Apache-licensed material. RegaLabs and Govtûgô names and logos are not licensed for endorsement or branding; customary attribution remains permitted. Citation below is appreciated, rather than an additional license condition.

The license covers RegaLabs contributions to the model release and model card. Training datasets retain their own terms and are not distributed here. The accompanying Govtugo_Research_Proposal documents are separate project materials and are not covered by this model-release license.


Citation

@misc{govtugo2026kurdish_asr,
  title        = {Govtûgô: Automated Kurdish Speech-to-Text Model},
  author       = {{RegaLabs Team}},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/RegaLabs/Govtugo-ASR-Sorani}}
}
Downloads last month
2
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RegaLabs/Govtugo-ASR-Sorani

Finetuned
(1)
this model