You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen3-ASR Sorani for MLX

Sorani Kurdish speech recognition that runs on device on Apple silicon (Mac, iPhone and iPad), with punctuation.

This is rzgar/qwen3-asr-sorani-kurdish-ckb-v1, a version of Qwen/Qwen3-ASR-1.7B fine-tuned for Sorani (Central Kurdish), converted to Apple's MLX format. The language model weights are compressed to 6 bits, so the model is 2.05 GB instead of 4.09 GB, with the same accuracy.

All credit for the model goes to its authors. This repository only changes its file format.

At a glance

Detail Value
Language Sorani (Central Kurdish)
Version 6-bit: language model weights compressed from 16 to 6 bits. The audio encoder stays at full precision.
Size 2.05 GB (the original is 4.09 GB)
Format MLX safetensors
Runtime mlx-audio (Python) or mlx-audio-swift
Punctuation Yes
License Apache 2.0, the same as the original model

Accuracy

Word error rate (WER) and character error rate (CER) on 30 clips from the FLEURS Sorani test set, on a Mac with MLX. Lower is better.

Version Size WER CER
Original (16-bit) 4.09 GB 39.5% 10.0%
8-bit 2.48 GB 38.5% 9.7%
6-bit (this model) 2.05 GB 39.0% 9.9%
5-bit 1.83 GB 42.3% 10.8%
4-bit 1.60 GB 83.2% 29.9%

Below 6 bits the Sorani spelling degrades. At 4 bits the model writes many Sorani letters (ە، ێ، ۆ) in their Arabic or Persian forms, which is why there is no 4-bit version.

All versions used greedy decoding and the system prompt from the original model card (below).

Usage

pip install -U mlx-audio
from mlx_audio.stt.utils import load_model

SORANI_PROMPT = """You are an expert Sorani Kurdish (ckb) transcriptionist. 
Transcribe the audio into Sorani Kurdish with perfect orthography.
IMPORTANT: The speaker may use loanwords in Persian or Arabic. if you recognize any, 
please transcribe them in their standard writing form."""

model = load_model("UIDUser-NSB/Qwen3-ASR-Sorani-MLX")
result = model.generate("recording.wav", system_prompt=SORANI_PROMPT, temperature=0.0)
print(result.text)

The system prompt is the one the model was fine-tuned with; use it for the best Sorani spelling. Long recordings are best transcribed in chunks of about 30 seconds.

How it was converted

  1. Source: rzgar/qwen3-asr-sorani-kurdish-ckb-v1, revision d71490a623113b4b069ac07cfc85b409389dde4c.
  2. MLX: converted with mlx-audio 0.5.8: python -m mlx_audio.convert --hf-path <source> -q --q-bits 6 --q-group-size 64.
  3. Compression: affine 6-bit quantization in groups of 64 for the language model. The audio encoder is not quantized.

Notes

  • Strong on standard and formal Sorani (news, audiobooks, speeches). Heavy regional dialects and fast conversation are harder, as the original model card describes.
  • Tested on the FLEURS read-speech test set. Conversation, dialects and noisy audio may score differently.

Credits

Downloads last month
10
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for UIDUser-NSB/Qwen3-ASR-Sorani-MLX

Quantized
(1)
this model