Instructions to use UIDUser-NSB/BuzzASR-Sorani-CoreML with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- WhisperKit
How to use UIDUser-NSB/BuzzASR-Sorani-CoreML with WhisperKit:
# Install CLI with Homebrew on macOS device brew install whisperkit-cli # View all available inference options whisperkit-cli transcribe --help # Download and run inference using whisper base model whisperkit-cli transcribe --audio-path /path/to/audio.mp3 # Or use your preferred model variant whisperkit-cli transcribe --model "large-v3" --model-prefix "distil" --audio-path /path/to/audio.mp3 --verbose
- Notebooks
- Google Colab
- Kaggle
BuzzASR Sorani for Core ML (WhisperKit)
Sorani Kurdish speech recognition that runs on device on iPhone, iPad and Mac.
This is BuzzASR/sorani-kurdish, a Whisper Large v3 model fine-tuned for Sorani (Central Kurdish), converted to Apple's Core ML format for WhisperKit. The weights are compressed to 4 bits, so the whole model is 898 MB instead of 3 GB, with the same accuracy.
All credit for the model goes to its authors. This repository only changes its file format.
At a glance
| Detail | Value |
|---|---|
| Language | Sorani (Central Kurdish) only. It does not transcribe English or Arabic. |
| Version | 4-bit: weights compressed from 16 to 4 bits, with the same accuracy as the original |
| Size | 898 MB (the original is 3.1 GB) |
| Format | Core ML, iOS 18 / macOS 15 or newer |
| Runtime | WhisperKit |
| License | MIT, the same as the original model |
Accuracy
Word error rate (WER) and character error rate (CER) on 30 clips from the FLEURS Sorani test set. Lower is better.
| Model | WER | CER |
|---|---|---|
| Original BuzzASR (PyTorch) | 34.2% | 9.0% |
| This model (Core ML, 4-bit) | 33.9% | 8.8% |
Both used greedy decoding. One clip, where the 4-bit model repeated its sentence, is left out of both rows; see Notes.
Usage
Download the buzzasr-sorani-ckb-4bit folder and load it with WhisperKit. Set the language to fa: the model was trained with that language token for Sorani.
import WhisperKit
let config = WhisperKitConfig(
modelFolder: "/path/to/buzzasr-sorani-ckb-4bit",
tokenizerFolder: URL(filePath: "/path/to/buzzasr-sorani-ckb-4bit"),
download: false
)
let whisper = try await WhisperKit(config)
var options = DecodingOptions()
options.language = "fa"
options.usePrefillPrompt = true
let result = try await whisper.transcribe(audioPath: "recording.wav", decodeOptions: options)
print(result.map(\.text).joined(separator: " "))
Files
The folder is a complete WhisperKit model:
| File | What it is |
|---|---|
MelSpectrogram.mlmodelc |
Turns audio into the spectrogram the model reads |
AudioEncoder.mlmodelc |
The audio encoder, 4-bit |
TextDecoder.mlmodelc |
The text decoder, 4-bit |
tokenizer.json, tokenizer_config.json |
The model's own Sorani tokenizer |
config.json, generation_config.json |
Model settings |
How it was converted
- Source: BuzzASR/sorani-kurdish, revision
ce7e6a0f4d28c2f6a75815d2c9d79e0c81e917bf. - Core ML: converted with whisperkittools 0.4.2 for iOS 18 / macOS 15.
- Compression: weights palettized to 4 bits with coremltools (k-means, one table per 16 channels).
- Tokenizer:
tokenizer.jsonwas built from the model'svocab.json,merges.txtandadded_tokens.json. BuzzASR'svocab.jsonalso lists every special token (such as<|endoftext|>) at ids 0 to 1608, while the model uses the copies at 51865 and up. The low copies are renamed<|unused_N|>, so each special token has one id. Text encodes exactly as with the original tokenizer. - Word timestamps:
generation_config.jsonhas thealignment_headsofopenai/whisper-large-v3, the base model, which word timestamps need.
Notes
- Like other Whisper models, it can repeat a sentence on rare clips. That happened on 1 of 30 test clips with the 4-bit model. WhisperKit's temperature fallback, which is on by default, retries such output.
- It was tested on the FLEURS read-speech test set. Conversation, dialects and noisy audio may score differently.
Citation
Please cite the original BuzzASR paper:
@misc{buzzasr2026,
title = {BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models},
author = {Shivam Singh and Aditya Yadavalli and Catherine Arnett and Alex Warstadt},
year = {2026},
eprint = {2609.09554},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.09554}
}
- Downloads last month
- -