How to use from the
Use from the
WhisperKit library
# Gated model: Login with a HF token with gated access permission
hf auth login
# Install CLI with Homebrew on macOS device
brew install whisperkit-cli

# View all available inference options
whisperkit-cli transcribe --help

# Download and run inference using whisper base model
whisperkit-cli transcribe --audio-path /path/to/audio.mp3

# Or use your preferred model variant
whisperkit-cli transcribe --model "large-v3" --model-prefix "distil" --audio-path /path/to/audio.mp3 --verbose

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

BuzzASR Sorani for Core ML (WhisperKit)

Sorani Kurdish speech recognition that runs on device on iPhone, iPad and Mac.

This is BuzzASR/sorani-kurdish, a Whisper Large v3 model fine-tuned for Sorani (Central Kurdish), converted to Apple's Core ML format for WhisperKit. The weights are compressed to 4 bits, so the whole model is 898 MB instead of 3 GB, with the same accuracy.

All credit for the model goes to its authors. This repository only changes its file format.

At a glance

Detail Value
Language Sorani (Central Kurdish) only. It does not transcribe English or Arabic.
Version 4-bit: weights compressed from 16 to 4 bits, with the same accuracy as the original
Size 898 MB (the original is 3.1 GB)
Format Core ML, iOS 18 / macOS 15 or newer
Runtime WhisperKit
License MIT, the same as the original model

Accuracy

Word error rate (WER) and character error rate (CER) on 30 clips from the FLEURS Sorani test set. Lower is better.

Model WER CER
Original BuzzASR (PyTorch) 34.2% 9.0%
This model (Core ML, 4-bit) 33.9% 8.8%

Both used greedy decoding. One clip, where the 4-bit model repeated its sentence, is left out of both rows; see Notes.

Usage

Download the buzzasr-sorani-ckb-4bit folder and load it with WhisperKit. Set the language to fa: the model was trained with that language token for Sorani.

import WhisperKit

let config = WhisperKitConfig(
    modelFolder: "/path/to/buzzasr-sorani-ckb-4bit",
    tokenizerFolder: URL(filePath: "/path/to/buzzasr-sorani-ckb-4bit"),
    download: false
)
let whisper = try await WhisperKit(config)

var options = DecodingOptions()
options.language = "fa"
options.usePrefillPrompt = true
let result = try await whisper.transcribe(audioPath: "recording.wav", decodeOptions: options)
print(result.map(\.text).joined(separator: " "))

Files

The folder is a complete WhisperKit model:

File What it is
MelSpectrogram.mlmodelc Turns audio into the spectrogram the model reads
AudioEncoder.mlmodelc The audio encoder, 4-bit
TextDecoder.mlmodelc The text decoder, 4-bit
tokenizer.json, tokenizer_config.json The model's own Sorani tokenizer
config.json, generation_config.json Model settings

How it was converted

  1. Source: BuzzASR/sorani-kurdish, revision ce7e6a0f4d28c2f6a75815d2c9d79e0c81e917bf.
  2. Core ML: converted with whisperkittools 0.4.2 for iOS 18 / macOS 15.
  3. Compression: weights palettized to 4 bits with coremltools (k-means, one table per 16 channels).
  4. Tokenizer: tokenizer.json was built from the model's vocab.json, merges.txt and added_tokens.json. BuzzASR's vocab.json also lists every special token (such as <|endoftext|>) at ids 0 to 1608, while the model uses the copies at 51865 and up. The low copies are renamed <|unused_N|>, so each special token has one id. Text encodes exactly as with the original tokenizer.
  5. Word timestamps: generation_config.json has the alignment_heads of openai/whisper-large-v3, the base model, which word timestamps need.

Notes

  • Like other Whisper models, it can repeat a sentence on rare clips. That happened on 1 of 30 test clips with the 4-bit model. WhisperKit's temperature fallback, which is on by default, retries such output.
  • It was tested on the FLEURS read-speech test set. Conversation, dialects and noisy audio may score differently.

Citation

Please cite the original BuzzASR paper:

@misc{buzzasr2026,
  title         = {BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models},
  author        = {Shivam Singh and Aditya Yadavalli and Catherine Arnett and Alex Warstadt},
  year          = {2026},
  eprint        = {2609.09554},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2609.09554}
}
Downloads last month
26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for UIDUser-NSB/BuzzASR-Sorani-CoreML

Finetuned
(1)
this model

Paper for UIDUser-NSB/BuzzASR-Sorani-CoreML