UIDUser-NSB's picture
BuzzASR Sorani for Core ML (WhisperKit), 4-bit
36db787
|
Raw History Blame Contribute Delete
4.27 kB
---
license: mit
language:
- ckb
library_name: whisperkit
pipeline_tag: automatic-speech-recognition
tags:
- whisperkit
- coreml
- whisper
- sorani
- central-kurdish
- kurdish
base_model:
- BuzzASR/sorani-kurdish
---
# BuzzASR Sorani for Core ML (WhisperKit)
Sorani Kurdish speech recognition that runs **on device** on iPhone, iPad and Mac.
This is [BuzzASR/sorani-kurdish](https://huggingface.co/BuzzASR/sorani-kurdish), a Whisper Large v3 model fine-tuned for Sorani (Central Kurdish), converted to Apple's Core ML format for [WhisperKit](https://github.com/argmaxinc/argmax-oss-swift). The weights are compressed to 4 bits, so the whole model is **898 MB** instead of 3 GB, with the same accuracy.
All credit for the model goes to its authors. This repository only changes its file format.
## At a glance
| Detail | Value |
|---|---|
| Language | Sorani (Central Kurdish) only. It does not transcribe English or Arabic. |
| Version | 4-bit: weights compressed from 16 to 4 bits, with the same accuracy as the original |
| Size | 898 MB (the original is 3.1 GB) |
| Format | Core ML, iOS 18 / macOS 15 or newer |
| Runtime | WhisperKit |
| License | MIT, the same as the original model |
## Accuracy
Word error rate (WER) and character error rate (CER) on 30 clips from the FLEURS Sorani test set. Lower is better.
| Model | WER | CER |
|---|---|---|
| Original BuzzASR (PyTorch) | 34.2% | 9.0% |
| **This model (Core ML, 4-bit)** | **33.9%** | **8.8%** |
Both used greedy decoding. One clip, where the 4-bit model repeated its sentence, is left out of both rows; see Notes.
## Usage
Download the `buzzasr-sorani-ckb-4bit` folder and load it with WhisperKit. Set the language to `fa`: the model was trained with that language token for Sorani.
```swift
import WhisperKit
let config = WhisperKitConfig(
modelFolder: "/path/to/buzzasr-sorani-ckb-4bit",
tokenizerFolder: URL(filePath: "/path/to/buzzasr-sorani-ckb-4bit"),
download: false
)
let whisper = try await WhisperKit(config)
var options = DecodingOptions()
options.language = "fa"
options.usePrefillPrompt = true
let result = try await whisper.transcribe(audioPath: "recording.wav", decodeOptions: options)
print(result.map(\.text).joined(separator: " "))
```
## Files
The folder is a complete WhisperKit model:
| File | What it is |
|---|---|
| `MelSpectrogram.mlmodelc` | Turns audio into the spectrogram the model reads |
| `AudioEncoder.mlmodelc` | The audio encoder, 4-bit |
| `TextDecoder.mlmodelc` | The text decoder, 4-bit |
| `tokenizer.json`, `tokenizer_config.json` | The model's own Sorani tokenizer |
| `config.json`, `generation_config.json` | Model settings |
## How it was converted
1. **Source:** BuzzASR/sorani-kurdish, revision `ce7e6a0f4d28c2f6a75815d2c9d79e0c81e917bf`.
2. **Core ML:** converted with whisperkittools 0.4.2 for iOS 18 / macOS 15.
3. **Compression:** weights palettized to 4 bits with coremltools (k-means, one table per 16 channels).
4. **Tokenizer:** `tokenizer.json` was built from the model's `vocab.json`, `merges.txt` and `added_tokens.json`. BuzzASR's `vocab.json` also lists every special token (such as `<|endoftext|>`) at ids 0 to 1608, while the model uses the copies at 51865 and up. The low copies are renamed `<|unused_N|>`, so each special token has one id. Text encodes exactly as with the original tokenizer.
5. **Word timestamps:** `generation_config.json` has the `alignment_heads` of `openai/whisper-large-v3`, the base model, which word timestamps need.
## Notes
- Like other Whisper models, it can repeat a sentence on rare clips. That happened on 1 of 30 test clips with the 4-bit model. WhisperKit's temperature fallback, which is on by default, retries such output.
- It was tested on the FLEURS read-speech test set. Conversation, dialects and noisy audio may score differently.
## Citation
Please cite the original BuzzASR paper:
```bibtex
@misc{buzzasr2026,
title = {BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models},
author = {Shivam Singh and Aditya Yadavalli and Catherine Arnett and Alex Warstadt},
year = {2026},
eprint = {2609.09554},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.09554}
}
```