Instructions to use aitestinguserpg/whisper-small-gujarati-coreml with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- WhisperKit
How to use aitestinguserpg/whisper-small-gujarati-coreml with WhisperKit:
# Install CLI with Homebrew on macOS device brew install whisperkit-cli # View all available inference options whisperkit-cli transcribe --help # Download and run inference using whisper base model whisperkit-cli transcribe --audio-path /path/to/audio.mp3 # Or use your preferred model variant whisperkit-cli transcribe --model "large-v3" --model-prefix "distil" --audio-path /path/to/audio.mp3 --verbose
- Notebooks
- Google Colab
- Kaggle
Atlas Gujarati repair (Core ML)
A Core ML build of vasista22/whisper-gujarati-small, a
Whisper small fine-tuned on Gujarati speech. The Atlas app uses it for a
second listen: when English dictation has a word or two that sound like Gujarati and the English model was unsure of
them, Atlas plays just that stretch (at most 4 seconds) to this model, which writes it in Gujarati script, and Atlas then
spells it in Latin letters.
Atlas downloads and runs it on the Mac through WhisperKit. The audio never leaves the device.
Files
whisper-small-gujarati/
AudioEncoder.mlmodelc/ MelSpectrogram.mlmodelc/ TextDecoder.mlmodelc/
tokenizer.json, config.json, generation_config.json, …
About 470 MB in 16-bit precision.
How it was made
- Weights:
vasista22/whisper-gujarati-smallat revisionf08036e339f9d1a53b1e28f556e4b436d77f7762, unchanged. Its card reports 14.73% WER on the FLEURS Gujarati test set. - Config: the checkpoint's
generation_config.jsonpredates Whisper's language and task tokens, so the one fromopenai/whisper-small(same tokenizer) was used, for the language tokens and alignment heads. - Conversion: to Core ML with
whisperkittools. - Prompting: Atlas asks for Gujarati (
gu), no timestamps.
Limits
- It writes Gujarati script, and sometimes Devanagari; Atlas maps Devanagari to Gujarati letter for letter and marks those words as unsure.
- Whisper's output limit fills at around 11 seconds of Gujarati, so long passages get cut off. Atlas only sends short stretches and rejects a repair that was cut off.
- It is a Gujarati recogniser, not an English one: on English speech its output is not useful.
Licence and credit
Apache-2.0, as the base model. The weights are the work of Vasista Sai Lodagala
(whisper-finetune), fine-tuned from OpenAI's Whisper small (MIT).
Core ML conversion tooling is from Argmax (whisperkittools, WhisperKit). The only changes here are the format and the
generation config described above.
Model tree for aitestinguserpg/whisper-small-gujarati-coreml
Base model
vasista22/whisper-gujarati-small