Automatic Speech Recognition
MLX
Safetensors
Chinese
English
audio8_asr_infinite
mlx-audio
speech-recognition
streaming
4-bit precision
Instructions to use cavi-ai/Audio8-ASR-Infinite-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use cavi-ai/Audio8-ASR-Infinite-MLX-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Audio8-ASR-Infinite-MLX-4bit cavi-ai/Audio8-ASR-Infinite-MLX-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Audio8-ASR-Infinite-MLX-4bit
MLX conversion of Edge0/Audio8-ASR-Infinite for Apple Silicon.
- Converted from the base model at revision
7476824bc222e4ad509d286e8cae8b8d3f371129. - Not affiliated with or endorsed by Edge0.
- Converted by Sasan Sotoodehfar, CAVI AI (https://cavi-ai.xyz).
Contents
model.safetensors: 3.74 GB (3.48 GiB).- Quantized to 4-bit (affine, group size 64): the text decoder (
language_model.*), including the tied token embedding. - Kept in bf16: Voxtral Realtime audio tower, multi-modal projector, frame-length embedding, semantic VAD heads,
ada_rms_normMLPs. audio8_asr_infinite/: MLX model code. mlx-audio 0.5.7 does not include this architecture; the snippet below registers it.
Requirements
- Apple Silicon Mac.
- Python 3.12.
pip install mlx-audio==0.5.7
Usage
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("cavi-ai/Audio8-ASR-Infinite-MLX-4bit")
sys.path.insert(0, path)
import audio8_asr_infinite
sys.modules["mlx_audio.stt.models.audio8_asr_infinite"] = audio8_asr_infinite
from mlx_audio.stt.utils import load_model
model = load_model(path)
print(model.generate("audio.wav", language="en").text)
generate(audio, language, transcription_delay_ms=480):audiois a file path or a 16 kHz mono array.language:"en"or"zh"; required.transcription_delay_ms: positive multiple of 80; default 480.- Decoding: streaming greedy decode, one text token per 80 ms step.
- Long audio: 30 s rolling decoder window.
Measured results
Hardware: Apple M5 Max.
| Test | 4-bit (this repo) | bf16 weights, same MLX code |
|---|---|---|
LibriSpeech validation-clean subset, hf-internal-testing/librispeech_asr_dummy (73 utterances, 1,150 words), WER |
8.35% | 7.30% |
| 96 s English recording, WER | 7.11% (decoded in 23.4 s) | |
| Mandarin recording | exact transcript | |
| Peak memory, 6 s clip | 4.2 GB |
- WER scoring: uppercase; punctuation removed except apostrophes; no number or spelling normalization.
- Port check: MLX code vs a PyTorch fp32 reference built from
transformersclasses produced identical transcripts on 4 clips.
License
- Model weights and configuration files: Apache-2.0, inherited from the base model. See
LICENSE. - Code in
audio8_asr_infinite/: MIT. Seeaudio8_asr_infinite/LICENSE.
Links
- Base model: https://huggingface.co/Edge0/Audio8-ASR-Infinite
- Upstream code: https://github.com/Edge0-AI/Audio8-ASR-Infinite
- MLX port source: https://github.com/cavi-ai/mlx-agent
- Downloads last month
- -
Model size
4B params
Tensor type
BF16
·
U32 ·
Hardware compatibility
Log In to add your hardware
4-bit
Model tree for cavi-ai/Audio8-ASR-Infinite-MLX-4bit
Base model
Edge0/Audio8-ASR-Infinite