Audar-ASR-V1-Turbo · GGUF

Audar's Arabic-first speech-recognition model — leaderboard-grade, dialect-aware.

From Arabic to the world.

License Task Format Params Open-AR-ASR Avg CER Emirati GitHub

🧭 Overview · 📊 Benchmarks · 💻 GGUF Deploy · 🤗 Transformers · 🎙️ Streaming · ⚡ vLLM · 📄 Tech Report · 🐙 GitHub · ☁️ Audar API · 📜 License


🧭 What it is

Audar-ASR-V1-Turbo is an Arabic-first generative speech-recognition model — the accuracy tier of the Audar-ASR family. It recasts transcription as audio-conditioned next-token prediction over a unified text vocabulary (a language-model decoder rather than a CTC or transducer objective), and is built on a permissively-licensed open-weight audio-LLM foundation and adapted in-house — the contribution is the adaptation (the data curriculum and the alignment rubric), not the foundation:

  • 🧱 Large-scale bilingual pretraining — 300,000+ hours of labeled audio, primarily Arabic and English, spanning MSA, Gulf, Egyptian, Levantine and Maghrebi speech, code-switching, and diverse acoustic channels.
  • 🎯 Dialect-targeted fine-tuning — hardness sampling and multi-task conditioning focused on proper nouns, code-switching, and dialect-faithful orthography.
  • 🧠 KTO preference alignment — Kahneman-Tversky Optimization on accented dialectal Arabic, with unpaired binary-desirability labels from trained native annotators across the Gulf, Levantine, Egyptian, and Maghrebi dialects, along five axes: verbatim accuracy, diacritic correctness, code-switch handling, named-entity preservation, and output formatting.

The result is state-of-the-art dialectal Arabic ASR — the lowest average WER and CER of any evaluated system on the Open Universal Arabic ASR Leaderboard. It transcribes MSA and every major Arabic dialect, code-switched Arabic–English, and English, across 30 languages in total.

Built on a permissively-licensed open-weight audio-LLM foundation; the adaptation, data, and alignment are Audar's. Full method and results: Audar-ASR-V1 Technical Report.

Model summary

ModelAudar-ASR-V1-Turbo — Arabic-first generative ASR (accuracy tier)
TaskAutomatic speech recognition (audio → text)
ApproachGenerative ASR — audio encoder + language-model decoder (audio-conditioned next-token prediction)
Trainingbuilt on an open-weight audio-LLM foundation; adapted via a 4-stage curriculum — 300k+ hrs bilingual pretraining → multi-task fine-tuning → dialect PEFT → KTO alignment
Decoder parameters2,031,739,904 (2.03B)
Audio encoder parameters317,477,504 (0.32B)
Total parameters2,349,217,408 (2.35B, bf16)
Audio input16 kHz mono; 30 s context (longer audio is chunked/streamed)
LanguagesArabic (MSA + Gulf/Egyptian/Levantine/Maghrebi dialects) + English + 28 more
RuntimeGGUF / llama.cpp — CPU · GPU · edge
LicenseAudarAI Community License v1.0

📊 Benchmarks

Arabic dialectal ASR is hard — heavily dialectal, conversational, code-switched speech is the frontier for every system. On the Open Universal Arabic ASR Leaderboard, Audar-ASR-V1-Turbo ranks #1 of 35 systems with the lowest average WER (24.8 %) and the lowest average CER (9.5 %) of any model evaluated — ahead of the next-best system by 1.1 pp WER — and it is the single best system on SADA and MGB-2.

Open Universal Arabic ASR Leaderboard — full standings

Per-dataset WER % across all six leaderboard test sets, plus the two composite averages. Lower is better; Avg WER is the ranking metric. Numbers are our end-to-end re-run on the full test sets with the leaderboard-equivalent normalizer — calibrated to the public leaderboard within 0.03 pp (bit-exact on 4 of 6 sets for Qwen3-ASR-1.7B), so every row is directly comparable. Ours in bold.

# Model Avg WER Avg CER SADA CV-18 MASC-clean MASC-noisy MGB-2 Casablanca
1 Audar-ASR-V1-Turbo (Ours) 24.78 9.49 29.41 8.60 19.60 28.35 11.13 51.58
2 CohereLabs/cohere-transcribe-arabic-07-2026 25.87 11.80 37.47 5.82 19.60 27.07 15.54 49.71
3 omnilingual-asr/omniASR_LLM_7B 28.32 12.52 41.61 8.75 19.69 29.29 14.13 56.46
4 omnilingual-asr/omniASR_LLM_3B 29.96 13.77 46.18 9.15 19.90 30.03 14.22 60.27
5 omnilingual-asr/omniASR_LLM_1B 29.96 13.40 43.84 9.55 20.03 30.26 15.34 60.68
6 CohereLabs/cohere-transcribe-03-2026 30.67 16.37 60.11 8.17 8.66 19.01 25.33 62.71
7 Qwen/Qwen3-Omni-30B-A3B-Instruct 30.71 13.67 44.82 11.46 21.47 30.85 13.09 62.55
8 nvidia-conformer-ctc-large-arabic (lm) 32.91 13.84 44.52 8.80 23.74 34.29 17.20 68.90
9 omnilingual-asr/omniASR_LLM_300M 32.96 14.84 51.38 12.03 20.66 32.45 16.58 64.64
10 google/gemma-4-E4B-it 32.98 13.71 43.40 19.65 24.86 33.59 17.72 58.63
11 Qwen/Qwen3-ASR-1.7B 33.36 12.33 45.53 16.90 24.37 34.29 16.57 64.47
12 mistralai/Voxtral-Small-24B-2507 34.47 15.29 50.82 15.25 23.96 34.43 16.03 66.30
13 nvidia-conformer-ctc-large-arabic (greedy) 34.74 13.37 47.26 10.60 24.12 35.64 19.69 71.13
14 google/gemma-4-E2B-it 35.87 15.34 46.23 23.76 27.47 36.15 20.72 60.87
15 openai/whisper-large-v3 36.86 17.21 55.96 17.83 24.66 34.63 16.26 71.81
16 omnilingual-asr/omniASR_CTC_3B 37.78 19.79 69.85 14.19 21.48 34.60 18.96 67.58
17 omnilingual-asr/omniASR_CTC_7B 38.12 20.91 72.69 12.47 21.08 35.04 20.43 67.02
18 facebook/seamless-m4t-v2-large 38.16 17.03 62.52 21.70 25.04 33.24 20.23 66.25
19 omnilingual-asr/omniASR_CTC_1B 39.29 20.47 71.42 17.55 22.76 35.73 19.96 68.32
20 openai/whisper-large-v3-turbo 40.05 18.87 60.36 25.73 25.51 37.16 17.75 73.79
21 openai/whisper-large-v2 40.20 19.55 57.46 21.77 27.25 38.55 25.17 71.01
22 Qwen/Qwen3-ASR-0.6B 42.19 16.23 53.75 28.28 31.34 42.63 25.45 71.68
23 openai/whisper-large 42.57 20.49 63.24 26.04 28.89 40.79 24.28 72.18
24 mistralai/Voxtral-Mini-3B-2507 42.58 19.90 63.65 22.12 28.37 41.27 22.56 77.52
25 asafaya/hubert-large-arabic-transcribe 45.50 17.35 67.82 8.01 32.94 50.16 37.51 76.53
26 openai/whisper-medium 45.57 22.27 67.71 28.07 29.99 42.91 29.32 75.44
27 nvidia-Parakeet-ctc-1.1b-concat 46.54 23.88 70.70 26.34 30.49 45.95 24.94 80.80
28 omnilingual-asr/omniASR_CTC_300M 46.65 21.86 78.11 27.90 28.40 43.26 26.85 75.35
29 nvidia-Parakeet-ctc-1.1b-universal 51.96 25.19 73.58 40.01 36.16 50.03 30.68 81.30
30 microsoft/VibeVoice-ASR 52.99 28.95 69.83 44.25 32.95 52.43 25.10 93.37
31 facebook/mms-1b-all 54.54 21.45 77.48 26.52 38.82 57.33 39.16 87.95
32 openai/whisper-small 55.13 21.68 78.02 24.18 35.93 56.36 48.64 87.64
33 whitefox123/w2v-bert-2.0-arabic-4 58.13 27.62 87.34 41.79 37.82 53.28 40.66 87.88
34 jonatasgrosman/wav2vec2-large-xlsr-53-arabic 60.98 25.61 86.82 23.00 42.75 64.27 56.29 92.72
35 speechbrain/asr-wav2vec2-commonvoice-14-ar 65.74 30.93 88.54 29.17 49.10 69.57 64.37 93.68

Bold = best in column. Audar-ASR-V1-Turbo owns both composite averages and leads on SADA and MGB-2; the recent Cohere and OmniASR systems are the closest competitors, each strongest on a subset of the conversational and clean-read sets. Casablanca (Moroccan Darija) is the hardest set for every system.

Emirati Arabic

Set WER % CER %
Emirati (Mixat, full 1,585-clip test) 19.4 7.3

On Emirati, the real recognition error is ≈ 7.3 % — near-parity with spontaneous English — while the residual up to 19.4 % WER is largely orthographic convention (near-miss spelling of the same word, e.g. انتو↔انتوا, and Latin-vs-Arabic rendering of English loanwords), not misrecognition.

💻 GGUF inference (llama.cpp)

Turbo runs on llama.cpp via the multimodal (mtmd) path — a quantized decoder GGUF plus a BF16 audio projector (mmproj). Build a recent llama.cpp (with Qwen3-ASR support), then:

./llama-mtmd-cli \
  -m       Audar-ASR-V1-Turbo-Q8_0.gguf \
  --mmproj mmproj-Audar-ASR-V1-Turbo.gguf \
  --audio  clip.wav \
  -sys     "فرّغ الكلام العربي التالي." \
  --temp 0

⚠️ The audio projector (mmproj) must stay BF16 (its ClippableLinear is numerically sensitive). The decoder quantizes normally.

Prefer a managed endpoint? The Audar-ASR family is also available via the Audar API/SDK — streaming, speaker-attributed transcription, and diarization, production-hosted.

GGUF variants

File Approx. size Notes
Audar-ASR-V1-Turbo-Q4_K_M.gguf ~1.28 GB Smallest; constrained hardware
Audar-ASR-V1-Turbo-Q8_0.gguf ~2.16 GB Near-lossless (recommended)
Audar-ASR-V1-Turbo.gguf (BF16) ~4.07 GB Full precision decoder
mmproj-Audar-ASR-V1-Turbo.gguf ~0.64 GB BF16 audio encoder — required, keep BF16

🤗 Transformers (full-precision safetensors)

The full-precision bf16 weights are published at the repo root — the reference checkpoint the GGUF and W4A16 builds are derived from (2,349,217,408 params, safetensors). Standard 🤗 Transformers, loaded with trust_remote_code=True (the repo ships the self-contained Qwen3-ASR code).

# pip install "transformers==4.57.6" torch librosa
import torch, librosa
from transformers import AutoProcessor, AutoModelForCausalLM

repo  = "audarai/Audar-ASR-V1-Turbo"
proc  = AutoProcessor.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo, trust_remote_code=True,
    dtype=torch.bfloat16, device_map="cuda:0",
).eval()

SYSTEM = "فرّغ الكلام العربي التالي."          # "Transcribe the following Arabic speech."
audio, _ = librosa.load("clip.wav", sr=16000, mono=True)

conv = [{"role": "system", "content": SYSTEM},
        {"role": "user",   "content": [{"type": "audio"}]}]     # audio placeholder (a list, not "<audio>")
text   = proc.apply_chat_template(conv, tokenize=False, add_generation_prompt=True)
inputs = proc(text=text, audio=audio, sampling_rate=16000, return_tensors="pt").to(model.device)
inputs["input_features"] = inputs["input_features"].to(model.dtype)   # features are fp32 -> cast to bf16

out = model.generate(**inputs, max_new_tokens=440, do_sample=False)
print(proc.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0].strip())

The self-contained modeling code targets transformers==4.57.6 (the version this checkpoint was built and validated with). For version-independent, high-throughput serving, prefer vLLM — it implements Qwen3-ASR natively (no custom code); see below.

File (repo root) Approx. size Notes
model.safetensors ~4.7 GB Full bf16 weights (2,349,217,408 params)
config.json · *_audar_asr.py · __init__.py Config + self-contained Qwen3-ASR modeling code
tokenizer files · preprocessor_config.json Qwen3 tokenizer + Whisper-mel feature extractor

⚡ vLLM inference (GPU serving)

Turbo also runs on vLLM for high-throughput GPU serving with an OpenAI-compatible API. vLLM implements the Qwen3-ASR architecture natively (Qwen3ASRForConditionalGeneration + Qwen3ASRRealtimeGeneration) — no custom serving code: point vLLM at a checkpoint and it exposes /v1/chat/completions, /v1/audio/transcriptions, and a realtime /v1/realtime WebSocket.

vLLM serves quantized compressed-tensors checkpoints (not the GGUF files — those are for llama.cpp). A vLLM-ready 4-bit (W4A16) build is provided in the vllm-w4a16/ folder:

Build Folder Size Decoder Audio encoder / lm_head / embeddings Accuracy
W4A16 vllm-w4a16 ~2.6 GB INT4 (group-128) BF16 (kept) ~+1 pp CER vs BF16

Only the language-model decoder is quantized; the audio encoder + projector stay BF16 (the projector's ClippableLinear is numerically sensitive — the same rule as the GGUF mmproj), as do lm_head and the token embeddings. An FP8 build (lossless vs BF16, ~3.3 GB) can be produced with the same recipe — see the note at the end.

For full-precision GPU serving, point vLLM at the repo (the full bf16 root weights) instead of the 4-bit build — same native Qwen3-ASR support, no quantization.

1. Install (audio support required)

vLLM needs the audio extras (PyAV + librosa + soundfile) to decode audio; the stock image does not ship them:

FROM vllm/vllm-openai:v0.24.0
RUN pip install --no-cache-dir av librosa soundfile
docker build -t vllm-audio:0.24.0 .

(Or in a plain environment: pip install "vllm>=0.24" av librosa soundfile.)

2. Get the weights & serve

# download just the vLLM build
hf download audarai/Audar-ASR-V1-Turbo --include "vllm-w4a16/*" --local-dir ./turbo

docker run -d --name audar-asr --gpus '"device=0"' \
  -v $PWD/turbo/vllm-w4a16:/model:ro -p 8000:8000 \
  vllm-audio:0.24.0 \
  --model /model --served-model-name audar-asr-v1-turbo \
  --trust-remote-code --max-model-len 8192 --gpu-memory-utilization 0.4

vLLM auto-detects the compressed-tensors quantization (Marlin INT4 kernel). Weights + KV cache fit on any ≥12 GB GPU.

3. Transcribe

Turbo is prompt-steerable: the system message sets the task/language. For Arabic use فرّغ الكلام العربي التالي.; steer other languages with the equivalent instruction. Send 16 kHz mono audio as base64 input_audio and decode greedily (temperature: 0).

import base64, requests

audio = base64.b64encode(open("clip.wav", "rb").read()).decode()   # 16 kHz mono wav
r = requests.post("http://localhost:8000/v1/chat/completions", json={
    "model": "audar-asr-v1-turbo",
    "temperature": 0,
    "max_tokens": 320,
    "messages": [
        {"role": "system", "content": "فرّغ الكلام العربي التالي."},
        {"role": "user", "content": [
            {"type": "input_audio", "input_audio": {"data": audio, "format": "wav"}}
        ]},
    ],
})
print(r.json()["choices"][0]["message"]["content"])

The OpenAI-style POST /v1/audio/transcriptions (multipart file upload) endpoint is also available for Whisper-style clients.

4. Accuracy (FLEURS Arabic, greedy)

Character Error Rate vs the BF16 source — CER is the stable cross-precision metric for Arabic, where minor و-segmentation differences inflate WER without changing the characters:

Build AR CER Δ vs BF16
BF16 source 2.46 %
W4A16 (this build) 3.73 % +1.27
FP8 (optional) 2.46 % +0.00 (lossless)

Leaderboard-grade full-test-set numbers are in the Benchmarks section above; 4-bit quantization keeps them within ~1 pp CER (FP8 keeps them exactly).

Notes

  • Realtime streaming: vLLM also registers Qwen3ASRRealtimeGeneration, exposing an OpenAI-Realtime-compatible /v1/realtime WebSocket; pair it with VAD/endpointing for stable incremental output.
  • Long audio: the audio encoder is a 30 s window; chunk longer inputs client-side.
  • Producing other precisions (needs the BF16 source weights): quantize the decoder Linears only via llm-compressor model_free_ptq, ignoring the audio tower, lm_head, and embeddings — scheme="W4A16" (4-bit) or "FP8_DYNAMIC" (lossless), ignore=["re:.*lm_head.*","re:.*embed_tokens.*","re:.*audio_tower.*"].

🎙️ Real-time streaming

Audar-ASR streams via LocalAgreement-2: as audio arrives the trailing window is re-decoded each hop and a word is committed only once two consecutive decodes agree on it — giving stable, low-latency incremental output over the GGUF runtime. Audar's production realtime engine serves the same policy over an OpenAI-Realtime-compatible WebSocket with model-based endpointing and ≥64 concurrent streams on a single A100-80GB.

🌍 Languages, dialects & tasks

  • Primary: Arabic — MSA and dialectal (Gulf/Emirati, Egyptian, Levantine, Maghrebi), plus code-switched Arabic–English; emits dialect-faithful orthography from audio alone.
  • Also: English + 28 additional languages.
  • Task: transcription (audio → UTF-8 text), prompt-steerable for language and formatting.

Intended use & limitations

Intended use. Broadcast/media transcription, meeting & contact-center intelligence, voice agents, captioning, and accessibility — cloud or on-prem.

Limitations.

  • Maghrebi / Moroccan Darija (Casablanca) remains the hardest condition (~63 % WER) for all systems.
  • Heavily code-switched telephony and low-SNR audio degrade accuracy relative to clean MSA.
  • Long-form audio can drift on very long recordings.
  • Not evaluated for, and must not be used for, covert speaker identification.

📜 License

Released under the AudarAI Community License v1.0 — research and limited commercial use for qualifying Community Entities; enterprise / large-scale / MaaS use requires an AudarAI Enterprise License. See audarai.com/license/audarai-community-license-v1.0.

Citation

@misc{audar-asr-turbo-2026,
  title  = {Audar-ASR-V1: A Multilingual, Arabic-First Generative Speech Recognition Foundation Model},
  author = {AudarAI},
  year   = {2026},
  note   = {Audar-ASR-V1-Turbo},
  url    = {https://github.com/AudarAI/Audar-ASR-V1/blob/main/report/Audar-ASR-V1-Technical-Report.pdf}
}

About AudarAI

Leading Arabic-First Multilingual Audio Intelligence

AudarAI starts with Arabic — and expands to the world.

We are building advanced multilingual audio intelligence that helps individuals, enterprises, and governments communicate across languages, cultures, and borders. By combining Arabic-first speech technology with global multilingual AI, AudarAI transforms voice into understanding, interaction, and connection.

Our work spans speech recognition, speech understanding, voice-enabled digital assistants, human-computer interaction, and intelligent audio systems designed for real-world impact. From empowering people to access technology in their native language to helping organizations communicate globally, AudarAI is shaping a future where every voice can be heard, understood, and connected.

Arabic-first. Multilingual by design. Human-centered at heart.

🌐 www.audarai.com · 🤗 Hugging Face · GitHub · contact@audarai.com

© 2026 AUDARAI PTE. LTD. · Licensed under the AudarAI Community License v1.0

Downloads last month
530
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for audarai/Audar-ASR-V1-Turbo

Quantizations
1 model