rapidchat GGUF

Quantized builds of rapidchat, the 9B airline-support fine-tune. These builds were not separately benchmarked: the 86.5 tau2-bench airline pass^1 is for the full bf16 model, and 4-bit quantization usually costs a little accuracy. Training data: rapidchat-data.

file size class use
rapidchat-Q4_K_M.gguf ~5-6 GB phones and laptops
rapidchat-Q8_0.gguf ~9-10 GB desktops, near-lossless
rapidchat-Q3_K_M.gguf 4.6 GB phones with less memory, recommended small build
rapidchat-Q3_K_S.gguf 4.3 GB tighter phones
rapidchat-Q2_K.gguf 3.8 GB smallest, noticeably lower quality

run it

llama-server -hf badr7/rapidchat-GGUF:Q4_K_M --jinja

Works in any GGUF app (LM Studio, Ollama, phone apps that load GGUF).

Smoke test: Q4_K_M loads in llama.cpp and answers a flight-change request coherently (about 6 tok/s on 12 CPU threads).

Downloads last month
212
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for badr7/rapidchat-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(2)
this model