Ling-3.0-tiny — Q4_K / MXFP4 / Q8_0 GGUF

Fast GGUF for mainstream llama.cpp. Works out of the box on any build that ships MXFP4 support (llama.cpp ≥ b3900).

Recipe (antirez style)

Tensor group Format
Routed expert gate / up Q4_K (calibrated)
Routed expert down MXFP4
Shared experts, dense FFN (layer 0), attention, KDA, MLA Q-LoRA, output Q8_0
Token embeddings BF16
Routers, norms, SSM, gate biases F32

Same quantization strategy as Salvatore Sanfilippo's Qwen3.8 Flash Next Q4: calibrated Q4_K on the large expert projections, MXFP4 on the down path, Q8_0 everywhere the signal matters most.

Files

File Size
Ling-3.0-tiny-Q4K-MXFP4-Q8.gguf 4.74 GiB

Speed (llama.cpp)

Measured on a 12th-gen Intel laptop (4P+8E cores, 32 GB RAM).

Backend Threads Prefill pp512 Decode tg128
CPU 6 ~57 tok/s ~20 tok/s
Vulkan (Intel Iris Xe) 2 ~247 tok/s ~28 tok/s

Fast enough for real-time chat on CPU alone, no discrete GPU needed.

Usage

# interactive chat
llama-cli -m Ling-3.0-tiny-Q4K-MXFP4-Q8.gguf -t 4

# server
llama-server -m Ling-3.0-tiny-Q4K-MXFP4-Q8.gguf -t 4 --host 0.0.0.0 --port 8080

Model

Ling-3.0-tiny is a 128-expert MoE with 1.3 B active parameters out of 7.9 B total. Architecture: BailingMoE3 with KDA attention, MLA Q-LoRA compression, shared experts, and a hybrid SSM + attention layer 0. Context: 131 072 tokens.

Original model: inclusionAI/Ling-3.0-tiny. Original BF16 GGUF: bloomer010/Ling-3.0-tiny-GGUF. License: MIT.

Downloads last month
141
GGUF
Model size
8B params
Architecture
bailingmoe3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ninnix96/Ling-3.0-tiny-gguf

Quantized
(33)
this model

Collection including Ninnix96/Ling-3.0-tiny-gguf