Instructions to use ZeroWw/Ling-3.0-tiny-quant-variants-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ZeroWw/Ling-3.0-tiny-quant-variants-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF # Run inference directly in the terminal: llama cli -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF # Run inference directly in the terminal: llama cli -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF # Run inference directly in the terminal: ./llama-cli -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
Use Docker
docker model run hf.co/ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
- LM Studio
- Jan
- vLLM
How to use ZeroWw/Ling-3.0-tiny-quant-variants-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ZeroWw/Ling-3.0-tiny-quant-variants-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZeroWw/Ling-3.0-tiny-quant-variants-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
- Ollama
How to use ZeroWw/Ling-3.0-tiny-quant-variants-GGUF with Ollama:
ollama run hf.co/ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
- Unsloth Desktop
- Pi
How to use ZeroWw/Ling-3.0-tiny-quant-variants-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ZeroWw/Ling-3.0-tiny-quant-variants-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ZeroWw/Ling-3.0-tiny-quant-variants-GGUF with Docker Model Runner:
docker model run hf.co/ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
- Lemonade
How to use ZeroWw/Ling-3.0-tiny-quant-variants-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
Run and chat with the model
lemonade run user.Ling-3.0-tiny-quant-variants-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use ZeroWw/Ling-3.0-tiny-quant-variants-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ZeroWw/Ling-3.0-tiny-quant-variants-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ZeroWw/Ling-3.0-tiny-quant-variants-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Ling-3.0-tiny โ Q8_0 embeddings + Q4_K variants
- Quick start
- Model shape (measured from the GGUF, not copied from a config)
- Why variant 2 beats variant 1 for +1.0% size
- Variant 3 โ the arbiter experiment (read this before using it)
- What "is each expert for?" looks like from the outside
- Reproducing
- Verification performed
- Known issues
- License & attribution
- Quick start
Ling-3.0-tiny โ Q8_0 embeddings + Q4_K variants
Three ~4.7 GiB quantizations of inclusionAI/Ling-3.0-tiny,
made from the official bf16 GGUF, with per-variant pros/cons and everything below measured on this machine.
All three are recommendable; they differ in how the bit budget is spent, not in whether they work. Variant 2 is the one to use unless you have a reason not to.
| file | size | PPL | embedding/output | everything else |
|---|---|---|---|---|
Ling-3.0-tiny-Q8EMB-Q4K.gguf |
4.661 GiB | 4.170 | Q8_0 | Q4_K + llama.cpp's built-in promotions |
Ling-3.0-tiny-Q8EMB-mixed.gguf โ recommended |
4.706 GiB | 4.047 | Q8_0 | Q4_K, but Q5_K/Q6_K on attention, KDA/SSM and the shared expert |
Ling-3.0-tiny-Q8EMB-mixed-topk16.gguf |
4.706 GiB | not measured | Q8_0 | same weights as variant 2, expert_used_count 8 โ 16 |
Reference points: the bf16 original measures 3.969, so variant 1 is +5.1% and variant 2 +2.0% against
unquantized weights. For calibration, the upstream Q4_K_M is 4.493 GiB โ smaller than these, because it keeps the
embeddings at Q4_K. That saving is exactly what variant 2 spends on quality.
Quick start
# llama.cpp b11374 or newer (the bailingmoe3 arch + this GGUF's K-quants are required)
llama-cli -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF:Ling-3.0-tiny-Q8EMB-mixed.gguf \
-p "Question: A train leaves at 14:35 and arrives at 19:12. How long is the trip?" -n 1024
This is a hybrid reasoning model: its chat template has an enable_thinking flag and emits
[Start thinking] โฆ [End thinking]. Thinking is on by default. Give it a token budget โ the reasoning budget
matters far more than any inference trick, and we did not find evidence that any arbiter modification improves it.
Model shape (measured from the GGUF, not copied from a config)
bailingmoe3, 24 blocks, hidden_size 1536, vocab 157184, 7.893 B parameters, โ1.4 B active per token.
- Block 0 is dense; blocks 1-23 are MoE.
- 128 experts per layer, top-8 of 128, sigmoid gating (
expert_gating_func=2), group-limited routing (8 groups โ keep the best 4),expert_weights_norm=True, routed scale 2.5. - Hybrid attention: 18 blocks use KDA linear attention (
ssm_*tensors), 6 use full attention with LoRA-compressed q/kv (attn_q_a,attn_k_b,attn_v_b,attn_kv_a_mqa).layer_group_size=4. - 88% of all parameters are the routed experts โ the three fused tensors
ffn_{gate,up,down}_exps.weight, shape[1536, 512, 128], in each of 23 layers. - The MoE router (
ffn_gate_inp.weight) is F32 in the source and stays F32 in every variant here, so expert selection is never degraded by quantization. - Norms, the router, the load-balancing bias and the short-conv kernels are 1-D and stay F32 regardless.
Why variant 2 beats variant 1 for +1.0% size
Because the experts are 88% of the parameters, everything else is cheap to keep at higher precision. Variant 2 moves Q8_0 to the embedding/output as requested, and spends Q5_K on the families where precision matters most:
| family | params | share | V1 | V2 | why |
|---|---|---|---|---|---|
routed experts (*_exps) |
6.95 B | 88.0% | Q4_K/Q6_K | Q4_K | sheer volume; Q4_K is the right trade here |
token_embd + output |
483 M | 6.1% | Q8_0 | Q8_0 | logits and input lookup are the least forgiving tensors |
| attention | 270 M | 3.4% | Q4_K | Q5_K | error here propagates through every layer |
KDA/SSM (ssm_f_a, ssm_g_a) |
114 M | 1.45% | Q4_K | Q5_K | recurrent path โ error integrates over sequence position, unlike softmax attention |
shared expert (*_shexp) |
54 M | 0.69% | Q4_K/Q6_K | Q5_K/Q6_K | active on every token, unlike a routed expert |
Two notes on variant 1, which is kept as the honest baseline: on this build plain Q4_K is not pure Q4_K โ
llama.cpp silently promotes ffn_down_exps to Q6_K on 11 of 23 layers and attn_v on 9 of 18, by tensor index,
which is arbitrary. And 6 tensors (attn_k_b, ne[0]=128) fall back to q5_0 because that shape is not
K-quant-compatible. Variant 2 assigns one type per family instead, which is why it is both better and
smaller.
Variant 3 โ the arbiter experiment (read this before using it)
expert_used_count is a GGUF metadata value that llama.cpp reads at load time, so raising it from 8 to 16 is a
one-u32 edit that every llama.cpp build honors โ the only model surgery here that stays distribution-compatible.
Bit-identical weights to variant 2: same 5,052,457,888 bytes, same 526 tensors at identical offsets, same
data_start. Only one uint32 in the header differs.
What we measured: it loads, it generates, and its reasoning traces differ from variant 2's (it routes to 16 experts per token instead of 8, roughly doubling active expert FLOPs per token โ expect it to be slower than variant 2 despite the identical file size). A 5-item smoke test (two arithmetic, three factual) passes 5/5 on both variants.
What we did not measure: perplexity, and any real reasoning benchmark. Perplexity takes ~30 min per model on 4 CPU cores, which did not fit our compute budget. So we cannot tell you whether top-k 16 is better or worse โ only that it is not obviously broken. Treat it as an experiment, not a recommendation. The honest default is variant 2.
For completeness, the ideas we considered and rejected:
- Per-expert bit allocation (some experts at Q5_K, others at Q4_K) is not possible in a GGUF that mainline
llama.cpp can load. A GGUF tensor carries exactly one
ggml_type, and 88% of these parameters live in fused 3-D tensors that the loader reconstructs from canonical names at exactly[n_embd, n_ff_exp, 128]. You would have to split them into per-expert tensors, which requires patching the loader โ the resulting file would only load on a fork. The maximum granularity a stock file allows is per layer ร role (526 individual decisions, reachable viallama-quantize --tensor-type-file). - Adding an expert from another model is blocked three ways over:
n_expert=128is global metadata and 128 = 8 groups ร 16, so 129 breaks the group math; a foreign MLP's outputs are near-orthogonal to this model's residual stream; and the router is frozen, so it could not learn to select the new expert even if it loaded. - A true two-pass "arbiter loop" (route โ apply experts โ re-route on the updated hidden state) needs a
patched
bailingmoe3.cppand has no training signal behind it. Not attempted.
What "is each expert for?" looks like from the outside
Worth recording, since it constrained the design above. Per-expert router-weight norms correlate โ0.02 between layers, and the top-8 expert sets of any two layers overlap only 0-3 of 8:
blk.1 top-8: [ 9 12 30 46 63 64 80 91] blk.12 top-8: [ 57 66 76 86 92 94 118 123]
blk.6 top-8: [ 2 7 38 42 43 59 65 91] blk.23 top-8: [ 11 13 25 44 48 49 53 58]
Expert identity is layer-local โ there is no global "important expert" set, so any scheme that protected a
fixed subset would be protecting noise. Within a layer, per-expert norms vary only 1.6-2.2ร, i.e. importance is
spread fairly evenly. The trained aux-loss-free load-balancing bias (exp_probs_b.bias) confirms load is not
uniform and that it worsens with depth (std 0.013 at block 1 โ 0.049 at block 23, range โ0.155โฆ+0.042; its
correlation with router norms rises +0.17 โ +0.67).
Resolving which individual expert fires on which tokens is possible but needs a patched llama.cpp โ
llama-imatrix has no per-expert awareness (zero n_expert references in its disassembly; it stores one
importance vector per tensor, pooled over all 128 experts). We did not build that.
Reproducing
# 0. source weights (official, sha256 verified)
huggingface-cli download inclusionAI/Ling-3.0-tiny-GGUF Ling-3.0-tiny-bf16.gguf --local-dir .
# 1. variant 1 โ literal spec
llama-quantize --token-embedding-type q8_0 --output-tensor-type q8_0 \
Ling-3.0-tiny-bf16.gguf Ling-3.0-tiny-Q8EMB-Q4K.gguf Q4_K 4
# 2. variant 2 โ recommended
llama-quantize --token-embedding-type q8_0 --output-tensor-type q8_0 \
--tensor-type ssm_f_a.weight=q5_K --tensor-type ssm_g_a.weight=q5_K \
--tensor-type ssm_beta.weight=q5_K --tensor-type attn_q.weight=q5_K \
--tensor-type attn_k.weight=q5_K --tensor-type attn_v.weight=q5_K \
--tensor-type attn_output.weight=q5_K --tensor-type attn_q_a.weight=q5_K \
--tensor-type attn_q_b.weight=q5_K --tensor-type attn_kv_a_mqa.weight=q5_K \
--tensor-type attn_v_b.weight=q5_K --tensor-type attn_k_b.weight=q5_K \
--tensor-type ffn_down_shexp.weight=q6_K --tensor-type ffn_gate_shexp.weight=q5_K \
--tensor-type ffn_up_shexp.weight=q5_K \
--tensor-type ffn_down_exps.weight=q5_K \
--tensor-type ffn_gate_exps.weight=q4_K --tensor-type ffn_up_exps.weight=q4_K \
--tensor-type ffn_gate.weight=q4_K --tensor-type ffn_up.weight=q4_K \
--tensor-type ffn_down.weight=q4_K \
Ling-3.0-tiny-bf16.gguf Ling-3.0-tiny-Q8EMB-mixed.gguf Q4_K 4
# 3. variant 3 โ same weights, top-k 16 (header u32 patch; value width unchanged)
# key "bailingmoe3.expert_used_count" -> value u32, 8 -> 16
Requirements: llama.cpp with bailingmoe3 support (b11374 tested), no torch or gguf-py needed.
Verification performed
Source bf16 GGUF sha256
b11d4a45d3ad9e0dc3ea592a85f950beb53bd8a97d0f229db1f86ce3dee8d4c3, matching the Hub's LFS oid.All three outputs: 526 tensors, 53 KV keys, tokenizer vocab/merges and the 6,049-byte chat template byte-identical to the source. The only KV that changes is
general.file_type(32 = BF16 โ the quantizer's label).Variant 3 verified as a surgical edit: one
uint32differs, and every tensor name, shape, type and byte offset plusdata_startis unchanged.Generation smoke-tested on all variants (llama.cpp b11374, 4 threads); both non-topk16 variants produce coherent
[Start thinking]traces.Perplexity, 3 chunks of 512 tokens on an 8,868-byte English fixture:
model PPL bf16 (reference) 3.9692 ยฑ 0.30642 variant 1 4.1700 ยฑ 0.32828 variant 2 4.0471 ยฑ 0.31702
Caveats on that table, which matter
- The error bars overlap. Variant 2 is better in direction and is 61% closer to bf16, but 3 chunks is not enough to call it a statistically established win. Treat it as a strong hint, not proof.
- The fixture is 424 tokens repeated 4ร, so chunks are correlated. Fine for a paired comparison, weak as an absolute PPL.
- PPL measures next-token fit on English prose. It is not a reasoning benchmark, and the model is multilingual (157k vocab) โ an imatrix-driven variant chosen on English-only statistics could easily be mis-tuned. We built that variant and then cut it for time; it is the obvious next improvement.
Known issues
- Loading any variant prints
special_eos_id is not in special_eog_ids. This comes from inclusionAI's own conversion, not from our quantization; it was harmless in every test here, but if generation does not stop, look there first. - Variant 3's top-k 16 is not a configuration the model was trained with (it was trained at exactly 8), so expect a distribution shift and lower throughput.
License & attribution
MIT, inherited from inclusionAI/Ling-3.0-tiny and its
official GGUF release. These are third-party quantizations; please report quality problems against
inclusionAI/Ling-3.0-tiny first. GGUF format and llama.cpp are by the llama.cpp authors (MIT).
- Downloads last month
- 455
We're not able to determine the quantization variants.
Model tree for ZeroWw/Ling-3.0-tiny-quant-variants-GGUF
Base model
inclusionAI/Ling-3.0-tiny