Ling-3.0-tiny โ€” Q8_0 embeddings + Q4_K variants

Three ~4.7 GiB quantizations of inclusionAI/Ling-3.0-tiny, made from the official bf16 GGUF, with per-variant pros/cons and everything below measured on this machine.

All three are recommendable; they differ in how the bit budget is spent, not in whether they work. Variant 2 is the one to use unless you have a reason not to.

file size PPL embedding/output everything else
Ling-3.0-tiny-Q8EMB-Q4K.gguf 4.661 GiB 4.170 Q8_0 Q4_K + llama.cpp's built-in promotions
Ling-3.0-tiny-Q8EMB-mixed.gguf โ† recommended 4.706 GiB 4.047 Q8_0 Q4_K, but Q5_K/Q6_K on attention, KDA/SSM and the shared expert
Ling-3.0-tiny-Q8EMB-mixed-topk16.gguf 4.706 GiB not measured Q8_0 same weights as variant 2, expert_used_count 8 โ†’ 16

Reference points: the bf16 original measures 3.969, so variant 1 is +5.1% and variant 2 +2.0% against unquantized weights. For calibration, the upstream Q4_K_M is 4.493 GiB โ€” smaller than these, because it keeps the embeddings at Q4_K. That saving is exactly what variant 2 spends on quality.

Quick start

# llama.cpp b11374 or newer (the bailingmoe3 arch + this GGUF's K-quants are required)
llama-cli -hf ZeroWw/Ling-3.0-tiny-quant-variants-GGUF:Ling-3.0-tiny-Q8EMB-mixed.gguf \
  -p "Question: A train leaves at 14:35 and arrives at 19:12. How long is the trip?" -n 1024

This is a hybrid reasoning model: its chat template has an enable_thinking flag and emits [Start thinking] โ€ฆ [End thinking]. Thinking is on by default. Give it a token budget โ€” the reasoning budget matters far more than any inference trick, and we did not find evidence that any arbiter modification improves it.

Model shape (measured from the GGUF, not copied from a config)

bailingmoe3, 24 blocks, hidden_size 1536, vocab 157184, 7.893 B parameters, โ‰ˆ1.4 B active per token.

  • Block 0 is dense; blocks 1-23 are MoE.
  • 128 experts per layer, top-8 of 128, sigmoid gating (expert_gating_func=2), group-limited routing (8 groups โ†’ keep the best 4), expert_weights_norm=True, routed scale 2.5.
  • Hybrid attention: 18 blocks use KDA linear attention (ssm_* tensors), 6 use full attention with LoRA-compressed q/kv (attn_q_a, attn_k_b, attn_v_b, attn_kv_a_mqa). layer_group_size=4.
  • 88% of all parameters are the routed experts โ€” the three fused tensors ffn_{gate,up,down}_exps.weight, shape [1536, 512, 128], in each of 23 layers.
  • The MoE router (ffn_gate_inp.weight) is F32 in the source and stays F32 in every variant here, so expert selection is never degraded by quantization.
  • Norms, the router, the load-balancing bias and the short-conv kernels are 1-D and stay F32 regardless.

Why variant 2 beats variant 1 for +1.0% size

Because the experts are 88% of the parameters, everything else is cheap to keep at higher precision. Variant 2 moves Q8_0 to the embedding/output as requested, and spends Q5_K on the families where precision matters most:

family params share V1 V2 why
routed experts (*_exps) 6.95 B 88.0% Q4_K/Q6_K Q4_K sheer volume; Q4_K is the right trade here
token_embd + output 483 M 6.1% Q8_0 Q8_0 logits and input lookup are the least forgiving tensors
attention 270 M 3.4% Q4_K Q5_K error here propagates through every layer
KDA/SSM (ssm_f_a, ssm_g_a) 114 M 1.45% Q4_K Q5_K recurrent path โ€” error integrates over sequence position, unlike softmax attention
shared expert (*_shexp) 54 M 0.69% Q4_K/Q6_K Q5_K/Q6_K active on every token, unlike a routed expert

Two notes on variant 1, which is kept as the honest baseline: on this build plain Q4_K is not pure Q4_K โ€” llama.cpp silently promotes ffn_down_exps to Q6_K on 11 of 23 layers and attn_v on 9 of 18, by tensor index, which is arbitrary. And 6 tensors (attn_k_b, ne[0]=128) fall back to q5_0 because that shape is not K-quant-compatible. Variant 2 assigns one type per family instead, which is why it is both better and smaller.

Variant 3 โ€” the arbiter experiment (read this before using it)

expert_used_count is a GGUF metadata value that llama.cpp reads at load time, so raising it from 8 to 16 is a one-u32 edit that every llama.cpp build honors โ€” the only model surgery here that stays distribution-compatible. Bit-identical weights to variant 2: same 5,052,457,888 bytes, same 526 tensors at identical offsets, same data_start. Only one uint32 in the header differs.

What we measured: it loads, it generates, and its reasoning traces differ from variant 2's (it routes to 16 experts per token instead of 8, roughly doubling active expert FLOPs per token โ€” expect it to be slower than variant 2 despite the identical file size). A 5-item smoke test (two arithmetic, three factual) passes 5/5 on both variants.

What we did not measure: perplexity, and any real reasoning benchmark. Perplexity takes ~30 min per model on 4 CPU cores, which did not fit our compute budget. So we cannot tell you whether top-k 16 is better or worse โ€” only that it is not obviously broken. Treat it as an experiment, not a recommendation. The honest default is variant 2.

For completeness, the ideas we considered and rejected:

  • Per-expert bit allocation (some experts at Q5_K, others at Q4_K) is not possible in a GGUF that mainline llama.cpp can load. A GGUF tensor carries exactly one ggml_type, and 88% of these parameters live in fused 3-D tensors that the loader reconstructs from canonical names at exactly [n_embd, n_ff_exp, 128]. You would have to split them into per-expert tensors, which requires patching the loader โ€” the resulting file would only load on a fork. The maximum granularity a stock file allows is per layer ร— role (526 individual decisions, reachable via llama-quantize --tensor-type-file).
  • Adding an expert from another model is blocked three ways over: n_expert=128 is global metadata and 128 = 8 groups ร— 16, so 129 breaks the group math; a foreign MLP's outputs are near-orthogonal to this model's residual stream; and the router is frozen, so it could not learn to select the new expert even if it loaded.
  • A true two-pass "arbiter loop" (route โ†’ apply experts โ†’ re-route on the updated hidden state) needs a patched bailingmoe3.cpp and has no training signal behind it. Not attempted.

What "is each expert for?" looks like from the outside

Worth recording, since it constrained the design above. Per-expert router-weight norms correlate โˆ’0.02 between layers, and the top-8 expert sets of any two layers overlap only 0-3 of 8:

blk.1  top-8: [  9  12  30  46  63  64  80  91]     blk.12 top-8: [ 57  66  76  86  92  94 118 123]
blk.6  top-8: [  2   7  38  42  43  59  65  91]     blk.23 top-8: [ 11  13  25  44  48  49  53  58]

Expert identity is layer-local โ€” there is no global "important expert" set, so any scheme that protected a fixed subset would be protecting noise. Within a layer, per-expert norms vary only 1.6-2.2ร—, i.e. importance is spread fairly evenly. The trained aux-loss-free load-balancing bias (exp_probs_b.bias) confirms load is not uniform and that it worsens with depth (std 0.013 at block 1 โ†’ 0.049 at block 23, range โˆ’0.155โ€ฆ+0.042; its correlation with router norms rises +0.17 โ†’ +0.67).

Resolving which individual expert fires on which tokens is possible but needs a patched llama.cpp โ€” llama-imatrix has no per-expert awareness (zero n_expert references in its disassembly; it stores one importance vector per tensor, pooled over all 128 experts). We did not build that.

Reproducing

# 0. source weights (official, sha256 verified)
huggingface-cli download inclusionAI/Ling-3.0-tiny-GGUF Ling-3.0-tiny-bf16.gguf --local-dir .

# 1. variant 1 โ€” literal spec
llama-quantize --token-embedding-type q8_0 --output-tensor-type q8_0 \
  Ling-3.0-tiny-bf16.gguf Ling-3.0-tiny-Q8EMB-Q4K.gguf Q4_K 4

# 2. variant 2 โ€” recommended
llama-quantize --token-embedding-type q8_0 --output-tensor-type q8_0 \
  --tensor-type ssm_f_a.weight=q5_K      --tensor-type ssm_g_a.weight=q5_K \
  --tensor-type ssm_beta.weight=q5_K     --tensor-type attn_q.weight=q5_K \
  --tensor-type attn_k.weight=q5_K       --tensor-type attn_v.weight=q5_K \
  --tensor-type attn_output.weight=q5_K  --tensor-type attn_q_a.weight=q5_K \
  --tensor-type attn_q_b.weight=q5_K     --tensor-type attn_kv_a_mqa.weight=q5_K \
  --tensor-type attn_v_b.weight=q5_K     --tensor-type attn_k_b.weight=q5_K \
  --tensor-type ffn_down_shexp.weight=q6_K --tensor-type ffn_gate_shexp.weight=q5_K \
  --tensor-type ffn_up_shexp.weight=q5_K \
  --tensor-type ffn_down_exps.weight=q5_K \
  --tensor-type ffn_gate_exps.weight=q4_K --tensor-type ffn_up_exps.weight=q4_K \
  --tensor-type ffn_gate.weight=q4_K     --tensor-type ffn_up.weight=q4_K \
  --tensor-type ffn_down.weight=q4_K \
  Ling-3.0-tiny-bf16.gguf Ling-3.0-tiny-Q8EMB-mixed.gguf Q4_K 4

# 3. variant 3 โ€” same weights, top-k 16 (header u32 patch; value width unchanged)
#    key "bailingmoe3.expert_used_count" -> value u32, 8 -> 16

Requirements: llama.cpp with bailingmoe3 support (b11374 tested), no torch or gguf-py needed.

Verification performed

  • Source bf16 GGUF sha256 b11d4a45d3ad9e0dc3ea592a85f950beb53bd8a97d0f229db1f86ce3dee8d4c3, matching the Hub's LFS oid.

  • All three outputs: 526 tensors, 53 KV keys, tokenizer vocab/merges and the 6,049-byte chat template byte-identical to the source. The only KV that changes is general.file_type (32 = BF16 โ†’ the quantizer's label).

  • Variant 3 verified as a surgical edit: one uint32 differs, and every tensor name, shape, type and byte offset plus data_start is unchanged.

  • Generation smoke-tested on all variants (llama.cpp b11374, 4 threads); both non-topk16 variants produce coherent [Start thinking] traces.

  • Perplexity, 3 chunks of 512 tokens on an 8,868-byte English fixture:

    model PPL
    bf16 (reference) 3.9692 ยฑ 0.30642
    variant 1 4.1700 ยฑ 0.32828
    variant 2 4.0471 ยฑ 0.31702

Caveats on that table, which matter

  • The error bars overlap. Variant 2 is better in direction and is 61% closer to bf16, but 3 chunks is not enough to call it a statistically established win. Treat it as a strong hint, not proof.
  • The fixture is 424 tokens repeated 4ร—, so chunks are correlated. Fine for a paired comparison, weak as an absolute PPL.
  • PPL measures next-token fit on English prose. It is not a reasoning benchmark, and the model is multilingual (157k vocab) โ€” an imatrix-driven variant chosen on English-only statistics could easily be mis-tuned. We built that variant and then cut it for time; it is the obvious next improvement.

Known issues

  • Loading any variant prints special_eos_id is not in special_eog_ids. This comes from inclusionAI's own conversion, not from our quantization; it was harmless in every test here, but if generation does not stop, look there first.
  • Variant 3's top-k 16 is not a configuration the model was trained with (it was trained at exactly 8), so expect a distribution shift and lower throughput.

License & attribution

MIT, inherited from inclusionAI/Ling-3.0-tiny and its official GGUF release. These are third-party quantizations; please report quality problems against inclusionAI/Ling-3.0-tiny first. GGUF format and llama.cpp are by the llama.cpp authors (MIT).

Downloads last month
455
GGUF
Model size
8B params
Architecture
bailingmoe3
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ZeroWw/Ling-3.0-tiny-quant-variants-GGUF

Quantized
(32)
this model