Clef-Flash mixed-precision GGUF (measured-KL, text only)

Mixed precision, not ternary. Tensors use IQ2_XXS, IQ2_S, Q2_K, IQ3_XXS, Q3_K, Q4_K and Q8_0; no ternary format (TQ1_0/TQ2_0) is used. This repo was renamed from Jakevin/clef-flash-ternary-GGUF on 2026-10-09; the old URL redirects here and the v1.0 tag is unchanged.

Version v1.0 (2026-10-09). Unofficial post-training quantization of Cloudflare/clef-flash revision 17f0b0ad64efb65d273590632833508766b2aae6. This is not an official Cloudflare release and is not endorsed by Cloudflare or the Qwen team. Licensed Apache-2.0 like the original (LICENSE). NOTICE.md lists the changes.

The file is the text backbone only. There is no vision tower. llama.cpp by itself does not produce Clef decisions: the joint schema head runs in Python on the backbone's final hidden states, and the option rows come from the original bf16 lm_head. The GGUF output.weight is Q2_K and the head does not read it.

general.architecture is qwen35. The llama.cpp clef graph does not export t_h_nextn, which is the hidden state this recipe reads, so the file is the Qwen3.5 text model rather than architecture clef. general.name is Snap Text because the conversion directory had that name. The weights are Clef-Flash text.

Size

bytes
clef-flash-mkl-3.30GB.gguf 3.300 GB (3,299,992,800)
original bf16 release 19.06 GB (18.82 GB shards + 0.24 GB head)

Scores

Author's private frozen development splits (kev). One seed. No confidence interval was computed.

suite n (clean) bf16 acc this model retained
decision-v7 1264 0.8861 0.8813 99.5%
transfer-v9 1046 0.8011 0.7859 98.1%
suite Brier NLL ECE
decision-v7 0.1755 0.3382 0.0278
transfer-v9 0.3061 0.6042 0.0421

Retention is this model's clean accuracy divided by the bf16 clean accuracy on the same suite.

decision-v7 can be optimistic. The imatrix and the KL allocation both used decision-v7 records: 128 calibration records, seed 1234, for the imatrix, and the first 32 of that draw (10,383 tokens) for the per-tensor KL. transfer-v9 was not used for calibration.

These are not the public Decision Index or Typesafe numbers.

Method

Per-tensor measured-KL sensitivity, then a knapsack that assigns each measured tensor a ggml type from IQ2_XXS, IQ2_S, Q2_K, IQ3_XXS, Q3_K, Q4_K, and Q8_0. The shipped file is llama-quantize --tensor-type with that imatrix. alloc/selection.json is the knapsack result. alloc/tensor-types.txt is the pattern file passed to the quantizer (it pins ssm_alpha and ssm_beta to bf16; the other assignments are exact tensor names).

Counts in the GGUF (427 tensors):

type tensors what
IQ2_XXS 68 body linears
IQ2_S 23 body linears
IQ3_XXS 35 body linears and token_embd
Q2_K 16 15 body linears plus output.weight
Q3_K 13 body linears
Q4_K 36 body linears
Q8_0 11 body linears
BF16 48 ssm_alpha, ssm_beta (24 each)
F32 177 norms, ssm_conv1d, ssm_a, ssm_dt

Norms, conv, ssm_alpha, and ssm_beta are not quantized. In the file, norms and conv are F32; ssm_alpha and ssm_beta are BF16. output.weight is fixed at Q2_K and is not read by the classification head. token_embd is IQ3_XXS (IQ2_XXS asserts an imatrix, and llama-quantize does not pass one for the embedding).

Sum of the per-tensor KL values used by the knapsack (not a model KL): 0.009015. Dry-run file estimate was 3.300 GB.

Other files of the same source model

Listed side by side. Different formats and different calibration. This card does not say which allocation is better.

release what it is size decision-v7 transfer-v9
this repo, v1.0 GGUF, measured-KL mixed types 3.300 GB 0.8813 (99.5% of bf16) 0.7859 (98.1% of bf16)
Jakevin/clef-flash-ternary-mlx v2.0 MLX packed mixed-bit GPTQ 3.13 GB 0.8766 (98.9% of bf16) 0.7361 (91.9% of bf16)
bartowski/Cloudflare_clef-flash-GGUF Cloudflare_clef-flash-Q2_K.gguf GGUF Q2_K 4.19 GB (4,194,031,776 bytes) not scored here not scored here

The MLX v2.0 scores are the ones on that repo's model card (same bf16 references, 0.8861 and 0.8011). Its calibration and packed format are not this GGUF's. The bartowski Q2_K file size is the file on that repo. We did not evaluate it.

Use

The joint head needs torch, safetensors, numpy, and transformers. hsdump needs llama.cpp b11407 or later with text-only qwen35. Build it from hsdump.cpp in this folder. The two embedding calls are already in that llama.cpp (src/llama-ext.h); they are not in the installed llama.h, so the cpp file declares them. No llama.cpp patch is required at b11407 or later.

output.weight inside the GGUF is not the classification head. Pass --base as a checkout of Cloudflare/clef-flash (the same revision as above). Only lm_head.weight is read from it.

c++ -std=c++17 -O2 \
  -I llama.cpp/include -I llama.cpp/ggml/include \
  hsdump.cpp -o hsdump \
  -L /path/to/llama.cpp/build/bin -lllama \
  -Wl,-rpath,/path/to/llama.cpp/build/bin

python run_clef_gguf.py \
  --base /path/to/Cloudflare/clef-flash \
  --hsdump ./hsdump

The default record is the invoice example from the Cloudflare card (one choice question and one noul question). --record file.json scores another text record of the same shape. Hidden states are llama.cpp final-norm embeddings (llama_set_embeddings_nextn, masked false), the read used by Livesport/clef-flash-GGUF.

Limitations

  • Text only. No images, no video.
  • Post-training quantization. No recovery training.
  • One calibration seed (1234). decision-v7 participated in the imatrix and the KL; treat that score as possibly high. transfer-v9 did not.
  • No paired bootstrap and no confidence interval.
  • llama.cpp chat or completion on this file is not a Clef decision. The head is joint_schema_model.py.

License and attribution

Derived from Cloudflare/clef-flash (© Cloudflare, Apache-2.0), itself post-trained from Qwen/Qwen3.5-9B (Apache-2.0). The hidden-state read follows Livesport/clef-flash-GGUF. Inference runs on llama.cpp.

繁中摘要

這是 Cloudflare/clef-flash 文字主幹的 measured-KL 混合精度 GGUF,3.300 GB(3,299,992,800 bytes)。decision-v7 正確率 0.8813(bf16 0.8861,保留 99.5%),transfer-v9 0.7859(bf16 0.8011,保留 98.1%)。單一 seed,沒有信賴區間。imatrix 與 KL 用了 decision-v7 的校正資料(imatrix 128 筆、seed 1234;KL 用其中前 32 筆),所以 decision-v7 可能偏樂觀;transfer-v9 沒有參與校正。沒有視覺。分類要跑 Python 的 JointSchemaHead,選項列取自原模型 bf16 lm_head;llama.cpp 單獨跑不出 clef 的判斷。output.weight 是 Q2_K,分類頭不讀它。norms、conv 在檔案裡是 F32,ssm_alpha / ssm_beta 是 BF16,都沒有再量化。

Downloads last month
33
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jakevin/clef-flash-mixed-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(42)
this model