Clef — NVFP4

An unofficial, community NVFP4 quantization of Cloudflare/clef, a 27B multimodal decision model that turns a state and a schema of typed questions into a probability for every allowed option in a single forward pass.

Quantized so the model runs on a single 16 GB consumer GPU (NVIDIA Blackwell, sm_120). In bf16 Clef is 55 GB and does not fit; this checkpoint holds 13.2 GiB of weights in VRAM.

Not affiliated with or endorsed by Cloudflare. The model, its architecture and joint_schema_model.py are theirs; this repo only re-encodes the weights. Both the base model and this derivative are Apache-2.0.

Usage

The quantized weights are not loaded by from_pretrained or by the base model's load_release_model — those build the full-precision graph and need ≈55 GB (they OOM on a 16 GB card). Use the included clef_rt.py, which keeps the decoder in NVFP4 on the GPU and the embeddings / lm_head in host RAM:

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("myroslavtryhubets/clef-NVFP4")
sys.path.insert(0, path)
from clef_rt import Clef                       # ships in this repo

model = Clef(path, "nvfp4")                    # ≈13.2 GiB VRAM on a Blackwell card

out = model.probs([model.encode({
    "state": "Our checkout is down and orders are blocked.",
    "questions": {
        "outage": {"type": "noul", "instructions": "Is a service down?"},
        "team": {"type": "choice", "instructions": "Which team should handle it?",
                 "criteria": {"billing": "payments", "tech": "outages"}},
    },
})])[0]
print(out)   # {'outage': {'true': 0.98, ...}, 'team': {'billing': 0.63, ...}}

Clef(path, "nvfp4-w4a16") loads the same weights but keeps activations in bf16 (weights-only 4-bit) — slower, slightly more accurate, same VRAM.

What was quantized

  • Scheme: NVFP4 — FP4 (e2m1) weights, an FP8 (e4m3) scale per 16-element block, one FP32 global scale per tensor; W4A4 (weights and activations in 4-bit). Format: compressed-tensors nvfp4-pack-quantized.
  • Tooling: llm-compressor 0.14 model-free PTQ for the weights; static activation global scales from layer-by-layer bf16 calibration on 256 records (UltraChat + security-log text in Clef's own prompt format, static_minmax).
  • Kept in bf16: lm_head, embed_tokens, the vision tower, the Gated-DeltaNet in_proj_a / in_proj_b / conv1d, all norms, and the joint schema head.
  • Fused global scales are shared across q/k/v, gate/up and the GDN in_proj_qkv+in_proj_z pair (the last so vLLM's in_proj_qkvz fusion stays consistent).

Quality — Decision Index 0.2.1

Measured on 11 benchmarks of the Decision Index 0.2.1 suite (20,337 requests), rebuilt byte-identical with the official reproduction kit and scored with its own scorer. The Clef bf16 (CF) column is Cloudflare's published result; This NVFP4 is this checkpoint on one RTX 5060 Ti 16 GB. Values are the coverage-adjusted percentages the board uses.

Benchmark Metric Clef bf16 (CF) This NVFP4 Δ
BFCL case exact accuracy 98.5 98.6 +0.1
BANKING77 macro-F1 94.2 93.7 -0.5
CLINC150+OOS macro-F1 97.4 97.3 -0.1
ContractNLI macro-F1 81.4 79.3 (2 OOM) -2.1
ANLI macro-F1 69.8 69.5 -0.3
ARC-Challenge accuracy 97.7 97.4 -0.3
WinoGrande accuracy 93.5 92.0 -1.5
MuSR accuracy 83.5 82.2 -1.3
FinEntity macro-F1 96.2 96.1 -0.1
CRUXEval accuracy 86.7 84.0 -2.7
PhishNChips accuracy 79.6 76.7 -2.9

Mean absolute gap to Cloudflare: 1.08 points (worst −2.9, PhishNChips). Two of the 123 ContractNLI requests (≈10k tokens) do not fit 16 GB and are scored as wrong per the kit's coverage rule — about 1.6 pts of the ContractNLI gap.

Requirements

  • GPU: NVIDIA Blackwell (sm_120), >=16 GB. NVFP4 matmul runs on the FP4 tensor cores via torch._scaled_mm. Tested on torch 2.14.0+cu130, transformers 5.17.0, driver 580.95.05.
  • Verified by downloading this repo into a clean environment and loading with clef_rt.py (13.2 GiB weights, correct decisions).

Caveats

  • W4A4. The clef-flash sibling loses ≈17 pts on CLINC150 under W4A4 (150 near-tied classes vs 4-bit activations). On this 27B model CLINC150 is unaffected (97.4 -> 97.3), but activation quantization is the main risk if you add tasks with many near-tied options.
  • vLLM: not supported for decisions. vLLM registers the Qwen3_5 backbone but has no Clef joint-head / decision path (checked against its model registry and qwen3_5.py), so it cannot produce typed decisions. vLLM's structured-generation work (PR #57250) targets DiffusionGemma's token-logprob mechanism, which is unrelated to Clef's routing head. Serving Clef in vLLM would require porting JointSchemaHead as a custom model. Use clef_rt.py (above).
  • Quality was verified only through this runtime, not a third-party serving stack.

Credits

Model, architecture and decision API by Cloudflare (blog). Base model Qwen/Qwen3.8-27B. Quantization by @myroslavtryhubets.

Downloads last month
35
Safetensors
Model size
27B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for myroslavtryhubets/clef-NVFP4

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Quantized
(18)
this model