Clef — NVFP4

A mixed NVFP4 / FP8 quantization of Cloudflare/clef, the 27B multimodal decision model. Unofficial — not affiliated with or endorsed by Cloudflare.

Clef answers typed questions about a state (text, JSON, images, video) with one probability per allowed option, in a single prefill pass. Its value is the probabilities, so this quant was checked option by option against the BF16 release, not only on accuracy.

23 GB on disk (BF16 release: 55 GB); 20.2 GiB of weights in vLLM.

What is quantized

Part Format
MLP gate/up/down_proj, layers 0–55 NVFP4 (W4A4, group 16, FP8 scales, static activation global scale)
MLP gate/up/down_proj, layers 56–63 FP8 (per-channel weights, dynamic per-token activations)
Full-attention q/k/v/o_proj; linear-attention in_proj_qkv/in_proj_z/out_proj FP8 (as above)
Linear-attention in_proj_a/in_proj_b, norms, conv, embeddings BF16
Vision encoder BF16
lm_head BF16 — the joint head reads its rows directly as option embeddings
Joint schema head (joint_head.safetensors) BF16, unchanged from the release

The layout follows the community recipe for Qwen3.8-27B with one change: lm_head stays BF16. The last eight layers stay FP8 on purpose: comparing the release against Qwen/Qwen3.8-27B, the vision encoder and early layers are byte-identical, and Clef's post-training lives in the upper layers.

Method: llm-compressor QuantizationModifier (round-to-nearest, no GPTQ), calibrated on 374 Clef-format records — 128 BANKING77 intent questions (77-way and 10-way schemas), 150 game-agent state records with a choice question, and 96 Flickr30k photos with noul/choice/score questions. Calibration and evaluation records are disjoint. recipe.yaml is the exact recipe.

Drift vs BF16

Same 615 records through the BF16 release and through this checkpoint on vLLM's real NVFP4/FP8 kernels (FlashInfer CUTLASS FP4 GEMM). agreement = top option identical to BF16. TV = mean total-variation distance between the per-question distributions (0 = identical, 1 = disjoint).

Eval set Questions Agreement Mean TV Accuracy BF16 → NVFP4
BANKING77 test, full 77-way schema 400 98.5% 0.019 94.25% → 94.50%
Flickr30k photos, 4 mixed-type questions each 256 96.9% 0.017 (no labels)
Game-agent states, 2–4 option choice 150 96.7% 0.047 68.7% → 67.3% *
README invoice example 2 100% 0.0007 —

* graded against synthetic labels that the BF16 model itself matches only 69% of the time; read the agreement column, not the accuracy, for that row.

Disagreements are near-ties: in the calibration-time (fake-quant) run, the 15 flipped questions had a median BF16 top-1/top-2 margin of 0.13 (11 of 15 under 0.2), against 0.95 for questions that agreed.

Speed

One decision at a time (batch 1, the latency case) on an NVIDIA DGX Spark (GB10, 273 GB/s unified memory), vLLM 0.23.1, CUDA graphs on:

Record Input tokens Median latency
text state, 1 question 288 177 ms
640 px image + 4 questions 621 257 ms
BANKING77, 77 options 1834 705 ms

The joint head adds 3–6 ms of that. On this machine the floor (~170 ms) is set by reading the weights once per pass; GPUs with more memory bandwidth should be considerably quicker (not measured here).

Usage

The files, record format and systemone API are the same as the BF16 release; see the Clef model card for the full input format.

vLLM (recommended — real FP4 kernels)

clef_vllm.py runs the backbone in vLLM as a pooling model (final hidden state of every token) and Clef's joint head on top, in the same process. Needs a Blackwell GPU for NVFP4 and a vLLM build with Qwen3_5ForConditionalGeneration (tested: 0.23.1 nightly).

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("simonlehmann/clef-NVFP4")
sys.path.insert(0, path)
from clef_vllm import ClefVLLM

clef = ClefVLLM(path, max_model_len=16384, gpu_memory_utilization=0.4)
response = clef.systemone({
    "model": "clef",
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle the message?",
            "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
        },
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
        "outage": {"type": "noul", "instructions": "Is a service down?"},
    },
})
print(response["answers"])

# Many records at once, images included (PIL), batched through vLLM:
# clef.probabilities([record, ...]) -> [{question_id: {option_id: probability}}, ...]

max_images / max_videos in the constructor set vLLM's per-request media limits (default 1 image).

transformers (compatibility — decompresses to BF16)

transformers' compressed-tensors integration cannot run the mixed NVFP4/FP8 layers compressed, so load with run_compressed=False. This decompresses to BF16 at load time (~57 GB of GPU memory, no speed-up): useful to check outputs, not to save memory.

import sys
from huggingface_hub import snapshot_download
from transformers import CompressedTensorsConfig

path = snapshot_download("simonlehmann/clef-NVFP4")
sys.path.insert(0, path)
from joint_schema_model import load_release_model, systemone

model, processor = load_release_model(
    path, device="cuda", quantization_config=CompressedTensorsConfig(run_compressed=False)
)

Requires compressed-tensors in addition to torch and transformers.

License

Apache-2.0, following Cloudflare/clef and Qwen/Qwen3.8-27B. joint_schema_model.py, joint_head.safetensors, LICENSE and the tokenizer/processor files are redistributed unchanged from the Cloudflare release; clef_vllm.py and the quantized weights are new.

Downloads last month
-
Safetensors
Model size
20B params
Tensor type
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for simonlehmann/clef-NVFP4

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Quantized
(7)
this model