clef-NVFP4

Unofficial NVFP4 quantization of Cloudflare/clef, the 27B decision model (Qwen3.8-27B backbone) that scores every option of every typed question (choice / score / true-false) in one forward pass. This repo is not affiliated with or endorsed by Cloudflare.

At 19 GB (from 55 GB in BF16) it runs on two 16 GB GPUs with tensor parallelism, or on one larger Blackwell device such as a DGX Spark (GB10). It ships with a small vLLM plugin and a Jev/SystemOne-compatible HTTP server (POST /v1/systemone), because stock vLLM cannot run Clef's custom joint-schema head.

Read before using

  • Accuracy cost is real: -1.4 points on our 7-task eval (82.8% vs 84.2% BF16; 94.7% top-1 agreement with BF16). That is more than the -0.5 points of our Clef-Flash NVFP4. On these task types, kurcontko/clef-flash-NVFP4 is both more accurate (83.6%) and about 5x faster; pick this 27B model for the workloads where Cloudflare's card shows Clef ahead of Clef-Flash (e.g. GSM8K, CLINC150, RAGTruth), and measure on your own data.
  • Text and JSON state only. The vision tower is kept in BF16 but has not been tested; the server rejects images / videos with HTTP 400 and runs vLLM with the vision tower disabled.
  • Use the bundled vLLM plugin, pinned to vLLM 0.28.x (it uses vLLM pooling internals). The checkpoint does not run efficiently in plain transformers.
  • Hardware: NVIDIA Blackwell (FP4 tensor cores). Tested on 2x RTX 5070 Ti 16 GB (SM 12.0, PCIe, no NVLink, tensor parallelism 2) and on 1x DGX Spark GB10 (SM 12.1, 128 GB unified memory, single device). Other Blackwell GPUs with 24 GB+ (RTX 5090, RTX PRO 6000, B200) should run it on one device but are untested.
  • Evaluated on 1,400 records from 7 public tasks with our own harness, not on Cloudflare's Decision Index.

Quickstart: SystemOne API server

hf download kurcontko/clef-NVFP4 --local-dir clef-NVFP4

docker run --gpus all --ipc=host -p 8000:8000 -v $PWD/clef-NVFP4:/model:ro \
  --entrypoint bash vllm/vllm-openai:v0.28.0 -c \
  "cp -r /model/vllm_plugin /tmp/p && pip install -q /tmp/p && clef-systemone --model /model --tp 2"

On one large GPU (DGX Spark, RTX 5090, B200) drop --tp 2. With --tp 2 the server turns FlashInfer autotuning off (vLLM 0.28 can deadlock autotuning NVFP4 kernels across tensor-parallel ranks); --flashinfer-autotune on|off overrides that. It warms up its kernels before /health reports ready (a few minutes from a cold container).

curl localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
  "model": "clef",
  "state": "Our checkout started returning errors and orders are blocked.",
  "questions": {
    "department": {"type": "choice", "instructions": "Which team should handle the message?",
                   "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}},
    "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
    "outage": {"type": "noul", "instructions": "Is a service down?"}}}'

Request and response bodies match the systemone() function of the original release (answers keyed by question ID: choice with confidence and probabilities, score with the expected score and legend, noul with the probability of true; plus usage). Also served: GET /v1/models, GET /health.

Quickstart: Python (offline vLLM)

# pip install ./clef-NVFP4/vllm_plugin   (in an environment with vllm==0.28.*)
import sys
from transformers import AutoTokenizer
from vllm import LLM, PoolingParams
from clef_vllm import ARCHITECTURE, TASK, schema_layout

path = "./clef-NVFP4"
sys.path.insert(0, path)
from joint_schema_model import encode_record


def main():
    # Keep `llm` local: vLLM shuts its engine process down when the LLM object is freed.
    llm = LLM(model=path, hf_overrides={"architectures": [ARCHITECTURE]}, runner="pooling",
              tensor_parallel_size=2, kernel_config={"enable_flashinfer_autotune": False},
              enable_prefix_caching=False, language_model_only=True, max_model_len=16384)
    tokenizer = AutoTokenizer.from_pretrained(path)
    record = {
        "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
        "questions": {
            "status": {"type": "choice", "instructions": "What is the invoice status?",
                       "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."}},
            "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
        },
    }
    encoded = encode_record(tokenizer, record)
    params = PoolingParams(task=TASK, extra_kwargs={"clef_questions": schema_layout(encoded)})
    output = llm.encode([{"prompt_token_ids": list(encoded.input_ids)}], pooling_params=[params], pooling_task=TASK)[0]

    logits, offset = output.outputs.data.float(), 0   # every option logit of every question, in order
    for question in encoded.questions:
        n = len(question.option_ids)
        print(question.question_id, dict(zip(question.option_ids, logits[offset:offset + n].softmax(-1).tolist())))
        offset += n


if __name__ == "__main__":
    main()

Results

Accuracy and agreement with BF16

1,400 records, 200 per task, from held-out splits: BANKING77 (test), ARC-Challenge (test), MMLU (test), HellaSwag (validation), ANLI R3 (test), BoolQ (validation), and a multi-question ANLI R1 set (3 questions per record). Acc is against gold labels; agree is top-1 agreement with BF16; KL is mean KL(BF16 ‖ NVFP4) per question. BF16 is the original release in transformers 5.10.2 (run on a DGX Spark); NVFP4 was measured in vLLM with real FP4 kernels.

Task BF16 acc NVFP4 acc Agree KL
BANKING77 (77-way intent) 96.0 95.5 99.5 0.011
ARC-Challenge 98.5 97.0 98.5 0.012
MMLU 93.0 92.0 95.5 0.022
HellaSwag 97.5 97.0 98.5 0.006
ANLI R3 51.0 47.5 88.5 0.046
BoolQ 90.5 90.5 99.0 0.009
Multi-question ANLI R1 73.5 71.2 90.8 0.030
All 84.2 82.8 94.7 0.022

A simulated NVFP4 run in transformers (weights dequantized, activations fake-quantized) gave 82.6% / 94.6% / 0.021, matching vLLM, so the gap is quantization error, not the serving path. For comparison on the same set, the original Clef-Flash scores 84.1% in BF16 and 83.6% as NVFP4.

Throughput

Same 1,400 records (717k tokens, ~512 per record).

Setup Hardware Records/s Tokens/s
BF16, transformers 5.10.2 (original release code) 1x DGX Spark (GB10) 2.0 1,002
NVFP4, vLLM + plugin, single device 1x DGX Spark (GB10) 4.2 2,139
NVFP4, vLLM + plugin, tensor parallel 2 2x RTX 5070 Ti 16 GB 8.4 4,288

On the GB10 the NVFP4 model agreed with BF16 on 94.1% of questions (KL 0.022, accuracy 82.7%), the same picture as on the RTX 5070 Ti pair; it also leaves room for ~747k tokens of KV cache at --gpu-memory-utilization 0.6.

Through the HTTP server on the 2x RTX 5070 Ti pair (clean install from this repo): 64 concurrent clients 8.3 req/s, 4,258 tok/s (p50 4.5 s, p95 26 s, queueing-bound); 8 concurrent clients 7.8 req/s, p50 611 ms, p95 3.4 s. Answers agreed with BF16 on 94.4 to 94.6% of questions.

Quantization details

  • Tool: llm-compressor 0.14.0 (compressed-tensors 0.19.0), QuantizationModifier(targets="Linear", scheme="NVFP4"), run on a DGX Spark in about 23 minutes.
  • Scheme: NVFP4 W4A4: FP4 (E2M1) weights and activations, 16-element groups with FP8 (E4M3) group scales and one FP32 global scale per tensor; activation global scales calibrated, group scales dynamic.
  • Calibration: 672 records (96 from the train split of each eval task, encoded exactly as at inference).
  • Kept in BF16: lm_head (the joint head reads its rows as lexical option embeddings), the vision tower, the Gated DeltaNet gate projections linear_attn.in_proj_a / in_proj_b, embeddings, norms, and the joint schema head.
  • Gated DeltaNet fused scale: vLLM fuses linear_attn.in_proj_qkv and in_proj_z into one GEMM, which needs a shared NVFP4 global scale; scripts/quantize.py adds that pair to llm-compressor's fused groups (all 48 GDN layers verified). Without it, vLLM silently mis-scales the qkv weights.
  • recipe.yaml is the exact recipe llm-compressor saved. Post-training quantization only: no quantization-aware training.

How the vLLM plugin works

vllm_plugin/ (package clef-vllm) registers ClefFlashForDecision, a pooling-model subclass of vLLM's Qwen3_5ForConditionalGeneration (the same class serves Clef and Clef-Flash). Its pooler collects each request's final hidden states across chunked prefill and runs the original JointSchemaHead (from joint_schema_model.py) on the full prompt; each request carries its question/option token spans in PoolingParams.extra_kwargs["clef_questions"]. Under tensor parallelism the vocab-sharded lm_head is gathered once into host memory on rank 0 and only each request's option-token rows are copied to the GPU, which leaves ~3.7 GiB of KV cache per 16 GB card instead of ~0.6 GiB.

Files

Path Purpose
model.safetensors, config.json, recipe.yaml Quantized backbone (19 GB, compressed-tensors format)
clef_head/ Joint schema head, BF16, unchanged from the original release (subfolder so vLLM does not load it as backbone weights)
joint_schema_model.py Original release code: encode_record, systemone_answer, the head module (unchanged)
tokenizer*.json, chat_template.jinja, processor_config.json, generation_config.json Tokenizer and processor (unchanged)
vllm_plugin/ vLLM plugin and clef-systemone server
scripts/ build_data.py (calibration/eval sets) and quantize.py (how this checkpoint was made)

License and attribution

Apache-2.0, following Cloudflare/clef (post-trained from Qwen/Qwen3.8-27B). All credit for the model goes to its original authors; this repository only adds the quantization, the vLLM plugin and the server. Not an official Cloudflare release.

Downloads last month
16
Safetensors
Model size
27B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kurcontko/clef-NVFP4

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Quantized
(20)
this model