clef-flash-NVFP4

Model Overview

  • Model Architecture: Cloudflare/clef-flash (Qwen3.5-9B backbone with vision encoder, plus Clef's joint schema head)
    • Input: a state (text, JSON, images or video) and a schema of typed questions
    • Output: a probability for every allowed option of every question
  • Model Optimizations:
    • Weight quantization: NVFP4 and FP8
    • Activation quantization: NVFP4 and FP8
  • Out-of-scope: free-form text generation. Clef is not a chat model; loading the backbone with a generation API produces meaningless text. Use the bundled clef_vllm.py.
  • Release Date: 2026-10-02
  • Version: 1.0
  • Model Developers: alpha-x-ai (unofficial; not affiliated with or endorsed by Cloudflare)

This model is a quantized version of Cloudflare/clef-flash for Blackwell GPUs. On an RTX 5090 it takes 53% of the BF16 release's latency, and its probabilities stay close to BF16: mean KL divergence 0.0016 over 1,358 held-out questions, with the same top option on 98.9% of them.

Model Optimizations

The checkpoint is 10.3 GB (BF16 release: 19.1 GB), and the backbone takes 7.6 GiB of GPU memory in vLLM (BF16: 15.8 GiB).

Not every layer is quantized the same way. Each of the 64 quantizable units (32 MLPs, 24 Gated DeltaNet linear-attention blocks, 8 full-attention blocks) was quantized to NVFP4 on its own and scored by how far it moved the output probabilities from BF16. Units were then switched to NVFP4 in order of that damage per parameter, and the rest were kept at FP8. The sensitive units are in the early and middle layers; the last layers are among the least sensitive.

Part Format
MLP gate/up/down_proj, layers 0–4 and 15–31 NVFP4
MLP gate/up/down_proj, layers 5–14 FP8
Linear-attention in_proj_qkv/in_proj_z/out_proj, layers 17, 18, 20–22, 24–26, 28–30 NVFP4
Linear-attention in_proj_qkv/in_proj_z/out_proj, layers 0–2, 4–6, 8–10, 12–14, 16 FP8
Full-attention q/k/v/o_proj, layers 19, 23, 27, 31 NVFP4
Full-attention q/k/v/o_proj, layers 3, 7, 11, 15 FP8
Linear-attention in_proj_a/in_proj_b, conv, norms, embeddings BF16
Vision encoder BF16
lm_head BF16
Joint schema head (joint_head.safetensors) BF16, unchanged from the release
  • NVFP4: W4A4, FP4 E2M1 values in groups of 16 with FP8 E4M3 group scales. The weight global scale is per tensor and shared within each group of layers that vLLM runs as one GEMM. The activation global scale is static, from calibration.
  • FP8: per-channel weights, dynamic per-token activations.
  • lm_head stays BF16 because the joint head reads its rows directly as option embeddings. It is not executed when the backbone runs as a pooling model, so quantizing it would not save time.

layout.json lists the layout and recipe.yaml is the exact llm-compressor recipe.

Deployment

Use with vLLM

The NVFP4 kernels require a Blackwell GPU. Tested with vLLM 0.30.0 on an RTX 5090 (sm_120) and a DGX Spark (GB10, sm_121).

clef_vllm.py runs the backbone in vLLM as a pooling model (the final hidden state of every token) and the joint schema head on top, in the same process.

pip install vllm huggingface_hub pillow
import os
import sys

from huggingface_hub import snapshot_download

path = snapshot_download("alpha-x-ai/clef-flash-NVFP4")
os.environ["CLEF_MODEL_PATH"] = path
sys.path.insert(0, path)
from clef_vllm import ClefVLLM

clef = ClefVLLM(path, max_model_len=16384, gpu_memory_utilization=0.6)
response = clef.systemone({
    "model": "clef-flash",
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle the message?",
            "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
        },
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
        "outage": {"type": "noul", "instructions": "Is a service down?"},
    },
})
print(response["answers"])

# Several records, images included (PIL), batched through vLLM:
# clef.probabilities([record, ...]) -> [{question_id: {option_id: probability}}, ...]

The request format and the systemone response are the same as for the BF16 release; see the Clef-Flash model card for the input format and question types. ClefVLLM(path, kv_cache_gib=1.5) sets a fixed KV cache size instead of a memory fraction, which is useful when other processes share the GPU.

Use with transformers

transformers cannot run the mixed NVFP4/FP8 layers compressed. Loading with run_compressed=False decompresses the weights to BF16. This is useful for checking outputs, but it saves no memory and gives no speed-up. Requires compressed-tensors and accelerate.

import sys

from huggingface_hub import snapshot_download
from transformers import CompressedTensorsConfig

path = snapshot_download("alpha-x-ai/clef-flash-NVFP4")
sys.path.insert(0, path)
from joint_schema_model import load_release_model, systemone

model, processor = load_release_model(
    path, device="cuda", quantization_config=CompressedTensorsConfig(run_compressed=False)
)

Creation

This model was created with LLM Compressor 0.14.0 QuantizationModifier (round-to-nearest, no GPTQ) and recipe.yaml. Calibration used 374 records in Clef's own input format:

  • 128 BANKING77 train utterances, alternating the full 77-way schema and 10-way subsets
  • 96 Flickr30k photos with noul, choice and score questions
  • 150 synthetic game-agent state records with a 2–4 option choice question

vLLM runs in_proj_qkv and in_proj_z as one GEMM with a single NVFP4 weight global scale, but LLM Compressor 0.14.0 only shares that scale across q/k/v and gate/up. Without the extra fused group in the code below, vLLM dequantizes one of the two shards with the other's scale. In our measurements that doubled the KL divergence of a layout with NVFP4 linear attention.

Calibration and quantization code
import sys

import torch
from datasets import Dataset
from llmcompressor import oneshot
from llmcompressor.observers import helpers as observer_helpers
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

SRC = "path/to/Cloudflare/clef-flash"   # local snapshot
DST = "clef-flash-NVFP4"
sys.path.insert(0, SRC)
from joint_schema_model import collate_records, encode_record

observer_helpers.FUSED_LAYER_NAMES.append(("in_proj_qkv", "in_proj_z"))

processor = AutoProcessor.from_pretrained(SRC)
model = Qwen3_5ForConditionalGeneration.from_pretrained(SRC, dtype=torch.bfloat16, device_map={"": "cuda"})
model.config.use_cache = False

records = [...]  # 374 Clef records: {"state": ..., "questions": {...}, optional "images": [PIL.Image]}

rows = []
for record in records:
    encoded = encode_record(processor.tokenizer, record, processor=processor)
    batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cpu"))
    row = {"input_ids": batch["input_ids"][0].tolist(), "attention_mask": batch["attention_mask"][0].tolist()}
    row.update({key: value.tolist() for key, value in batch["media"].items()})
    rows.append(row)
keys = sorted({key for row in rows for key in row})
dataset = Dataset.from_list([{key: row.get(key) for key in keys} for row in rows])

def collate(batch):
    item = {}
    for key, value in batch[0].items():
        if value is None:
            continue
        if key == "pixel_values":
            item[key] = torch.tensor(value, dtype=torch.bfloat16)
        elif key in ("input_ids", "attention_mask"):
            item[key] = torch.tensor(value).unsqueeze(0)
        else:  # image_grid_thw, mm_token_type_ids
            item[key] = torch.tensor(value)
    return item

oneshot(model=model, dataset=dataset, recipe="recipe.yaml", data_collator=collate,
        num_calibration_samples=len(rows), max_seq_length=16384, pipeline="basic")
model.save_pretrained(DST, save_compressed=True)

Then copy joint_head.safetensors, joint_head_config.json, joint_schema_model.py, LICENSE and the tokenizer and processor files from the release unchanged, and write model.safetensors.index.json. A single-shard save has no index, and without one vLLM also tries to load joint_head.safetensors as backbone weights.

Evaluation

The BF16 release and this model were run through vLLM 0.30.0 on the same 686 test records (1,358 questions). None of the test records were used for calibration or for the per-unit sensitivity scores.

Test group Records Questions
BANKING77 test, full 77-way schema 300 300
CLINC150 test, 151-way schema (150 intents + out-of-scope) 150 150
Flickr30k photos, 4 mixed-type questions each 100 400
BANKING77 messages, 4 mixed-type questions each 100 400
Synthetic JSON event logs of 2k–12k tokens, 3 questions each 36 108

Agreement with BF16

KL is KL(BF16 ‖ this model) per question, and TV is the total-variation distance between the two distributions (0 = identical, 1 = disjoint). The first table compares each GPU against BF16 on the same GPU.

GPU Same top option Mean KL Mean TV
RTX 5090 98.9% 0.0016 0.012
DGX Spark 99.0% 0.0014 0.011

By test group, on the RTX 5090:

Test group Questions Same top option Mean KL
BANKING77, 77-way 300 99.0% 0.0016
CLINC150, 151-way 150 98.7% 0.0028
Flickr30k photos 400 99.3% 0.0007
BANKING77, mixed questions 400 98.8% 0.0016
JSON event logs, 2k–12k tokens 108 98.1% 0.0026

For reference, on the DGX Spark, quantizing every linear layer to FP8 gives mean KL 0.0006. A layout with MLP layers 0–27 in NVFP4 and the rest in FP8 gives 0.0038 at about the same speed as this model.

Accuracy

Eval set Metric Cloudflare/clef-flash clef-flash-NVFP4 (this model)
BANKING77 test (300) accuracy 94.67 93.33
CLINC150 test (150) accuracy 97.33 96.67

Measured on the DGX Spark. The differences are 4 of 300 questions (BANKING77) and 1 of 150 (CLINC150).

Latency

Batch 1 with vLLM 0.30.0, median of 30 runs after warmup. Times include preprocessing and the joint head.

Input Tokens RTX 5090, BF16 RTX 5090, this model DGX Spark, BF16 DGX Spark, this model
Text, 3 questions 362 48.8 ms 32.2 ms 107.7 ms 57.2 ms
640 px image, 4 questions 710 83.3 ms 52.2 ms 168.5 ms 93.0 ms
BANKING77, 77 options 1,838 163.7 ms 83.6 ms 405.5 ms 215.0 ms
JSON event log, 3 questions 3,672 303.5 ms 134.5 ms 777.2 ms 472.3 ms
JSON event log, 3 questions 6,961 585.8 ms 257.9 ms 1,403.6 ms 1,352.7 ms

On the DGX Spark the FP8 layers run slower than BF16 once a prefill chunk holds several thousand tokens, so the gain over BF16 shrinks on the longest input. The RTX 5090 shows no such effect.

Limitations

  • Calibration contains no inputs longer than about 2,000 tokens, and the test set has only 108 questions on long inputs. Long inputs have the lowest top-option agreement of all groups (98.1% on the RTX 5090, 96.3% on the DGX Spark).
  • Video inputs were not evaluated.
  • Round-to-nearest only; GPTQ was not tried.
  • Batched throughput was not measured.
  • Answers whose top two options are close in BF16 can change after quantization. Check low-margin answers on your own data if your application acts on them.

License

Apache-2.0, following Cloudflare/clef-flash and Qwen/Qwen3.5-9B.

joint_schema_model.py, joint_head.safetensors, LICENSE and the tokenizer and processor files are redistributed unchanged from the Cloudflare release. clef_vllm.py and the quantized weights are new.

Downloads last month
47
Safetensors
Model size
9B params
Tensor type
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alpha-x-ai/clef-flash-NVFP4

Finetuned
Qwen/Qwen3.5-9B
Quantized
(30)
this model