Instructions to use alpha-x-ai/clef-flash-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use alpha-x-ai/clef-flash-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="alpha-x-ai/clef-flash-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("alpha-x-ai/clef-flash-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("alpha-x-ai/clef-flash-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use alpha-x-ai/clef-flash-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "alpha-x-ai/clef-flash-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alpha-x-ai/clef-flash-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/alpha-x-ai/clef-flash-NVFP4
- SGLang
How to use alpha-x-ai/clef-flash-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "alpha-x-ai/clef-flash-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alpha-x-ai/clef-flash-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "alpha-x-ai/clef-flash-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alpha-x-ai/clef-flash-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use alpha-x-ai/clef-flash-NVFP4 with Docker Model Runner:
docker model run hf.co/alpha-x-ai/clef-flash-NVFP4
clef-flash-NVFP4
Model Overview
- Model Architecture: Cloudflare/clef-flash
(Qwen3.5-9B backbone with vision encoder, plus Clef's joint schema head)
- Input: a state (text, JSON, images or video) and a schema of typed questions
- Output: a probability for every allowed option of every question
- Model Optimizations:
- Weight quantization: NVFP4 and FP8
- Activation quantization: NVFP4 and FP8
- Out-of-scope: free-form text generation. Clef is not a chat model; loading the backbone with a
generation API produces meaningless text. Use the bundled
clef_vllm.py. - Release Date: 2026-10-02
- Version: 1.0
- Model Developers: alpha-x-ai (unofficial; not affiliated with or endorsed by Cloudflare)
This model is a quantized version of Cloudflare/clef-flash for Blackwell GPUs. On an RTX 5090 it takes 53% of the BF16 release's latency, and its probabilities stay close to BF16: mean KL divergence 0.0016 over 1,358 held-out questions, with the same top option on 98.9% of them.
Model Optimizations
The checkpoint is 10.3 GB (BF16 release: 19.1 GB), and the backbone takes 7.6 GiB of GPU memory in vLLM (BF16: 15.8 GiB).
Not every layer is quantized the same way. Each of the 64 quantizable units (32 MLPs, 24 Gated DeltaNet linear-attention blocks, 8 full-attention blocks) was quantized to NVFP4 on its own and scored by how far it moved the output probabilities from BF16. Units were then switched to NVFP4 in order of that damage per parameter, and the rest were kept at FP8. The sensitive units are in the early and middle layers; the last layers are among the least sensitive.
| Part | Format |
|---|---|
MLP gate/up/down_proj, layers 0–4 and 15–31 |
NVFP4 |
MLP gate/up/down_proj, layers 5–14 |
FP8 |
Linear-attention in_proj_qkv/in_proj_z/out_proj, layers 17, 18, 20–22, 24–26, 28–30 |
NVFP4 |
Linear-attention in_proj_qkv/in_proj_z/out_proj, layers 0–2, 4–6, 8–10, 12–14, 16 |
FP8 |
Full-attention q/k/v/o_proj, layers 19, 23, 27, 31 |
NVFP4 |
Full-attention q/k/v/o_proj, layers 3, 7, 11, 15 |
FP8 |
Linear-attention in_proj_a/in_proj_b, conv, norms, embeddings |
BF16 |
| Vision encoder | BF16 |
lm_head |
BF16 |
Joint schema head (joint_head.safetensors) |
BF16, unchanged from the release |
- NVFP4: W4A4, FP4 E2M1 values in groups of 16 with FP8 E4M3 group scales. The weight global scale is per tensor and shared within each group of layers that vLLM runs as one GEMM. The activation global scale is static, from calibration.
- FP8: per-channel weights, dynamic per-token activations.
lm_headstays BF16 because the joint head reads its rows directly as option embeddings. It is not executed when the backbone runs as a pooling model, so quantizing it would not save time.
layout.json lists the layout and recipe.yaml is the exact llm-compressor recipe.
Deployment
Use with vLLM
The NVFP4 kernels require a Blackwell GPU. Tested with vLLM 0.30.0 on an RTX 5090 (sm_120) and a DGX Spark (GB10, sm_121).
clef_vllm.py runs the backbone in vLLM as a pooling model (the final hidden state of every token) and
the joint schema head on top, in the same process.
pip install vllm huggingface_hub pillow
import os
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("alpha-x-ai/clef-flash-NVFP4")
os.environ["CLEF_MODEL_PATH"] = path
sys.path.insert(0, path)
from clef_vllm import ClefVLLM
clef = ClefVLLM(path, max_model_len=16384, gpu_memory_utilization=0.6)
response = clef.systemone({
"model": "clef-flash",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
})
print(response["answers"])
# Several records, images included (PIL), batched through vLLM:
# clef.probabilities([record, ...]) -> [{question_id: {option_id: probability}}, ...]
The request format and the systemone response are the same as for the BF16 release; see the
Clef-Flash model card for the input format and question
types. ClefVLLM(path, kv_cache_gib=1.5) sets a fixed KV cache size instead of a memory fraction, which
is useful when other processes share the GPU.
Use with transformers
transformers cannot run the mixed NVFP4/FP8 layers compressed. Loading with run_compressed=False
decompresses the weights to BF16. This is useful for checking outputs, but it saves no memory and gives
no speed-up. Requires compressed-tensors and accelerate.
import sys
from huggingface_hub import snapshot_download
from transformers import CompressedTensorsConfig
path = snapshot_download("alpha-x-ai/clef-flash-NVFP4")
sys.path.insert(0, path)
from joint_schema_model import load_release_model, systemone
model, processor = load_release_model(
path, device="cuda", quantization_config=CompressedTensorsConfig(run_compressed=False)
)
Creation
This model was created with LLM Compressor 0.14.0 QuantizationModifier (round-to-nearest, no GPTQ)
and recipe.yaml. Calibration used 374 records in Clef's own input format:
- 128 BANKING77 train utterances, alternating the full 77-way schema and 10-way subsets
- 96 Flickr30k photos with
noul,choiceandscorequestions - 150 synthetic game-agent state records with a 2–4 option
choicequestion
vLLM runs in_proj_qkv and in_proj_z as one GEMM with a single NVFP4 weight global scale, but
LLM Compressor 0.14.0 only shares that scale across q/k/v and gate/up. Without the extra fused group
in the code below, vLLM dequantizes one of the two shards with the other's scale. In our measurements
that doubled the KL divergence of a layout with NVFP4 linear attention.
Calibration and quantization code
import sys
import torch
from datasets import Dataset
from llmcompressor import oneshot
from llmcompressor.observers import helpers as observer_helpers
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
SRC = "path/to/Cloudflare/clef-flash" # local snapshot
DST = "clef-flash-NVFP4"
sys.path.insert(0, SRC)
from joint_schema_model import collate_records, encode_record
observer_helpers.FUSED_LAYER_NAMES.append(("in_proj_qkv", "in_proj_z"))
processor = AutoProcessor.from_pretrained(SRC)
model = Qwen3_5ForConditionalGeneration.from_pretrained(SRC, dtype=torch.bfloat16, device_map={"": "cuda"})
model.config.use_cache = False
records = [...] # 374 Clef records: {"state": ..., "questions": {...}, optional "images": [PIL.Image]}
rows = []
for record in records:
encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cpu"))
row = {"input_ids": batch["input_ids"][0].tolist(), "attention_mask": batch["attention_mask"][0].tolist()}
row.update({key: value.tolist() for key, value in batch["media"].items()})
rows.append(row)
keys = sorted({key for row in rows for key in row})
dataset = Dataset.from_list([{key: row.get(key) for key in keys} for row in rows])
def collate(batch):
item = {}
for key, value in batch[0].items():
if value is None:
continue
if key == "pixel_values":
item[key] = torch.tensor(value, dtype=torch.bfloat16)
elif key in ("input_ids", "attention_mask"):
item[key] = torch.tensor(value).unsqueeze(0)
else: # image_grid_thw, mm_token_type_ids
item[key] = torch.tensor(value)
return item
oneshot(model=model, dataset=dataset, recipe="recipe.yaml", data_collator=collate,
num_calibration_samples=len(rows), max_seq_length=16384, pipeline="basic")
model.save_pretrained(DST, save_compressed=True)
Then copy joint_head.safetensors, joint_head_config.json, joint_schema_model.py, LICENSE and the
tokenizer and processor files from the release unchanged, and write model.safetensors.index.json. A
single-shard save has no index, and without one vLLM also tries to load joint_head.safetensors as
backbone weights.
Evaluation
The BF16 release and this model were run through vLLM 0.30.0 on the same 686 test records (1,358 questions). None of the test records were used for calibration or for the per-unit sensitivity scores.
| Test group | Records | Questions |
|---|---|---|
| BANKING77 test, full 77-way schema | 300 | 300 |
| CLINC150 test, 151-way schema (150 intents + out-of-scope) | 150 | 150 |
| Flickr30k photos, 4 mixed-type questions each | 100 | 400 |
| BANKING77 messages, 4 mixed-type questions each | 100 | 400 |
| Synthetic JSON event logs of 2k–12k tokens, 3 questions each | 36 | 108 |
Agreement with BF16
KL is KL(BF16 ‖ this model) per question, and TV is the total-variation distance between the two distributions (0 = identical, 1 = disjoint). The first table compares each GPU against BF16 on the same GPU.
| GPU | Same top option | Mean KL | Mean TV |
|---|---|---|---|
| RTX 5090 | 98.9% | 0.0016 | 0.012 |
| DGX Spark | 99.0% | 0.0014 | 0.011 |
By test group, on the RTX 5090:
| Test group | Questions | Same top option | Mean KL |
|---|---|---|---|
| BANKING77, 77-way | 300 | 99.0% | 0.0016 |
| CLINC150, 151-way | 150 | 98.7% | 0.0028 |
| Flickr30k photos | 400 | 99.3% | 0.0007 |
| BANKING77, mixed questions | 400 | 98.8% | 0.0016 |
| JSON event logs, 2k–12k tokens | 108 | 98.1% | 0.0026 |
For reference, on the DGX Spark, quantizing every linear layer to FP8 gives mean KL 0.0006. A layout with MLP layers 0–27 in NVFP4 and the rest in FP8 gives 0.0038 at about the same speed as this model.
Accuracy
| Eval set | Metric | Cloudflare/clef-flash | clef-flash-NVFP4 (this model) |
|---|---|---|---|
| BANKING77 test (300) | accuracy | 94.67 | 93.33 |
| CLINC150 test (150) | accuracy | 97.33 | 96.67 |
Measured on the DGX Spark. The differences are 4 of 300 questions (BANKING77) and 1 of 150 (CLINC150).
Latency
Batch 1 with vLLM 0.30.0, median of 30 runs after warmup. Times include preprocessing and the joint head.
| Input | Tokens | RTX 5090, BF16 | RTX 5090, this model | DGX Spark, BF16 | DGX Spark, this model |
|---|---|---|---|---|---|
| Text, 3 questions | 362 | 48.8 ms | 32.2 ms | 107.7 ms | 57.2 ms |
| 640 px image, 4 questions | 710 | 83.3 ms | 52.2 ms | 168.5 ms | 93.0 ms |
| BANKING77, 77 options | 1,838 | 163.7 ms | 83.6 ms | 405.5 ms | 215.0 ms |
| JSON event log, 3 questions | 3,672 | 303.5 ms | 134.5 ms | 777.2 ms | 472.3 ms |
| JSON event log, 3 questions | 6,961 | 585.8 ms | 257.9 ms | 1,403.6 ms | 1,352.7 ms |
On the DGX Spark the FP8 layers run slower than BF16 once a prefill chunk holds several thousand tokens, so the gain over BF16 shrinks on the longest input. The RTX 5090 shows no such effect.
Limitations
- Calibration contains no inputs longer than about 2,000 tokens, and the test set has only 108 questions on long inputs. Long inputs have the lowest top-option agreement of all groups (98.1% on the RTX 5090, 96.3% on the DGX Spark).
- Video inputs were not evaluated.
- Round-to-nearest only; GPTQ was not tried.
- Batched throughput was not measured.
- Answers whose top two options are close in BF16 can change after quantization. Check low-margin answers on your own data if your application acts on them.
License
Apache-2.0, following Cloudflare/clef-flash and Qwen/Qwen3.5-9B.
joint_schema_model.py, joint_head.safetensors, LICENSE and the tokenizer and processor files are
redistributed unchanged from the Cloudflare release. clef_vllm.py and the quantized weights are new.
- Downloads last month
- 47