Instructions to use simonlehmann/clef-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use simonlehmann/clef-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="simonlehmann/clef-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("simonlehmann/clef-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("simonlehmann/clef-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use simonlehmann/clef-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "simonlehmann/clef-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "simonlehmann/clef-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/simonlehmann/clef-NVFP4
- SGLang
How to use simonlehmann/clef-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "simonlehmann/clef-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "simonlehmann/clef-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "simonlehmann/clef-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "simonlehmann/clef-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use simonlehmann/clef-NVFP4 with Docker Model Runner:
docker model run hf.co/simonlehmann/clef-NVFP4
Clef — NVFP4
A mixed NVFP4 / FP8 quantization of Cloudflare/clef, the 27B multimodal decision model. Unofficial — not affiliated with or endorsed by Cloudflare.
Clef answers typed questions about a state (text, JSON, images, video) with one probability per allowed option, in a single prefill pass. Its value is the probabilities, so this quant was checked option by option against the BF16 release, not only on accuracy.
23 GB on disk (BF16 release: 55 GB); 20.2 GiB of weights in vLLM.
What is quantized
| Part | Format |
|---|---|
MLP gate/up/down_proj, layers 0–55 |
NVFP4 (W4A4, group 16, FP8 scales, static activation global scale) |
MLP gate/up/down_proj, layers 56–63 |
FP8 (per-channel weights, dynamic per-token activations) |
Full-attention q/k/v/o_proj; linear-attention in_proj_qkv/in_proj_z/out_proj |
FP8 (as above) |
Linear-attention in_proj_a/in_proj_b, norms, conv, embeddings |
BF16 |
| Vision encoder | BF16 |
lm_head |
BF16 — the joint head reads its rows directly as option embeddings |
Joint schema head (joint_head.safetensors) |
BF16, unchanged from the release |
The layout follows the community recipe for Qwen3.8-27B with one change: lm_head stays BF16.
The last eight layers stay FP8 on purpose: comparing the release against
Qwen/Qwen3.8-27B, the vision encoder and early layers
are byte-identical, and Clef's post-training lives in the upper layers.
Method: llm-compressor QuantizationModifier (round-to-nearest, no GPTQ), calibrated on 374
Clef-format records — 128 BANKING77 intent questions (77-way and 10-way schemas), 150 game-agent
state records with a choice question, and 96 Flickr30k photos with noul/choice/score
questions. Calibration and evaluation records are disjoint. recipe.yaml is the exact recipe.
Drift vs BF16
Same 615 records through the BF16 release and through this checkpoint on vLLM's real NVFP4/FP8
kernels (FlashInfer CUTLASS FP4 GEMM). agreement = top option identical to BF16. TV = mean
total-variation distance between the per-question distributions (0 = identical, 1 = disjoint).
| Eval set | Questions | Agreement | Mean TV | Accuracy BF16 → NVFP4 |
|---|---|---|---|---|
| BANKING77 test, full 77-way schema | 400 | 98.5% | 0.019 | 94.25% → 94.50% |
| Flickr30k photos, 4 mixed-type questions each | 256 | 96.9% | 0.017 | (no labels) |
Game-agent states, 2–4 option choice |
150 | 96.7% | 0.047 | 68.7% → 67.3% * |
| README invoice example | 2 | 100% | 0.0007 | — |
* graded against synthetic labels that the BF16 model itself matches only 69% of the time; read the agreement column, not the accuracy, for that row.
Disagreements are near-ties: in the calibration-time (fake-quant) run, the 15 flipped questions had a median BF16 top-1/top-2 margin of 0.13 (11 of 15 under 0.2), against 0.95 for questions that agreed.
Speed
One decision at a time (batch 1, the latency case) on an NVIDIA DGX Spark (GB10, 273 GB/s unified memory), vLLM 0.23.1, CUDA graphs on:
| Record | Input tokens | Median latency |
|---|---|---|
| text state, 1 question | 288 | 177 ms |
| 640 px image + 4 questions | 621 | 257 ms |
| BANKING77, 77 options | 1834 | 705 ms |
The joint head adds 3–6 ms of that. On this machine the floor (~170 ms) is set by reading the weights once per pass; GPUs with more memory bandwidth should be considerably quicker (not measured here).
Usage
The files, record format and systemone API are the same as the BF16 release; see the
Clef model card for the full input format.
vLLM (recommended — real FP4 kernels)
clef_vllm.py runs the backbone in vLLM as a pooling model (final hidden state of every token)
and Clef's joint head on top, in the same process. Needs a Blackwell GPU for NVFP4 and a vLLM build
with Qwen3_5ForConditionalGeneration (tested: 0.23.1 nightly).
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("simonlehmann/clef-NVFP4")
sys.path.insert(0, path)
from clef_vllm import ClefVLLM
clef = ClefVLLM(path, max_model_len=16384, gpu_memory_utilization=0.4)
response = clef.systemone({
"model": "clef",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
})
print(response["answers"])
# Many records at once, images included (PIL), batched through vLLM:
# clef.probabilities([record, ...]) -> [{question_id: {option_id: probability}}, ...]
max_images / max_videos in the constructor set vLLM's per-request media limits (default 1 image).
transformers (compatibility — decompresses to BF16)
transformers' compressed-tensors integration cannot run the mixed NVFP4/FP8 layers compressed, so
load with run_compressed=False. This decompresses to BF16 at load time (~57 GB of GPU memory, no
speed-up): useful to check outputs, not to save memory.
import sys
from huggingface_hub import snapshot_download
from transformers import CompressedTensorsConfig
path = snapshot_download("simonlehmann/clef-NVFP4")
sys.path.insert(0, path)
from joint_schema_model import load_release_model, systemone
model, processor = load_release_model(
path, device="cuda", quantization_config=CompressedTensorsConfig(run_compressed=False)
)
Requires compressed-tensors in addition to torch and transformers.
License
Apache-2.0, following Cloudflare/clef and
Qwen/Qwen3.8-27B. joint_schema_model.py,
joint_head.safetensors, LICENSE and the tokenizer/processor files are redistributed unchanged
from the Cloudflare release; clef_vllm.py and the quantized weights are new.
- Downloads last month
- -