Instructions to use myroslavtryhubets/clef-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use myroslavtryhubets/clef-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="myroslavtryhubets/clef-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("myroslavtryhubets/clef-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("myroslavtryhubets/clef-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use myroslavtryhubets/clef-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "myroslavtryhubets/clef-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "myroslavtryhubets/clef-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/myroslavtryhubets/clef-NVFP4
- SGLang
How to use myroslavtryhubets/clef-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "myroslavtryhubets/clef-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "myroslavtryhubets/clef-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "myroslavtryhubets/clef-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "myroslavtryhubets/clef-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use myroslavtryhubets/clef-NVFP4 with Docker Model Runner:
docker model run hf.co/myroslavtryhubets/clef-NVFP4
Clef — NVFP4
An unofficial, community NVFP4 quantization of Cloudflare/clef,
a 27B multimodal decision model that turns a state and a schema of typed questions into a
probability for every allowed option in a single forward pass.
Quantized so the model runs on a single 16 GB consumer GPU (NVIDIA Blackwell, sm_120). In bf16 Clef is 55 GB and does not fit; this checkpoint holds 13.2 GiB of weights in VRAM.
Not affiliated with or endorsed by Cloudflare. The model, its architecture and
joint_schema_model.py are theirs; this repo only re-encodes the weights. Both the base model and
this derivative are Apache-2.0.
Usage
The quantized weights are not loaded by from_pretrained or by the base model's
load_release_model — those build the full-precision graph and need ≈55 GB (they OOM on a 16 GB
card). Use the included clef_rt.py, which keeps the decoder in NVFP4 on the GPU and the
embeddings / lm_head in host RAM:
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("myroslavtryhubets/clef-NVFP4")
sys.path.insert(0, path)
from clef_rt import Clef # ships in this repo
model = Clef(path, "nvfp4") # ≈13.2 GiB VRAM on a Blackwell card
out = model.probs([model.encode({
"state": "Our checkout is down and orders are blocked.",
"questions": {
"outage": {"type": "noul", "instructions": "Is a service down?"},
"team": {"type": "choice", "instructions": "Which team should handle it?",
"criteria": {"billing": "payments", "tech": "outages"}},
},
})])[0]
print(out) # {'outage': {'true': 0.98, ...}, 'team': {'billing': 0.63, ...}}
Clef(path, "nvfp4-w4a16") loads the same weights but keeps activations in bf16 (weights-only
4-bit) — slower, slightly more accurate, same VRAM.
What was quantized
- Scheme: NVFP4 — FP4 (e2m1) weights, an FP8 (e4m3) scale per 16-element block, one FP32 global
scale per tensor; W4A4 (weights and activations in 4-bit). Format: compressed-tensors
nvfp4-pack-quantized. - Tooling: llm-compressor 0.14 model-free PTQ for the weights; static activation global scales
from layer-by-layer bf16 calibration on 256 records (UltraChat + security-log text in Clef's own
prompt format,
static_minmax). - Kept in bf16:
lm_head,embed_tokens, the vision tower, the Gated-DeltaNetin_proj_a/in_proj_b/conv1d, all norms, and the joint schema head. - Fused global scales are shared across q/k/v, gate/up and the GDN
in_proj_qkv+in_proj_zpair (the last so vLLM'sin_proj_qkvzfusion stays consistent).
Quality — Decision Index 0.2.1
Measured on 11 benchmarks of the Decision Index 0.2.1 suite (20,337 requests), rebuilt byte-identical with the official reproduction kit and scored with its own scorer. The Clef bf16 (CF) column is Cloudflare's published result; This NVFP4 is this checkpoint on one RTX 5060 Ti 16 GB. Values are the coverage-adjusted percentages the board uses.
| Benchmark | Metric | Clef bf16 (CF) | This NVFP4 | Δ |
|---|---|---|---|---|
| BFCL | case exact accuracy | 98.5 | 98.6 | +0.1 |
| BANKING77 | macro-F1 | 94.2 | 93.7 | -0.5 |
| CLINC150+OOS | macro-F1 | 97.4 | 97.3 | -0.1 |
| ContractNLI | macro-F1 | 81.4 | 79.3 (2 OOM) | -2.1 |
| ANLI | macro-F1 | 69.8 | 69.5 | -0.3 |
| ARC-Challenge | accuracy | 97.7 | 97.4 | -0.3 |
| WinoGrande | accuracy | 93.5 | 92.0 | -1.5 |
| MuSR | accuracy | 83.5 | 82.2 | -1.3 |
| FinEntity | macro-F1 | 96.2 | 96.1 | -0.1 |
| CRUXEval | accuracy | 86.7 | 84.0 | -2.7 |
| PhishNChips | accuracy | 79.6 | 76.7 | -2.9 |
Mean absolute gap to Cloudflare: 1.08 points (worst −2.9, PhishNChips). Two of the 123 ContractNLI requests (≈10k tokens) do not fit 16 GB and are scored as wrong per the kit's coverage rule — about 1.6 pts of the ContractNLI gap.
Requirements
- GPU: NVIDIA Blackwell (sm_120), >=16 GB. NVFP4 matmul runs on the FP4 tensor cores via
torch._scaled_mm. Tested on torch 2.14.0+cu130, transformers 5.17.0, driver 580.95.05. - Verified by downloading this repo into a clean environment and loading with
clef_rt.py(13.2 GiB weights, correct decisions).
Caveats
- W4A4. The
clef-flashsibling loses ≈17 pts on CLINC150 under W4A4 (150 near-tied classes vs 4-bit activations). On this 27B model CLINC150 is unaffected (97.4 -> 97.3), but activation quantization is the main risk if you add tasks with many near-tied options. - vLLM: not supported for decisions. vLLM registers the
Qwen3_5backbone but has no Clef joint-head / decision path (checked against its model registry andqwen3_5.py), so it cannot produce typed decisions. vLLM's structured-generation work (PR #57250) targets DiffusionGemma's token-logprob mechanism, which is unrelated to Clef's routing head. Serving Clef in vLLM would require portingJointSchemaHeadas a custom model. Useclef_rt.py(above). - Quality was verified only through this runtime, not a third-party serving stack.
Credits
Model, architecture and decision API by Cloudflare (blog). Base model Qwen/Qwen3.8-27B. Quantization by @myroslavtryhubets.
- Downloads last month
- 35