--- base_model: - Cloudflare/clef license: apache-2.0 language: - en pipeline_tag: image-text-to-text library_name: transformers tags: - text-generation-inference - llm-compressor - vllm - clef - fp8 - w8a8 - cloudflare - systemone - qwen3.8 - post-train - image-text-to-typed-output - multimodal - structured-output - classification - custom-code --- # **clef-FP8** FP8 (W8A8, dynamic) quantization of [Cloudflare/clef](https://huggingface.co/Cloudflare/clef), a 27B multimodal model that turns a state and a schema of typed questions into decisions. Clef reads text, JSON, images, or video and returns a probability for every allowed option of every question in a single forward pass, with no free-form generation and no output parsing. This repo quantizes only the backbone's linear layers. The vision encoder, embeddings, `lm_head`, and linear-attention layers are left in their original precision. For model behavior, input format, and the Jev/SystemOne API, see the [original Clef card](https://huggingface.co/Cloudflare/clef). ## Quantization | | | |---|---| | **Modality** | Image-Text-to-Text | | **Quantization scheme** | FP8_DYNAMIC (W8A8) | | **Weights** | FP8, per-channel | | **Activations** | FP8, per-token, dynamic | | **Calibration data** | Not required | | **Format** | compressed-tensors (safetensors) | | **Tooling** | [LLM Compressor](https://github.com/vllm-project/llm-compressor) | | **License** | Apache 2.0 | | Setting | Value | |---|---| | **targets** | `Linear` | | **ignore** | `lm_head`, `embed_tokens`, `visual`, `model.visual`, `linear_attn` | | **scheme** | `FP8_DYNAMIC` | | **bypass_divisibility_checks** | `false` | | **requires_calibration_data** | `false` | ### recipe.yaml ```yaml default_stage: default_modifiers: QuantizationModifier: targets: [Linear] ignore: ['re:.*lm_head', 're:.*embed_tokens$', 're:.*visual.*', 're:.*model.visual.*', 're:.*linear_attn.*'] scheme: FP8_DYNAMIC bypass_divisibility_checks: false requires_calibration_data: false ``` Because the scheme is `FP8_DYNAMIC`, weight scales are computed directly from the weights and activation scales are computed per token at runtime. No calibration dataset is needed. ## Usage Install `compressed-tensors` alongside `transformers` so the FP8 checkpoint can be loaded: ```bash pip install torch transformers compressed-tensors pillow ``` Usage is the same as for Clef: ```python import sys import torch from huggingface_hub import snapshot_download path = snapshot_download("prithivMLmods/clef-FP8") sys.path.insert(0, path) from joint_schema_model import collate_records, encode_record, load_release_model model, processor = load_release_model(path, device="cuda") record = { "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}}, "questions": { "status": { "type": "choice", "instructions": "What is the invoice status?", "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."}, }, "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"}, }, } encoded = encode_record(processor.tokenizer, record, processor=processor) batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda")) with torch.inference_mode(): logits = model(batch)[0] for question, question_logits in zip(encoded.questions, logits): probabilities = question_logits.float().softmax(-1).tolist() print(question.question_id, dict(zip(question.option_ids, probabilities))) ``` The `systemone(model, processor, request)` helper and image/video inputs work as described in the Clef card. **Hardware:** FP8 W8A8 compute needs a GPU with FP8 support (Ada, Hopper, or newer). On older GPUs, FP8 weights may only give memory savings, depending on the runtime. ## Reproducing the quantization ```python from transformers import AutoModelForImageTextToText, AutoProcessor from llmcompressor import oneshot src = "Cloudflare/clef" # local path from snapshot_download works too dst = "clef-FP8" model = AutoModelForImageTextToText.from_pretrained(src, torch_dtype="auto", device_map="auto") processor = AutoProcessor.from_pretrained(src) oneshot(model=model, recipe="recipe.yaml") # data-free, no dataset argument model.save_pretrained(dst, save_compressed=True) processor.save_pretrained(dst) ``` Then copy these files from the original Clef repo into `clef-FP8/` unchanged: `joint_head.safetensors`, `joint_head_config.json`, `joint_schema_model.py`. ## Notes and limitations - **Joint head:** the joint schema head is stored separately and is kept in its original precision, because the recipe quantizes only the backbone. If you modify `load_release_model` or re-export the model, make sure the head is not quantized. - **Excluded modules:** `linear_attn`, the vision encoder, embeddings, and `lm_head` stay unquantized, so the size reduction is somewhat smaller than a full 2x versus BF16. - **Loading path:** this checkpoint is intended for the custom `joint_schema_model.py` loader. General-purpose serving engines will not run the joint head. ## License Apache-2.0, following [Cloudflare/clef](https://huggingface.co/Cloudflare/clef) and the base model [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B).