Image-Text-to-Text
Transformers
Safetensors
English
qwen3_5
text-generation-inference
llm-compressor
vllm
clef
fp8
w8a8
cloudflare
systemone
qwen3.8
post-train
image-text-to-typed-output
multimodal
structured-output
classification
custom-code
conversational
compressed-tensors
Instructions to use prithivMLmods/clef-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use prithivMLmods/clef-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="prithivMLmods/clef-FP8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("prithivMLmods/clef-FP8") model = AutoModelForMultimodalLM.from_pretrained("prithivMLmods/clef-FP8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use prithivMLmods/clef-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prithivMLmods/clef-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/clef-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/prithivMLmods/clef-FP8
- SGLang
How to use prithivMLmods/clef-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "prithivMLmods/clef-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/clef-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "prithivMLmods/clef-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/clef-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use prithivMLmods/clef-FP8 with Docker Model Runner:
docker model run hf.co/prithivMLmods/clef-FP8
|
Download README.md from prithivMLmods/clef-FP8: direct link, hf CLI and curl.
- Browser
- Download file 5.39 kB
-
https://huggingface.co/prithivMLmods/clef-FP8/resolve/main/README.md
- Command line
-
hf download hf://prithivMLmods/clef-FP8/README.md
-
curl -L -o README.md https://huggingface.co/prithivMLmods/clef-FP8/resolve/main/README.md
5.39 kB
| base_model: | |
| - Cloudflare/clef | |
| license: apache-2.0 | |
| language: | |
| - en | |
| pipeline_tag: image-text-to-text | |
| library_name: transformers | |
| tags: | |
| - text-generation-inference | |
| - llm-compressor | |
| - vllm | |
| - clef | |
| - fp8 | |
| - w8a8 | |
| - cloudflare | |
| - systemone | |
| - qwen3.8 | |
| - post-train | |
| - image-text-to-typed-output | |
| - multimodal | |
| - structured-output | |
| - classification | |
| - custom-code | |
| # **clef-FP8** | |
| FP8 (W8A8, dynamic) quantization of [Cloudflare/clef](https://huggingface.co/Cloudflare/clef), | |
| a 27B multimodal model that turns a state and a schema of typed questions into decisions. Clef | |
| reads text, JSON, images, or video and returns a probability for every allowed option of every | |
| question in a single forward pass, with no free-form generation and no output parsing. This repo quantizes only the backbone's linear layers. The vision encoder, embeddings, `lm_head`, | |
| and linear-attention layers are left in their original precision. For model behavior, input | |
| format, and the Jev/SystemOne API, see the | |
| [original Clef card](https://huggingface.co/Cloudflare/clef). | |
| ## Quantization | |
| | | | | |
| |---|---| | |
| | **Modality** | Image-Text-to-Text | | |
| | **Quantization scheme** | FP8_DYNAMIC (W8A8) | | |
| | **Weights** | FP8, per-channel | | |
| | **Activations** | FP8, per-token, dynamic | | |
| | **Calibration data** | Not required | | |
| | **Format** | compressed-tensors (safetensors) | | |
| | **Tooling** | [LLM Compressor](https://github.com/vllm-project/llm-compressor) | | |
| | **License** | Apache 2.0 | | |
| | Setting | Value | | |
| |---|---| | |
| | **targets** | `Linear` | | |
| | **ignore** | `lm_head`, `embed_tokens`, `visual`, `model.visual`, `linear_attn` | | |
| | **scheme** | `FP8_DYNAMIC` | | |
| | **bypass_divisibility_checks** | `false` | | |
| | **requires_calibration_data** | `false` | | |
| ### recipe.yaml | |
| ```yaml | |
| default_stage: | |
| default_modifiers: | |
| QuantizationModifier: | |
| targets: [Linear] | |
| ignore: ['re:.*lm_head', 're:.*embed_tokens$', 're:.*visual.*', 're:.*model.visual.*', | |
| 're:.*linear_attn.*'] | |
| scheme: FP8_DYNAMIC | |
| bypass_divisibility_checks: false | |
| requires_calibration_data: false | |
| ``` | |
| Because the scheme is `FP8_DYNAMIC`, weight scales are computed directly from the weights and | |
| activation scales are computed per token at runtime. No calibration dataset is needed. | |
| ## Usage | |
| Install `compressed-tensors` alongside `transformers` so the FP8 checkpoint can be loaded: | |
| ```bash | |
| pip install torch transformers compressed-tensors pillow | |
| ``` | |
| Usage is the same as for Clef: | |
| ```python | |
| import sys | |
| import torch | |
| from huggingface_hub import snapshot_download | |
| path = snapshot_download("prithivMLmods/clef-FP8") | |
| sys.path.insert(0, path) | |
| from joint_schema_model import collate_records, encode_record, load_release_model | |
| model, processor = load_release_model(path, device="cuda") | |
| record = { | |
| "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}}, | |
| "questions": { | |
| "status": { | |
| "type": "choice", | |
| "instructions": "What is the invoice status?", | |
| "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."}, | |
| }, | |
| "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"}, | |
| }, | |
| } | |
| encoded = encode_record(processor.tokenizer, record, processor=processor) | |
| batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda")) | |
| with torch.inference_mode(): | |
| logits = model(batch)[0] | |
| for question, question_logits in zip(encoded.questions, logits): | |
| probabilities = question_logits.float().softmax(-1).tolist() | |
| print(question.question_id, dict(zip(question.option_ids, probabilities))) | |
| ``` | |
| The `systemone(model, processor, request)` helper and image/video inputs work as described in the | |
| Clef card. | |
| **Hardware:** FP8 W8A8 compute needs a GPU with FP8 support (Ada, Hopper, or newer). On older | |
| GPUs, FP8 weights may only give memory savings, depending on the runtime. | |
| ## Reproducing the quantization | |
| ```python | |
| from transformers import AutoModelForImageTextToText, AutoProcessor | |
| from llmcompressor import oneshot | |
| src = "Cloudflare/clef" # local path from snapshot_download works too | |
| dst = "clef-FP8" | |
| model = AutoModelForImageTextToText.from_pretrained(src, torch_dtype="auto", device_map="auto") | |
| processor = AutoProcessor.from_pretrained(src) | |
| oneshot(model=model, recipe="recipe.yaml") # data-free, no dataset argument | |
| model.save_pretrained(dst, save_compressed=True) | |
| processor.save_pretrained(dst) | |
| ``` | |
| Then copy these files from the original Clef repo into `clef-FP8/` unchanged: | |
| `joint_head.safetensors`, `joint_head_config.json`, `joint_schema_model.py`. | |
| ## Notes and limitations | |
| - **Joint head:** the joint schema head is stored separately and is kept in its original | |
| precision, because the recipe quantizes only the backbone. If you modify `load_release_model` or | |
| re-export the model, make sure the head is not quantized. | |
| - **Excluded modules:** `linear_attn`, the vision encoder, embeddings, and `lm_head` stay | |
| unquantized, so the size reduction is somewhat smaller than a full 2x versus BF16. | |
| - **Loading path:** this checkpoint is intended for the custom `joint_schema_model.py` loader. | |
| General-purpose serving engines will not run the joint head. | |
| ## License | |
| Apache-2.0, following [Cloudflare/clef](https://huggingface.co/Cloudflare/clef) and the base model | |
| [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B). |