clef-FP8 / README.md
prithivMLmods's picture
Update README.md
cea2293 verified
|
Raw History Blame Contribute Delete
5.39 kB
---
base_model:
- Cloudflare/clef
license: apache-2.0
language:
- en
pipeline_tag: image-text-to-text
library_name: transformers
tags:
- text-generation-inference
- llm-compressor
- vllm
- clef
- fp8
- w8a8
- cloudflare
- systemone
- qwen3.8
- post-train
- image-text-to-typed-output
- multimodal
- structured-output
- classification
- custom-code
---
# **clef-FP8**
FP8 (W8A8, dynamic) quantization of [Cloudflare/clef](https://huggingface.co/Cloudflare/clef),
a 27B multimodal model that turns a state and a schema of typed questions into decisions. Clef
reads text, JSON, images, or video and returns a probability for every allowed option of every
question in a single forward pass, with no free-form generation and no output parsing. This repo quantizes only the backbone's linear layers. The vision encoder, embeddings, `lm_head`,
and linear-attention layers are left in their original precision. For model behavior, input
format, and the Jev/SystemOne API, see the
[original Clef card](https://huggingface.co/Cloudflare/clef).
## Quantization
| | |
|---|---|
| **Modality** | Image-Text-to-Text |
| **Quantization scheme** | FP8_DYNAMIC (W8A8) |
| **Weights** | FP8, per-channel |
| **Activations** | FP8, per-token, dynamic |
| **Calibration data** | Not required |
| **Format** | compressed-tensors (safetensors) |
| **Tooling** | [LLM Compressor](https://github.com/vllm-project/llm-compressor) |
| **License** | Apache 2.0 |
| Setting | Value |
|---|---|
| **targets** | `Linear` |
| **ignore** | `lm_head`, `embed_tokens`, `visual`, `model.visual`, `linear_attn` |
| **scheme** | `FP8_DYNAMIC` |
| **bypass_divisibility_checks** | `false` |
| **requires_calibration_data** | `false` |
### recipe.yaml
```yaml
default_stage:
default_modifiers:
QuantizationModifier:
targets: [Linear]
ignore: ['re:.*lm_head', 're:.*embed_tokens$', 're:.*visual.*', 're:.*model.visual.*',
're:.*linear_attn.*']
scheme: FP8_DYNAMIC
bypass_divisibility_checks: false
requires_calibration_data: false
```
Because the scheme is `FP8_DYNAMIC`, weight scales are computed directly from the weights and
activation scales are computed per token at runtime. No calibration dataset is needed.
## Usage
Install `compressed-tensors` alongside `transformers` so the FP8 checkpoint can be loaded:
```bash
pip install torch transformers compressed-tensors pillow
```
Usage is the same as for Clef:
```python
import sys
import torch
from huggingface_hub import snapshot_download
path = snapshot_download("prithivMLmods/clef-FP8")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model
model, processor = load_release_model(path, device="cuda")
record = {
"state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
"questions": {
"status": {
"type": "choice",
"instructions": "What is the invoice status?",
"criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."},
},
"large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
},
}
encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
logits = model(batch)[0]
for question, question_logits in zip(encoded.questions, logits):
probabilities = question_logits.float().softmax(-1).tolist()
print(question.question_id, dict(zip(question.option_ids, probabilities)))
```
The `systemone(model, processor, request)` helper and image/video inputs work as described in the
Clef card.
**Hardware:** FP8 W8A8 compute needs a GPU with FP8 support (Ada, Hopper, or newer). On older
GPUs, FP8 weights may only give memory savings, depending on the runtime.
## Reproducing the quantization
```python
from transformers import AutoModelForImageTextToText, AutoProcessor
from llmcompressor import oneshot
src = "Cloudflare/clef" # local path from snapshot_download works too
dst = "clef-FP8"
model = AutoModelForImageTextToText.from_pretrained(src, torch_dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained(src)
oneshot(model=model, recipe="recipe.yaml") # data-free, no dataset argument
model.save_pretrained(dst, save_compressed=True)
processor.save_pretrained(dst)
```
Then copy these files from the original Clef repo into `clef-FP8/` unchanged:
`joint_head.safetensors`, `joint_head_config.json`, `joint_schema_model.py`.
## Notes and limitations
- **Joint head:** the joint schema head is stored separately and is kept in its original
precision, because the recipe quantizes only the backbone. If you modify `load_release_model` or
re-export the model, make sure the head is not quantized.
- **Excluded modules:** `linear_attn`, the vision encoder, embeddings, and `lm_head` stay
unquantized, so the size reduction is somewhat smaller than a full 2x versus BF16.
- **Loading path:** this checkpoint is intended for the custom `joint_schema_model.py` loader.
General-purpose serving engines will not run the joint head.
## License
Apache-2.0, following [Cloudflare/clef](https://huggingface.co/Cloudflare/clef) and the base model
[Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B).