clef-flash-FP8

FP8 (W8A8, dynamic) quantization of Cloudflare/clef-flash, a 9B multimodal model that turns a state and a schema of typed questions into decisions. Clef-Flash reads text, JSON, images, or video and returns a probability for every allowed option of every question in a single forward pass, with no free-form generation and no output parsing. This repo quantizes only the backbone's linear layers. The vision encoder, embeddings, lm_head, and linear-attention layers are left in their original precision. For model behavior, input format, and the Jev/SystemOne API, see the original Clef-Flash card.

Quantization

Modality Image-Text-to-Text
Quantization scheme FP8_DYNAMIC (W8A8)
Weights FP8, per-channel
Activations FP8, per-token, dynamic
Calibration data Not required
Format compressed-tensors (safetensors)
Tooling LLM Compressor
License Apache 2.0
Setting Value
targets Linear
ignore lm_head, embed_tokens, visual, model.visual, linear_attn
scheme FP8_DYNAMIC
bypass_divisibility_checks false
requires_calibration_data false

recipe.yaml

default_stage:
  default_modifiers:
    QuantizationModifier:
      targets: [Linear]
      ignore: ['re:.*lm_head', 're:.*embed_tokens$', 're:.*visual.*', 're:.*model.visual.*',
        're:.*linear_attn.*']
      scheme: FP8_DYNAMIC
      bypass_divisibility_checks: false
      requires_calibration_data: false

Because the scheme is FP8_DYNAMIC, weight scales are computed directly from the weights and activation scales are computed per token at runtime. No calibration dataset is needed.

Usage

Install compressed-tensors alongside transformers so the FP8 checkpoint can be loaded:

pip install torch transformers compressed-tensors pillow

Usage is the same as for Clef-Flash:

import sys

import torch
from huggingface_hub import snapshot_download

path = snapshot_download("prithivMLmods/clef-flash-FP8")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model

model, processor = load_release_model(path, device="cuda")

record = {
    "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
    "questions": {
        "status": {
            "type": "choice",
            "instructions": "What is the invoice status?",
            "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."},
        },
        "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
    },
}

encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
    logits = model(batch)[0]

for question, question_logits in zip(encoded.questions, logits):
    probabilities = question_logits.float().softmax(-1).tolist()
    print(question.question_id, dict(zip(question.option_ids, probabilities)))

The systemone(model, processor, request) helper and image/video inputs work as described in the Clef-Flash card.

Hardware: FP8 W8A8 compute needs a GPU with FP8 support (Ada, Hopper, or newer). On older GPUs, FP8 weights may only give memory savings, depending on the runtime.

Reproducing the quantization

from transformers import AutoModelForImageTextToText, AutoProcessor
from llmcompressor import oneshot

src = "Cloudflare/clef-flash"   # local path from snapshot_download works too
dst = "clef-flash-FP8"

model = AutoModelForImageTextToText.from_pretrained(src, torch_dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained(src)

oneshot(model=model, recipe="recipe.yaml")   # data-free, no dataset argument

model.save_pretrained(dst, save_compressed=True)
processor.save_pretrained(dst)

Then copy these files from the original Clef-Flash repo into clef-flash-FP8/ unchanged: joint_head.safetensors, joint_head_config.json, joint_schema_model.py.

Notes and limitations

  • Joint head: the joint schema head is stored separately and is kept in its original precision, because the recipe quantizes only the backbone. If you modify load_release_model or re-export the model, make sure the head is not quantized.
  • Excluded modules: linear_attn, the vision encoder, embeddings, and lm_head stay unquantized, so the size reduction is somewhat smaller than a full 2x versus BF16.
  • Loading path: this checkpoint is intended for the custom joint_schema_model.py loader. General-purpose serving engines will not run the joint head.

License

Apache-2.0, following Cloudflare/clef-flash and the base model Qwen/Qwen3.5-9B.

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/clef-flash-FP8

Finetuned
Qwen/Qwen3.5-9B
Quantized
(8)
this model

Collections including prithivMLmods/clef-flash-FP8