Cloudflare Clef 27B EXL3

EXL3 conversion of Cloudflare/clef with a functional text/JSON Clef SystemOne adapter.

Source revision: 2f3de3dd85f379784083b0814d997ab627200f0c

Quantization

  • Qwen3.8 27B decoder: EXL3 4.00 bpw
  • Vision tower: EXL3 6 bpw
  • LM head: FP16 (head_bits=16)
  • Cloudflare joint schema head: original BF16 weights retained in joint_head.safetensors; loaded as FP16 by the adapter
  • Tested with ExLlamaV3 1.5.2+cu128.torch2.10.0 on an RTX 4090

The LM head is intentionally kept at FP16 because Clef uses its output embedding vectors when scoring schema options.

Validation

The bundled clef_exl3.py bridges ExLlamaV3 final hidden states into Cloudflare's original JointSchemaHead and returns the same noul, choice, and score answer structures used by SystemOne.

Check BF16 reference EXL3
Invoice status: overdue 0.9942 0.9944
Invoice total > $1000 0.9942 0.9945
Outage routing: technical 0.9161 0.9258
Urgency expected score 1.8176 1.8354
Service outage: true 0.8981 0.9078

Across all numeric values in the two bundled validation cases, mean absolute delta was 0.007055 and maximum absolute delta was 0.0178.

Discrete decisions matched the BF16 reference in both validation cases.

clef-exl3-smoke.json, clef-bf16-reference.json, and VALIDATION.json contain the validation outputs and provenance summary.

A standard generation probe on the RTX 4090 measured 26.337 tok/s. This is a loader/generation smoke benchmark, not SystemOne decision throughput.

Current scope

Text and JSON state inputs are validated.

The vision tower is included and quantized, but clef_exl3.py does not yet wire image/video inputs into the Clef decision path. Do not treat this release as validated multimodal SystemOne inference.

TabbyAPI compatibility is not validated in this release and is not used as a publish gate. Clef relies on its custom SystemOne decision path rather than a standard chat-completions path.

A preprocessor_config.json compatibility shim is included because this Cloudflare release stores image-processor metadata inside processor_config.json, while the tested ExLlamaV3 Qwen3.5/Qwen3.8 loader expects the standalone file.

Usage

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("ramgpt/clef-EXL3")
sys.path.insert(0, path)

from clef_exl3 import ClefEXL3

model = ClefEXL3(path)

response = model.systemone({
    "model": "clef-exl3",
    "state": {
        "invoice": {
            "vendor": "Acme",
            "total": 1250.0,
            "currency": "USD",
            "status": "overdue"
        }
    },
    "questions": {
        "status": {
            "type": "choice",
            "instructions": "What is the invoice status?",
            "criteria": {
                "paid": "Invoice is paid.",
                "overdue": "Invoice is past due.",
                "draft": "Not sent."
            }
        },
        "large": {
            "type": "noul",
            "instructions": "Is the total above 1000 USD?"
        }
    }
})

print(response["answers"])

Included validation artifacts

  • VALIDATION.json: quantization settings, runtime probe, BF16 parity summary, and numeric deltas
  • clef-exl3-smoke.json: EXL3 SystemOne outputs for choice/noul/score cases
  • clef-bf16-reference.json: official BF16 backbone + original joint-head outputs for the same cases
  • clef_exl3_smoke.py: reproducible EXL3 validation script
  • clef_exl3.py: EXL3-to-Clef joint-head adapter

Attribution

The base model, Clef joint schema head, and joint_schema_model.py originate from Cloudflare/clef and are provided under the source model's Apache-2.0 license.

This repository adds the EXL3 conversion, metadata compatibility shim, EXL3 adapter, and validation artifacts.

Downloads last month
25
Safetensors
Model size
9B params
Tensor type
BF16
路
F16
路
I16
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for ramgpt/clef-EXL3

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Quantized
(20)
this model