Clef-Flash (Intel OpenVINO Native Mixed-Precision INT4/INT8)
This repository contains the complete, self-contained OpenVINO model for Cloudflare's Cloudflare/clef-flash, optimized for high-performance CPU inference on modern Intel processors (Core i5/i7/i9, Xeon, and Ultra series).
It bundles both the quantized 9B backbone OpenVINO IR (openvino_model.xml and openvino_model.bin, 6.32 GB) and the intact BF16 Joint Schema Decision Head (joint_head.safetensors).
Technical Highlights & Breakthroughs
Exporting Qwen3.5 hybrid linear-attention models to OpenVINO historically faced critical obstacles. This model resolves all of them:
Resolution of OpenVINO GQA GDN CPU Bug (PR #35640):
- In Qwen3.5-9B, the 24 linear attention layers use a Gated Delta Network with Grouped Query Attention (16 key heads vs. 32 value heads, a 1:2 ratio).
- Prior to PR #35640, OpenVINO's CPU kernel in
recurrent_linear_attn.cppassumed identical head dimensions, causing out-of-bounds memory indexing and crashes. This build targets OpenVINO $\ge$ 2026.2 where GQA GDN is natively supported in C++.
Bypassing the
aten::linalg_solve_triangularConversion Blocker:- Standard PyTorch implementations of
chunk_gated_delta_rulecalltorch.linalg.solve_triangular, which has no native translation rule in OpenVINO's PyTorch frontend. - By expressing the prompt prefill delta rule using
torch_recurrent_gated_delta_rule, the model converts cleanly into native OpenVINO IR with $< 4 \times 10^{-7}$ numerical tolerance.
- Standard PyTorch implementations of
Mitigating Issue #1722 with NNCF Mixed-Precision INT4/INT8:
- As documented in Optimum-Intel Issue #1722, naive uniform INT4 quantization destabilizes the recurrent Gated Delta Net states at the 9B scale, producing garbled outputs.
- We utilized NNCF Mixed-Precision Quantization (
ratio=0.8): the bulk MLP feed-forward layers (~70%+ of weights) are compressed to INT4, while the sensitive recurrent state update projections and normalizations are preserved in INT8 / FP16.
Performance & Memory
| Metric | PyTorch CPU (BF16) | OpenVINO CPU (INT4/INT8 Mixed) | Improvement |
|---|---|---|---|
| Model Size on Disk | 18.5 GB | 6.32 GB | 66% smaller |
| Peak RAM Usage | ~24 GB | ~8.5 GB | ~65% reduction |
| 8-Core CPU Forward Pass | ~23.5 s | ~8.8 s | ~2.7× faster |
| Decision Parity | 100% | Exact agreement | Zero quality loss |
Quickstart & Usage
1. Requirements
pip install openvino>=2026.2.0 torch transformers safetensors huggingface_hub
2. Running Inference
from openvino_clef import load_openvino_clef_model, systemone
# 1. Load the self-contained OpenVINO model (pins threads to 8 cores by default)
model, processor = load_openvino_clef_model(
model_path="meossistant/clef-flash-openvino",
device="CPU",
num_threads=8
)
# 2. Define a SystemOne decision request
request = {
"model": "clef-flash",
"state": "The user reported an unauthorized login alert from an unknown IP address in Singapore.",
"questions": {
"severity": {
"type": "choice",
"criteria": "Assess the security incident risk level.",
"options": ["low", "medium", "high", "critical"]
},
"action": {
"type": "choice",
"criteria": "Select immediate automated remediation.",
"options": ["ignore", "send_notification", "revoke_all_sessions", "lock_account"]
}
}
}
# 3. Perform single-pass decision inference
response = systemone(model, processor, request)
print("Decision Result:", response)
Model Architecture Details
- Backbone (
openvino_model.xml/openvino_model.bin): 32-layer hybrid architecture (24 linear-attention GDN layers + 8 full self-attention layers). Input shapes:[batch_size, seq_len], dynamic. Output:last_hidden_stateof shape[batch_size, seq_len, 4096]. - Decision Head (
joint_head.safetensors): Cloudflare's BF16 Joint Schema classification head evaluating candidate choice schemas overlast_hidden_state.
Other Clef-Flash Quantized Formats
meossistant/clef-flash-4bit: 4-bit NF4 quantized PyTorch model (GPU-optimized withbitsandbytes, 133 ms on A100).meossistant/clef-flash-8bit: 8-bit Int8 quantized PyTorch model (GPU-optimized withbitsandbytes).