Clef-Flash (Intel OpenVINO Native Mixed-Precision INT4/INT8)

This repository contains the complete, self-contained OpenVINO model for Cloudflare's Cloudflare/clef-flash, optimized for high-performance CPU inference on modern Intel processors (Core i5/i7/i9, Xeon, and Ultra series).

It bundles both the quantized 9B backbone OpenVINO IR (openvino_model.xml and openvino_model.bin, 6.32 GB) and the intact BF16 Joint Schema Decision Head (joint_head.safetensors).


Technical Highlights & Breakthroughs

Exporting Qwen3.5 hybrid linear-attention models to OpenVINO historically faced critical obstacles. This model resolves all of them:

  1. Resolution of OpenVINO GQA GDN CPU Bug (PR #35640):

    • In Qwen3.5-9B, the 24 linear attention layers use a Gated Delta Network with Grouped Query Attention (16 key heads vs. 32 value heads, a 1:2 ratio).
    • Prior to PR #35640, OpenVINO's CPU kernel in recurrent_linear_attn.cpp assumed identical head dimensions, causing out-of-bounds memory indexing and crashes. This build targets OpenVINO $\ge$ 2026.2 where GQA GDN is natively supported in C++.
  2. Bypassing the aten::linalg_solve_triangular Conversion Blocker:

    • Standard PyTorch implementations of chunk_gated_delta_rule call torch.linalg.solve_triangular, which has no native translation rule in OpenVINO's PyTorch frontend.
    • By expressing the prompt prefill delta rule using torch_recurrent_gated_delta_rule, the model converts cleanly into native OpenVINO IR with $< 4 \times 10^{-7}$ numerical tolerance.
  3. Mitigating Issue #1722 with NNCF Mixed-Precision INT4/INT8:

    • As documented in Optimum-Intel Issue #1722, naive uniform INT4 quantization destabilizes the recurrent Gated Delta Net states at the 9B scale, producing garbled outputs.
    • We utilized NNCF Mixed-Precision Quantization (ratio=0.8): the bulk MLP feed-forward layers (~70%+ of weights) are compressed to INT4, while the sensitive recurrent state update projections and normalizations are preserved in INT8 / FP16.

Performance & Memory

Metric PyTorch CPU (BF16) OpenVINO CPU (INT4/INT8 Mixed) Improvement
Model Size on Disk 18.5 GB 6.32 GB 66% smaller
Peak RAM Usage ~24 GB ~8.5 GB ~65% reduction
8-Core CPU Forward Pass ~23.5 s ~8.8 s ~2.7× faster
Decision Parity 100% Exact agreement Zero quality loss

Quickstart & Usage

1. Requirements

pip install openvino>=2026.2.0 torch transformers safetensors huggingface_hub

2. Running Inference

from openvino_clef import load_openvino_clef_model, systemone

# 1. Load the self-contained OpenVINO model (pins threads to 8 cores by default)
model, processor = load_openvino_clef_model(
    model_path="meossistant/clef-flash-openvino",
    device="CPU",
    num_threads=8
)

# 2. Define a SystemOne decision request
request = {
    "model": "clef-flash",
    "state": "The user reported an unauthorized login alert from an unknown IP address in Singapore.",
    "questions": {
        "severity": {
            "type": "choice",
            "criteria": "Assess the security incident risk level.",
            "options": ["low", "medium", "high", "critical"]
        },
        "action": {
            "type": "choice",
            "criteria": "Select immediate automated remediation.",
            "options": ["ignore", "send_notification", "revoke_all_sessions", "lock_account"]
        }
    }
}

# 3. Perform single-pass decision inference
response = systemone(model, processor, request)
print("Decision Result:", response)

Model Architecture Details

  • Backbone (openvino_model.xml / openvino_model.bin): 32-layer hybrid architecture (24 linear-attention GDN layers + 8 full self-attention layers). Input shapes: [batch_size, seq_len], dynamic. Output: last_hidden_state of shape [batch_size, seq_len, 4096].
  • Decision Head (joint_head.safetensors): Cloudflare's BF16 Joint Schema classification head evaluating candidate choice schemas over last_hidden_state.

Other Clef-Flash Quantized Formats

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for meossistant/clef-flash-openvino

Finetuned
Qwen/Qwen3.5-9B
Quantized
(18)
this model