Clef 27B NF4

Quantized from Cloudflare/clef at 2f3de3dd85f379784083b0814d997ab627200f0c. The original model, schema encoder and joint head are by Cloudflare. This derived release packages NF4 with double quantization decoder weights with the original BF16 vision encoder, joint head and dense input/output vocabulary. It includes every inference asset; the original BF16 checkpoint is not required. The output vocabulary is in lm-head.safetensors and indexed separately.

Clef performs a single forward pass over a schema of choice, score and noul (boolean) questions. It does not generate free-form text. Probabilities are distributions over the supplied alternatives, not calibrated accuracy estimates.

Usage

Use Forge Neo's Clef setup and GUI, or this repository's standalone CLI:

python -m pip install torch==2.13.0 torchvision==0.28.0 --index-url https://download.pytorch.org/whl/cu130
python -m pip install -r requirements.txt
python inference.py --profile clef-24gb --image example.png
python inference.py --profile clef-24gb --image example.png --questions questions.json

questions.json is the original Clef schema: each question ID maps to a type, instructions, and, for choice/score, criteria. Use --state for context, --max-pixels for image processing (default 262144), and --max-length for the total input token cap (default 2048). Inputs exceeding the cap are refused. An NVIDIA CUDA GPU with BF16 support is required. The standalone loader reads weights locally and never downloads code during inference.

The 27B release is shared by clef-24gb and clef-16gb. The latter keeps the input embedding on CPU. Both keep the dense output vocabulary on CPU and move only requested rows to CUDA; decoder layers remain on CUDA. The 9B INT8 release uses flash-int8. Packed lm_head replacement and whole-layer CPU offload are not part of these profiles. The supplied loader is required for this placement.

Validation and limits

Verified on Windows, RTX 3090 24GB, RAM 64GB, with Transformers 5.10.2, bitsandbytes 0.50.2, safetensors 0.8.0 and PyTorch 2.13.0+cu130. We use SDPA, batch=1, no KV cache, and safetensors pread to avoid persistent Windows copy-on-write mappings of all shards.

Four images and a text-only request were checked with the same schemas. 27B 16GB and 24GB profiles produced identical answers and probabilities. Peak image allocations were 13.38 GiB and 15.73 GiB respectively. The 16GB profile passed with a 15 GiB PyTorch allocator budget on the 3090; an actual 16GB card was not available. The 9B INT8 peak was 9.82 GiB versus 16.26 GiB for BF16, with no choice changes in those four images and a maximum probability difference of 0.0423. This small comparison is not a general accuracy evaluation. INT8 was slower than BF16 on this test; its benefit here is lower memory use. Longer inputs or more image pixels require more memory.

The bundled loader uses image preprocessing exif-rgb-size-v2. It passes the per-call pixel bounds through images_kwargs.size and reports the actual processed image size and vision token count. Earlier releases passed the legacy max_pixels option, which Transformers 5.10.2 ignored. The allocations and probability comparisons above were measured with that earlier preprocessing, at the original test images' sizes, rather than a verified 262144-pixel cap. Do not compare old and new preprocessing as a quantization quality difference.

Provenance and license

Apache-2.0; see LICENSE and the upstream model card. The original joint_schema_model.py, head, vocabulary values and processor assets are retained. Decoder linear layers were converted with bitsandbytes and saved using Transformers. Vision and the judgment head remain BF16. The included clef_runtime and inference.py add CPU row gathering, profile selection, input validation and Windows pread loading. complete.json records the upstream commit, quantization settings, file sizes and SHA-256 hashes. No user images, evaluation histories, local configuration or credentials are included in this model release.

Downloads last month
19
Safetensors
Model size
26B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Aikimi/clef-nf4

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Quantized
(23)
this model