Download README.md from ollaya-dev/clef: direct link, hf CLI and curl.
- Browser
- Download file 1.79 kB
-
https://huggingface.co/ollaya-dev/clef/resolve/main/README.md
- Command line
-
hf download hf://ollaya-dev/clef/README.md
-
curl -L -o README.md https://huggingface.co/ollaya-dev/clef/resolve/main/README.md
license: apache-2.0
base_model:
- Cloudflare/clef-flash
library_name: onnx
tags:
- ollaya
- onnx
- decision-model
- system-one
pipeline_tag: text-classification
clef for Ollaya
Ollaya package of Cloudflare/clef-flash by Cloudflare (post-trained model and joint schema head) and the Qwen team (base model). Ollaya runs open decision models locally, the way Ollama runs LLMs: typed questions in, calibrated answers out, behind a TypeSafe-compatible API.
ollaya run clef
What is in this repository
This repository holds only the files Ollaya derives, with no weights. Each graph is an ONNX export of the
original model whose weights reference the authors' own weight files by byte offset,
so ollaya pull downloads the weights from the upstream repositories, unmodified and pinned to a
commit, and verifies their sha256.
| Tag | Upstream | Files |
|---|---|---|
clef:flash |
Cloudflare/clef-flash@17f0b0a | flash/model-fp32.onnx, flash/decision.json, flash/calibration.json |
Each tag has an fp32 graph, used on CPU and GPU. Each tag also has decision.json (sequence layout, special tokens) and
calibration.json (temperatures).
Parity
Ollaya's Rust runtime matches the authors' own code (joint_schema_model.py: their encoder, Qwen3.5 model and joint schema head, fp32) on 571 questions from 131 requests, on CUDA: identical token ids and spans, the same 13 rejected requests, the same decision on every question, logits within 4.3e-5 and probabilities within 6.3e-6.
License
Same as the upstream model (Apache-2.0). Ollaya itself is Apache-2.0.