|
Download README.md from ollaya-dev/clm: direct link, hf CLI and curl.
- Browser
- Download file 1.98 kB
-
https://huggingface.co/ollaya-dev/clm/resolve/main/README.md
- Command line
-
hf download hf://ollaya-dev/clm/README.md
-
curl -L -o README.md https://huggingface.co/ollaya-dev/clm/resolve/main/README.md
1.98 kB
| license: apache-2.0 | |
| base_model: | |
| - Contrastive-LM/CLM-v0.1-8B | |
| - Qwen/Qwen3-8B | |
| library_name: onnx | |
| tags: | |
| - ollaya | |
| - onnx | |
| - decision-model | |
| - system-one | |
| pipeline_tag: text-classification | |
| # clm for Ollaya | |
| [Ollaya](https://github.com/ollaya-dev/ollaya) package of **[Contrastive-LM/CLM-v0.1-8B](https://huggingface.co/Contrastive-LM/CLM-v0.1-8B)** and **[Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)** by Contrastive-LM (projection heads) and the Qwen team (Qwen3-8B encoder). | |
| Ollaya runs open decision models locally, the way Ollama runs LLMs: typed questions in, | |
| calibrated answers out, behind a TypeSafe-compatible API. | |
| ```sh | |
| ollaya run clm | |
| ``` | |
| ## What is in this repository | |
| This repository holds only the files Ollaya derives, with no weights. Each graph is an ONNX export of the | |
| original model whose weights **reference the authors' own weight files by byte offset**, | |
| so `ollaya pull` downloads the weights from the upstream repositories, unmodified and pinned to a | |
| commit, and verifies their sha256. | |
| | Tag | Upstream | Files | | |
| |---|---|---| | |
| | `clm:8b` | [Contrastive-LM/CLM-v0.1-8B@e939398](https://huggingface.co/Contrastive-LM/CLM-v0.1-8B/tree/e939398d4556fcd9400c76fa8c5a513202f42b0a), [Qwen/Qwen3-8B@b968826](https://huggingface.co/Qwen/Qwen3-8B/tree/b968826d9c46dd6066d109eabc6255188de91218) | `8b/model-fp32.onnx`, `8b/decision.json`, `8b/calibration.json` | | |
| Each tag has an fp32 graph, used on CPU and GPU. Each tag also has `decision.json` (sequence layout, special tokens) and | |
| `calibration.json` (temperatures). | |
| ## Parity | |
| Ollaya's Rust runtime matches the reference (upstream clm.schema and heads, the Qwen3-8B encoder in fp32 on its BF16 weights) on 480 questions from 117 requests, on CPU and CUDA: the same state and option texts and token ids, the same rejections, the same decision on every question, option logits within 1.3e-4 and probabilities within 2.9e-5. | |
| ## License | |
| Same as the upstream model (Apache-2.0). Ollaya itself is Apache-2.0. | |