|
Download README.md from ollaya-dev/laya: direct link, hf CLI and curl.
- Browser
- Download file 2.59 kB
-
https://huggingface.co/ollaya-dev/laya/resolve/main/README.md
- Command line
-
hf download hf://ollaya-dev/laya/README.md
-
curl -L -o README.md https://huggingface.co/ollaya-dev/laya/resolve/main/README.md
2.59 kB
| license: apache-2.0 | |
| base_model: | |
| - convaiinnovations/laya | |
| library_name: onnx | |
| tags: | |
| - ollaya | |
| - onnx | |
| - decision-model | |
| - system-one | |
| pipeline_tag: text-classification | |
| # laya for Ollaya | |
| [Ollaya](https://github.com/ollaya-dev/ollaya) package of **[convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya)** by Convai Innovations. | |
| Ollaya runs open decision models locally, the way Ollama runs LLMs: typed questions in, | |
| calibrated answers out, behind a TypeSafe-compatible API. | |
| ```sh | |
| ollaya run laya | |
| ``` | |
| ## What is in this repository | |
| This repository holds only the files Ollaya derives, with no weights. Each graph is an ONNX export of the | |
| original model whose weights **reference the authors' own weight files by byte offset**, | |
| so `ollaya pull` downloads the weights from the upstream repositories, unmodified and pinned to a | |
| commit, and verifies their sha256. | |
| | Tag | Upstream | Files | | |
| |---|---|---| | |
| | `laya:en` | [convaiinnovations/laya@aa8c91c](https://huggingface.co/convaiinnovations/laya/tree/aa8c91ca088ec597df95a0d1c76b3063cb2ae5e8) | `en/model-fp32.onnx`, `en/model-fp16.onnx`, `en/decision.json`, `en/calibration.json` | | |
| | `laya:multilingual` | [convaiinnovations/laya@aa8c91c](https://huggingface.co/convaiinnovations/laya/tree/aa8c91ca088ec597df95a0d1c76b3063cb2ae5e8) | `multilingual/model-fp32.onnx`, `multilingual/model-fp16.onnx`, `multilingual/decision.json`, `multilingual/calibration.json` | | |
| | `laya:typed-decisions` | [convaiinnovations/laya@aa8c91c](https://huggingface.co/convaiinnovations/laya/tree/aa8c91ca088ec597df95a0d1c76b3063cb2ae5e8) | `typed-decisions/model-fp32.onnx`, `typed-decisions/model-fp16.onnx`, `typed-decisions/decision.json`, `typed-decisions/calibration.json` | | |
| Each tag has an fp32 graph (CPU) and an fp16 graph (GPU). Each tag also has `decision.json` (sequence layout, special tokens) and | |
| `calibration.json` (temperatures). | |
| ## Parity | |
| The exports are checked against the PyTorch reference on 2,383 questions per checkpoint. | |
| The checks use typed-decisions plus multilingual and edge cases: | |
| - **fp32:** the same decision on 100% of questions, with probabilities within 1.1e-4. | |
| - **fp16:** the same decision on 99.1–99.6% of questions. Nearly all of the differences are | |
| near-ties between the top two options. | |
| The fp32 graphs (CPU) are opset 23: attention runs as fused `Attention` nodes, which ONNX Runtime's | |
| CPU provider runs faster than the decomposed ops, so they need ONNX Runtime 1.23 or newer. The fp16 | |
| graphs (GPU) are opset 20. | |
| ## License | |
| Same as the upstream model (Apache-2.0). Ollaya itself is Apache-2.0. | |