laya / README.md
mertcobanov's picture
Ollaya package for convaiinnovations/laya
32fdc8c verified
|
Raw History Blame Contribute Delete
2.59 kB
---
license: apache-2.0
base_model:
- convaiinnovations/laya
library_name: onnx
tags:
- ollaya
- onnx
- decision-model
- system-one
pipeline_tag: text-classification
---
# laya for Ollaya
[Ollaya](https://github.com/ollaya-dev/ollaya) package of **[convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya)** by Convai Innovations.
Ollaya runs open decision models locally, the way Ollama runs LLMs: typed questions in,
calibrated answers out, behind a TypeSafe-compatible API.
```sh
ollaya run laya
```
## What is in this repository
This repository holds only the files Ollaya derives, with no weights. Each graph is an ONNX export of the
original model whose weights **reference the authors' own weight files by byte offset**,
so `ollaya pull` downloads the weights from the upstream repositories, unmodified and pinned to a
commit, and verifies their sha256.
| Tag | Upstream | Files |
|---|---|---|
| `laya:en` | [convaiinnovations/laya@aa8c91c](https://huggingface.co/convaiinnovations/laya/tree/aa8c91ca088ec597df95a0d1c76b3063cb2ae5e8) | `en/model-fp32.onnx`, `en/model-fp16.onnx`, `en/decision.json`, `en/calibration.json` |
| `laya:multilingual` | [convaiinnovations/laya@aa8c91c](https://huggingface.co/convaiinnovations/laya/tree/aa8c91ca088ec597df95a0d1c76b3063cb2ae5e8) | `multilingual/model-fp32.onnx`, `multilingual/model-fp16.onnx`, `multilingual/decision.json`, `multilingual/calibration.json` |
| `laya:typed-decisions` | [convaiinnovations/laya@aa8c91c](https://huggingface.co/convaiinnovations/laya/tree/aa8c91ca088ec597df95a0d1c76b3063cb2ae5e8) | `typed-decisions/model-fp32.onnx`, `typed-decisions/model-fp16.onnx`, `typed-decisions/decision.json`, `typed-decisions/calibration.json` |
Each tag has an fp32 graph (CPU) and an fp16 graph (GPU). Each tag also has `decision.json` (sequence layout, special tokens) and
`calibration.json` (temperatures).
## Parity
The exports are checked against the PyTorch reference on 2,383 questions per checkpoint.
The checks use typed-decisions plus multilingual and edge cases:
- **fp32:** the same decision on 100% of questions, with probabilities within 1.1e-4.
- **fp16:** the same decision on 99.1–99.6% of questions. Nearly all of the differences are
near-ties between the top two options.
The fp32 graphs (CPU) are opset 23: attention runs as fused `Attention` nodes, which ONNX Runtime's
CPU provider runs faster than the decomposed ops, so they need ONNX Runtime 1.23 or newer. The fp16
graphs (GPU) are opset 20.
## License
Same as the upstream model (Apache-2.0). Ollaya itself is Apache-2.0.