File size: 2,592 Bytes
d7629d5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
32fdc8c
 
d7629d5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
32fdc8c
 
 
 
d7629d5
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
---
license: apache-2.0
base_model:
- convaiinnovations/laya
library_name: onnx
tags:
- ollaya
- onnx
- decision-model
- system-one
pipeline_tag: text-classification
---

# laya for Ollaya

[Ollaya](https://github.com/ollaya-dev/ollaya) package of **[convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya)** by Convai Innovations.
Ollaya runs open decision models locally, the way Ollama runs LLMs: typed questions in,
calibrated answers out, behind a TypeSafe-compatible API.

```sh
ollaya run laya
```

## What is in this repository

This repository holds only the files Ollaya derives, with no weights. Each graph is an ONNX export of the
original model whose weights **reference the authors' own weight files by byte offset**,
so `ollaya pull` downloads the weights from the upstream repositories, unmodified and pinned to a
commit, and verifies their sha256.

| Tag | Upstream | Files |
|---|---|---|
| `laya:en` | [convaiinnovations/laya@aa8c91c](https://huggingface.co/convaiinnovations/laya/tree/aa8c91ca088ec597df95a0d1c76b3063cb2ae5e8) | `en/model-fp32.onnx`, `en/model-fp16.onnx`, `en/decision.json`, `en/calibration.json` |
| `laya:multilingual` | [convaiinnovations/laya@aa8c91c](https://huggingface.co/convaiinnovations/laya/tree/aa8c91ca088ec597df95a0d1c76b3063cb2ae5e8) | `multilingual/model-fp32.onnx`, `multilingual/model-fp16.onnx`, `multilingual/decision.json`, `multilingual/calibration.json` |
| `laya:typed-decisions` | [convaiinnovations/laya@aa8c91c](https://huggingface.co/convaiinnovations/laya/tree/aa8c91ca088ec597df95a0d1c76b3063cb2ae5e8) | `typed-decisions/model-fp32.onnx`, `typed-decisions/model-fp16.onnx`, `typed-decisions/decision.json`, `typed-decisions/calibration.json` |

Each tag has an fp32 graph (CPU) and an fp16 graph (GPU). Each tag also has `decision.json` (sequence layout, special tokens) and
`calibration.json` (temperatures).

## Parity

The exports are checked against the PyTorch reference on 2,383 questions per checkpoint.
The checks use typed-decisions plus multilingual and edge cases:

- **fp32:** the same decision on 100% of questions, with probabilities within 1.1e-4.
- **fp16:** the same decision on 99.1–99.6% of questions. Nearly all of the differences are
  near-ties between the top two options.

The fp32 graphs (CPU) are opset 23: attention runs as fused `Attention` nodes, which ONNX Runtime's
CPU provider runs faster than the decomposed ops, so they need ONNX Runtime 1.23 or newer. The fp16
graphs (GPU) are opset 20.

## License

Same as the upstream model (Apache-2.0). Ollaya itself is Apache-2.0.