File size: 2,124 Bytes
d467983
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
---
license: other
license_name: apache-2.0-and-gemma
license_link: https://ai.google.dev/gemma/terms
base_model:
- hiteshluke/arbiter-4b
- unsloth/gemma-3-4b-it
library_name: onnx
tags:
- ollaya
- onnx
- decision-model
- system-one
pipeline_tag: text-classification
---

# arbiter for Ollaya

[Ollaya](https://github.com/ollaya-dev/ollaya) package of **[hiteshluke/arbiter-4b](https://huggingface.co/hiteshluke/arbiter-4b)** and **[unsloth/gemma-3-4b-it](https://huggingface.co/unsloth/gemma-3-4b-it)** by Codekins Pvt Ltd · Zyot Lab.
Ollaya runs open decision models locally, the way Ollama runs LLMs: typed questions in,
calibrated answers out, behind a TypeSafe-compatible API.

```sh
ollaya run arbiter
```

## What is in this repository

This repository holds only the files Ollaya derives, with no weights. Each graph is an ONNX export of the
original model whose weights **reference the authors' own weight files by byte offset**,
so `ollaya pull` downloads the weights from the upstream repositories, unmodified and pinned to a
commit, and verifies their sha256.

| Tag | Upstream | Files |
|---|---|---|
| `arbiter:4b` | [hiteshluke/arbiter-4b@0c44271](https://huggingface.co/hiteshluke/arbiter-4b/tree/0c44271c59f89758e3cae17b032e98a9140093e9), [unsloth/gemma-3-4b-it@bf46152](https://huggingface.co/unsloth/gemma-3-4b-it/tree/bf46152c47f5dd20b896357cb51abc4c03b8ee8c) | `4b/model-fp32.onnx`, `4b/decision.json`, `4b/calibration.json` |

Each tag has an fp32 graph, used on CPU and GPU. Each tag also has `decision.json` (sequence layout, special tokens) and
`calibration.json` (temperatures).

## Parity

Ollaya's Rust runtime matches the reference (transformers' Gemma 3 with the authors' LoRA and head, fp32, the training script's prompts) on 420 questions from 127 requests, on CPU and CUDA (RTX 4090): identical token rows, the same 121 rejected requests, the same decision on every question, slot scores within 8.5e-5 and probabilities within 1.0e-5.

## License

Same as the upstream model (Apache-2.0 (LoRA adapter and head) and the Gemma Terms of Use (Gemma 3 base model)). Ollaya itself is Apache-2.0.