laya-onnx

ONNX export of the Laya non-autoregressive System 1 decision engine. Laya answers typed choice, score and yes/no questions over arbitrary text in a single forward pass β€” no token-by-token generation β€” and returns a full probability distribution rather than a sampled string.

This repository is a mirror, hosted so that the ElBruno.LocalLLMs.Decisions .NET package has a stable, owner-controlled default model source.

Provenance

Upstream model convaiinnovations/laya (PyTorch / safetensors)
ONNX export inferenceprince/laya-onnx
This repo Byte-identical mirror of that export
License Apache-2.0 (unchanged from upstream)

No weights were retrained, quantized or otherwise modified. Full credit for the model belongs to the Laya authors, and for the ONNX conversion to inferenceprince. Precision is fp16; weights live in the external-data file model.onnx.data.

Files

File Purpose
model.onnx Graph (weights are external)
model.onnx.data fp16 weights β€” required, the graph is unusable without it
config.json Encoder configuration
rl_agent_config.json Sequence budgets and fitted temperatures
tokenizer/tokenizer.json ByteLevel BPE vocabulary and merges
tokenizer/tokenizer_config.json Tokenizer settings

Graph contract

Inputs

Name Shape Type
input_ids [batch, seq] int64
attention_mask [batch, seq] int64
marker_pos [batch, k] int64
marker_mask [batch, k] bool
qtype [batch] int64 β€” choice=0, score=1, noul=2

Outputs

Name Shape Meaning
logits [batch, k] Per-option scores; masked slots are -1e4
act_logits [batch, 2] Escalation head

Batch, sequence and marker axes are all dynamic.

Each option is rendered as text preceded by a [MASK] marker, and a shared scorer scores every marker's hidden state independently. Nothing in the weights is indexed by label, so arbitrary option sets work against this static graph.

Usage from .NET

dotnet add package ElBruno.LocalLLMs.Decisions
var client = new LayaOnnxDecisionClient(new DecisionOptions
{
    ModelRepository = "elbruno/laya-onnx"
});

var result = await client.ChooseAsync(
    "My invoice charged me twice this month.",
    new[] { "billing", "technical", "sales" },
    "Route this support ticket");

Console.WriteLine(result.Choice);     // billing
Console.WriteLine(result.Confidence); // 0.9696

Runs fully in-process on ONNX Runtime β€” no Python, no server, no network at inference time. See the decisions guide.

Calibration caveat

rl_agent_config.json ships a fitted temperature of 0.1006 for the choice:11+ bucket (choice questions with 11 or more options), far below the 0.5 minimum Laya itself defines. That value sharpens the distribution roughly tenfold.

The .NET client clamps temperatures into Laya's own valid 0.5 – 5.0 range and reports it per answer via a CalibrationClamped flag.

Clamping bounds the damage but does not restore calibration. Any temperature below 1.0 still sharpens, so an 11+ option choice still saturates. Measured on this checkpoint with one support ticket:

Options Bucket Choice Reported confidence Clamped
5 choice:3-5 billing 86.9% no
12 choice:11+ billing 100.0% yes

Both answers are correct. For 11 or more options, use the ranking and ignore the magnitude, or fit your own temperature on labelled data. Consumers using this checkpoint through other runtimes should apply their own guard.

Verification

Reproduces the source export's published numbers exactly:

question: "Route this support ticket"  options: billing | technical | sales
raw logits : 4.2919, -3.0214, -3.0209
temperature: 1.760152
billing 0.9696 | technical 0.0152 | sales 0.0152

Note that fp16 accumulation makes batched results differ from single-question results by roughly 1e-5 to 1e-4, depending on the input and batch size. The ranking is unaffected β€” compare probabilities with a tolerance rather than ==.

Citation

@software{laya,
  title  = {Laya: Non-autoregressive System 1 decision engine},
  author = {Convai Innovations},
  url    = {https://github.com/NandhaKishorM/laya},
  license = {Apache-2.0}
}
Downloads last month
40
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support