Laya — ONNX (fp32, dynamic shapes)

An ONNX export of convaiinnovations/laya, the calibrated typed-decision model (ModernBERT-large encoder, 421M parameters). Laya answers choice, score and noul (calibrated yes/no) questions about a text or JSON state in one forward pass.

This export keeps batch, sequence length and option count dynamic. The existing community export bakes the sequence length into a Reshape inside ModernBERT's attention, so it only runs at 512 tokens. Here you can use it at any length up to 1,024 tokens (the export's bound; Laya itself uses 512), on any ONNX Runtime backend. It's also the starting point for running Laya on the Snapdragon X NPU: piffie/laya-snapdragon downloads it, pins a few fixed-shape buckets and compiles them for the Qualcomm HTP (17 ms per question vs 97 ms on the CPU, same answers).

Files

file what sha256
laya_fp32.onnx graph, opset 18 f6485a05ffd8ec1222fa7c78de8835dee2e0b9265e1436e1f81175e2af5e9b60
laya_fp32.onnx.data fp32 weights (external data, 1.69 GB) 487746363a8da57bcadb4345352997d22a0fb90d70aa22c6856668d023242aba
rl_agent_config.json temperatures for calibration, max_len, head_max_len (from the upstream checkpoint) ae287b56bbcf5f8c4f4541ae9dfd00c914c4c48b940b8398c3058af37ba92bbd
tokenizer/tokenizer.json, tokenizer/tokenizer_config.json upstream tokenizer, unchanged 6c8aaa9a…, 50044de6…

Inputs: input_ids [batch, seq] int64, attention_mask [batch, seq] int64, marker_pos [batch, markers] int64, marker_mask [batch, markers] bool, qtype [batch] int64 (0 choice, 1 score, 2 noul). Outputs: logits [batch, markers] (one per option, before temperature calibration) and act [batch, 2].

The raw logits are only half the model. You also need Laya's prompt layout (question, option markers, state) and its temperature calibration. Use them from laya-snapdragon (laya_snapdragon.Agent, which also runs this file on the CPU with device="cpu") or from laya-mlx. Both are torch-free and match upstream.

Provenance

  • Weights: convaiinnovations/laya at revision 1c5edc17a7acd8701df6fc341c0d179f1c62c982
  • Model code: NandhaKishorM/laya at 6a5819129eb220570792e417e49723d697efd76f
  • Export: torch.onnx.export(dynamo=True, dynamic_shapes=…, opset_version=18, optimize=True) with torch 2.14 (CPU) on native Windows ARM64. The script is laya_snapdragon/build.py (export) in the GitHub repo.
  • Nothing is quantized, pruned or retrained.

Fidelity

Against upstream PyTorch Laya (CPU), on laya-mlx's parity fixtures: 16 cases, 63 questions, 8 languages, up to 512 tokens and 20 options, with identical input tensors:

result
same argmax 63/63
max calibrated-probability difference 2.46e-6
max logit difference 2.86e-5
token ids from the torch-free prompt code identical, 16/16 cases
public result (rounded to 4 decimals, as upstream returns it) identical, 16/16 cases

License

Apache-2.0, like the upstream model and code. Laya and its weights are by Convai Innovations. This repo only re-packages them as ONNX.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for piffie/laya-onnx

Quantized
(63)
this model