Instructions to use piffie/laya-onnx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use piffie/laya-onnx with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Laya — ONNX (fp32, dynamic shapes)
An ONNX export of convaiinnovations/laya, the calibrated
typed-decision model (ModernBERT-large encoder, 421M parameters). Laya answers choice, score and
noul (calibrated yes/no) questions about a text or JSON state in one forward pass.
This export keeps batch, sequence length and option count dynamic. The existing community export bakes the sequence length into a Reshape inside ModernBERT's attention, so it only runs at 512 tokens. Here you can use it at any length up to 1,024 tokens (the export's bound; Laya itself uses 512), on any ONNX Runtime backend. It's also the starting point for running Laya on the Snapdragon X NPU: piffie/laya-snapdragon downloads it, pins a few fixed-shape buckets and compiles them for the Qualcomm HTP (17 ms per question vs 97 ms on the CPU, same answers).
Files
| file | what | sha256 |
|---|---|---|
laya_fp32.onnx |
graph, opset 18 | f6485a05ffd8ec1222fa7c78de8835dee2e0b9265e1436e1f81175e2af5e9b60 |
laya_fp32.onnx.data |
fp32 weights (external data, 1.69 GB) | 487746363a8da57bcadb4345352997d22a0fb90d70aa22c6856668d023242aba |
rl_agent_config.json |
temperatures for calibration, max_len, head_max_len (from the upstream checkpoint) |
ae287b56bbcf5f8c4f4541ae9dfd00c914c4c48b940b8398c3058af37ba92bbd |
tokenizer/tokenizer.json, tokenizer/tokenizer_config.json |
upstream tokenizer, unchanged | 6c8aaa9a…, 50044de6… |
Inputs: input_ids [batch, seq] int64, attention_mask [batch, seq] int64, marker_pos [batch, markers] int64,
marker_mask [batch, markers] bool, qtype [batch] int64 (0 choice, 1 score, 2 noul).
Outputs: logits [batch, markers] (one per option, before temperature calibration) and act [batch, 2].
The raw logits are only half the model. You also need Laya's prompt layout (question, option markers,
state) and its temperature calibration. Use them from laya-snapdragon
(laya_snapdragon.Agent, which also runs this file on the CPU with device="cpu") or from
laya-mlx. Both are torch-free and match upstream.
Provenance
- Weights:
convaiinnovations/layaat revision1c5edc17a7acd8701df6fc341c0d179f1c62c982 - Model code: NandhaKishorM/laya at
6a5819129eb220570792e417e49723d697efd76f - Export:
torch.onnx.export(dynamo=True, dynamic_shapes=…, opset_version=18, optimize=True)with torch 2.14 (CPU) on native Windows ARM64. The script islaya_snapdragon/build.py(export) in the GitHub repo. - Nothing is quantized, pruned or retrained.
Fidelity
Against upstream PyTorch Laya (CPU), on laya-mlx's parity fixtures: 16 cases, 63 questions, 8 languages, up to 512 tokens and 20 options, with identical input tensors:
| result | |
|---|---|
| same argmax | 63/63 |
| max calibrated-probability difference | 2.46e-6 |
| max logit difference | 2.86e-5 |
| token ids from the torch-free prompt code | identical, 16/16 cases |
| public result (rounded to 4 decimals, as upstream returns it) | identical, 16/16 cases |
License
Apache-2.0, like the upstream model and code. Laya and its weights are by Convai Innovations. This repo only re-packages them as ONNX.
Model tree for piffie/laya-onnx
Base model
convaiinnovations/laya