Instructions to use desplega/laya-onnx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use desplega/laya-onnx with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
laya-onnx
fp32 ONNX exports of the three laya decision checkpoints, for use from Node or Bun with laya-js.
What laya is
laya is a non-autoregressive decision model by Convai Innovations (NandhaKishorM/laya). You give it a state (text or a JSON object) and typed questions, and it answers each one with probabilities in one encoder pass plus a small head:
choice: one label out of a fixed set, with a probability per label.noul: yes/no, as P(true).score: a level on an ordered scale, with a probability per level.
There is no text generation, so the model can only answer with the options you allowed.
The model design, training and weights are upstream's work, released under Apache-2.0. This repo only converts the published checkpoints to ONNX; the weights are not retrained or changed.
Checkpoints
| Folder | Exported from | Encoder | Context | Use it for |
|---|---|---|---|---|
english/fp32 |
convaiinnovations/laya @ 55cf4c4e |
ModernBERT-large | 512 tokens | English text |
multilingual/fp32 |
convaiinnovations/laya-multilingual @ e4e9ddf2 |
mmBERT-base | 1,024 tokens | Non-English or mixed text, or when latency matters (about 2x faster) |
typed-decisions/fp32 |
convaiinnovations/laya-typed-decisions @ 1a793eb5 |
ModernBERT-large, fine-tuned | 1,024 tokens | Upstream's four typed-decisions workflows: agent-trace observability, customer service, invoice processing, security incidents |
All three were exported from upstream laya v0.3.21 (commit 9d955671) at opset 18. Sizes: english and typed-decisions 1.69 GB each, multilingual 1.32 GB.
Files
Each checkpoint folder has the same layout:
<checkpoint>/fp32/
encoder.onnx # the encoder graph, weights inlined
head.onnx # the decision head graph
tokenizer.json # the upstream tokenizer
rl_agent_config.json # upstream config: encoder id, token budgets, temperatures
manifest.json # source repo and revision, SHA-256 and size of every file, tool versions
manifest.json ties each bundle to the exact upstream revision it came from, so you can verify a download file by file.
Use it from laya-js
laya-js is a typed TypeScript runtime for these bundles (@desplega.ai/laya) plus an HTTP server with upstream's POST /v1/systemone API (@desplega.ai/laya-server). Neither package is on npm yet, so install from a clone (Node 22+ or Bun 1.4+):
git clone https://github.com/desplega-ai/laya-js && cd laya-js
bun install
bun run build
# Download bundles from this repo into ./bundles, verified against the pinned SHA-256
node packages/laya-server/dist/fetch-models.js --dest ./bundles multilingual english
LAYA_MODEL_DIR=$PWD/bundles LAYA_THREADS=4 bun examples/01-support-triage.ts
laya-js pins a revision of this repo and the SHA-256 of every file in packages/laya/src/artifacts.ts, so a download that does not match fails instead of loading.
In code:
import { createAgent, defineQuestions } from "@desplega.ai/laya";
const agent = await createAgent({ checkpoint: "english", modelDir: "./bundles/english/fp32" });
const questions = defineQuestions({
team: { type: "choice", instructions: "Which team handles `message`?",
criteria: { billing: "charges, invoices", technical: "bugs, outages", sales: "pricing" } },
urgent: { type: "noul", instructions: "Is the customer blocked or on a deadline?" },
});
const r = await agent.predict({ message: "Production orders stopped syncing, we ship in two hours." }, questions);
r.answers.team.choice; // "billing" | "technical" | "sales"
r.answers.urgent.noul; // P(true), in [0, 1]
await agent.dispose();
See the usage guide for batching, long documents, zod-typed decide, routing between checkpoints, and the server.
Export them yourself
You do not need this repo to run laya-js. The source checkpoints are public, so tools/export rebuilds the same bundles with Python and uv, on CPU, with no token. It exports the split graphs from a pinned upstream revision and checks torch against ONNX to 1e-4:
cd tools/export
uv sync --frozen
CKPT=multilingual REPO=convaiinnovations/laya-multilingual REV=e4e9ddf21a7b1903b7acffd8814ad4307bf63a67 LEN=1024
uv run python export_split.py --repo $REPO --revision $REV --verify-len $LEN --out-dir ../../bundles/$CKPT/fp32
uv run python finalize_bundle.py --bundle-dir ../../bundles/$CKPT/fp32 --checkpoint $CKPT --repo $REPO --revision $REV
The export README lists the --repo, --revision and --verify-len values for all three checkpoints. Point LAYA_MODEL_DIR at the resulting bundles/ directory.
License and attribution
Apache License 2.0, the same license as the upstream checkpoints and code. laya is developed by Convai Innovations: github.com/NandhaKishorM/laya. The ONNX conversion and laya-js are by Desplega Labs.
Model tree for desplega/laya-onnx
Base model
convaiinnovations/laya