Laya Decision Router (bytesbrains/naderu-laya-150m, v0.2.0)
A lightweight, sub-50ms CPU/edge-friendly "System-1" decision and routing model released under the naderu.com specialized model fleet.
Fine-tuned from the answerdotai/ModernBERT-base foundation model and built on top of the upstream laya foundation framework (Apache-2.0), calibrated with post-training temperature scaling (holdout $ECE = 0.0356$ on templates never seen in training).
Model Details
- Model ID:
bytesbrains/naderu-laya-150m— version v0.2.0 (Hub tagv0.2.0) - Weights: bytesbrains/naderu-laya-150m on the Hugging Face Hub, published 2026-09-28. v0.1.0 was never published.
- Foundation Model:
answerdotai/ModernBERT-base(149M). v0.1.0 (never published) was namednaderu/laya-decision-large-v0.1.0but was also aModernBERT-basefine-tune. - Adaptation: LoRA ($r=16$, $\alpha=32$ on
Wqkv/Wo) merged into the released weights; seeversions.md. - Upstream Framework:
laya(Apache-2.0) - Architecture:
ModernBertForSequenceClassificationwith calibrated Platt scaling - Primary Domain: Fast zero-overhead decision routing across customer operations, technical incident triage, sales qualification, and offboarding workflows, plus identity/provenance questions about the model itself.
- License: Apache-2.0
Decision Space
The model classifies input sequences into 5 decision buckets — 4 operational, 1 identity. Boundary cases follow the
labeling guide in training/datasets/laya_decisions/README.md.
| Label | Decision Category | Target Triggers |
|---|---|---|
account_billing |
Invoicing & Payments | Charge disputes, renewal queries, VAT receipts, wire details, tier upgrade charges. |
technical_support |
System & Infrastructure | OOM errors, 5xx outages, database latency spikes, telemetry lag, API timeouts. |
sales_inquiry |
Commercial & Enterprise | Product demos, bespoke volume pricing, RFP vendor questionnaires, custom SLAs. |
cancellation |
Offboarding & Churn | Account termination, GDPR/CCPA data purge requests, auto-renew cancellations. |
naderu_identity |
Identity & Provenance | "Who made you?", "Is this Naderu by BytesBrains?", base-model and lineage questions. |
naderu_identity fires on questions about the model. Operational queries that merely name Naderu or
BytesBrains ("cancel our Naderu subscription") route to their operational label; the evaluation gate checks this.
Breaking change from v0.1.0: logits went from 4 to 5 columns — update any hard-coded label map.
Quickstart
1. Standard Hugging Face Transformers
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_id = "bytesbrains/naderu-laya-150m"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
device = "mps" if torch.backends.mps.is_available() else ("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)
prompt = "System event: disk threshold warning above 92% capacity."
inputs = tokenizer(prompt, return_tensors="pt").to(device)
with torch.no_grad():
logits = model(**inputs).logits
decision = torch.argmax(logits, dim=-1).item()
print(f"Routing Decision: {model.config.id2label[decision]}")
2. Upstream Laya Pipeline
import laya
# Load the Laya agent or router
router = laya.LayaRouter(
criteria={
"account_billing": "Billing inquiries, invoice adjustments, and payment questions",
"technical_support": "Technical defects, outage warnings, and infrastructure errors",
"sales_inquiry": "Enterprise sales, demo requests, and contract quotes",
"cancellation": "Subscription cancellations and account offboarding",
"naderu_identity": "Questions about the model's identity, creator, and provenance",
},
model="bytesbrains/naderu-laya-150m",
)
result = router.invoke("System event: disk threshold warning above 92% capacity.")
print(f"Laya Routed to: {result}")
Performance & Evaluation Gate
Evaluated by eval/suites/laya_decision_v1/ on a template-disjoint holdout: 140 records from 32 independent
units (templates or edge-case texts), none of which appears in training. Full output:
eval/suites/laya_decision_v1/RESULTS.md.
| Metric | Target Gate | Measured |
|---|---|---|
| Top-1 Accuracy (unseen templates) | — (reported) | 93.57% (131/140) |
| Macro-F1 Lift over Base | $\ge +12.0%$ | +83.31% (base 10.24% → 93.55%) |
| Expected Calibration Error (ECE) | $< 0.06$ | 0.0356 ($T = 1.1306$) |
| Apple Silicon MPS Latency (p95) | $< 50$ ms | 11.42 ms (M4 Mac mini) |
| CPU Latency (p95), reported | $< 45$ ms SLA | 15.84 ms PyTorch, 9.93 ms ONNX Runtime (M4, arm64) |
Where it misses. The 9 errors come from 4 test templates, all near a class boundary: asking to email invoices to two addresses (billing → sales), a pro-forma invoice for renewal (billing → sales, 1 of 4 renders), "we're staying, but move us to the smaller plan" (billing → cancellation), and an RFP asking for pricing (sales → billing). Confidence on misses is 0.63–0.97, lower than on hits. With 32 test units, one template moves accuracy by ~3 points; treat 93.57% as roughly ±5 points. x86_64 CPU latency has not been measured.
How v0.2.0 got here: a template-disjoint split exposed that the earlier 20-template-per-class bank did not
generalise (72.86% top-1, ECE 0.2151); the bank was tripled with a labeling guide and boundary cases, the splits
were hash-recorded before training, and the recipe was trained and gated once. Every run, and the gate's
identity-routing checks, are recorded in versions.md.
Edge Runbook & Quantization
ONNX Runtime Edge Serving
# Run ONNX export with dynamic batching
python training/recipes/laya_decision/export_onnx.py \
--model-dir output/laya-decision-base-v0.2.0 \
--output-onnx output/laya-decision-base-v0.2.0/model.onnx
Run edge inference using ONNX Runtime with CoreML / CPU providers:
import onnxruntime as ort
from transformers import AutoTokenizer
import numpy as np
session = ort.InferenceSession("model.onnx", providers=["CoreMLExecutionProvider", "CPUExecutionProvider"])
tokenizer = AutoTokenizer.from_pretrained("bytesbrains/naderu-laya-150m")
inputs = tokenizer("Payment failed on renewal invoice.", return_tensors="np")
outputs = session.run(None, {"input_ids": inputs["input_ids"], "attention_mask": inputs["attention_mask"]})
pred_idx = np.argmax(outputs[0], axis=-1)[0]
Apple CoreML Conversion (via coremltools)
import coremltools as ct
import torch
from transformers import AutoModelForSequenceClassification
model = AutoModelForSequenceClassification.from_pretrained("output/laya-decision-base-v0.2.0")
model.train(False)
example_input = torch.randint(0, 1000, (1, 128))
example_mask = torch.ones((1, 128), dtype=torch.int64)
traced = torch.jit.trace(model, (example_input, example_mask))
mlmodel = ct.convert(
traced,
inputs=[
ct.TensorType(name="input_ids", shape=(1, ct.RangeDim(1, 512)), dtype=np.int64),
ct.TensorType(name="attention_mask", shape=(1, ct.RangeDim(1, 512)), dtype=np.int64),
],
compute_precision=ct.precision.FLOAT16,
)
mlmodel.save("LayaDecision.mlpackage")
Provenance & Attribution
- Foundation Backbone:
answerdotai/ModernBERT-base(Benjamin Warner, Antoine Chaffin, Benjamin Clavié, et al. Apache-2.0). - Upstream Framework:
laya(Apache-2.0). - Fine-Tuning & Gate: Naderu, a venture of BytesBrains Pte. Ltd. (naderu.com).
- Downloads last month
- -
Model tree for bytesbrains/naderu-laya-150m
Base model
answerdotai/ModernBERT-base