Laya Decision Router (bytesbrains/naderu-laya-150m, v0.2.0)

A lightweight, sub-50ms CPU/edge-friendly "System-1" decision and routing model released under the naderu.com specialized model fleet.

Fine-tuned from the answerdotai/ModernBERT-base foundation model and built on top of the upstream laya foundation framework (Apache-2.0), calibrated with post-training temperature scaling (holdout $ECE = 0.0356$ on templates never seen in training).

Model Details

  • Model ID: bytesbrains/naderu-laya-150m — version v0.2.0 (Hub tag v0.2.0)
  • Weights: bytesbrains/naderu-laya-150m on the Hugging Face Hub, published 2026-09-28. v0.1.0 was never published.
  • Foundation Model: answerdotai/ModernBERT-base (149M). v0.1.0 (never published) was named naderu/laya-decision-large-v0.1.0 but was also a ModernBERT-base fine-tune.
  • Adaptation: LoRA ($r=16$, $\alpha=32$ on Wqkv/Wo) merged into the released weights; see versions.md.
  • Upstream Framework: laya (Apache-2.0)
  • Architecture: ModernBertForSequenceClassification with calibrated Platt scaling
  • Primary Domain: Fast zero-overhead decision routing across customer operations, technical incident triage, sales qualification, and offboarding workflows, plus identity/provenance questions about the model itself.
  • License: Apache-2.0

Decision Space

The model classifies input sequences into 5 decision buckets — 4 operational, 1 identity. Boundary cases follow the labeling guide in training/datasets/laya_decisions/README.md.

Label Decision Category Target Triggers
account_billing Invoicing & Payments Charge disputes, renewal queries, VAT receipts, wire details, tier upgrade charges.
technical_support System & Infrastructure OOM errors, 5xx outages, database latency spikes, telemetry lag, API timeouts.
sales_inquiry Commercial & Enterprise Product demos, bespoke volume pricing, RFP vendor questionnaires, custom SLAs.
cancellation Offboarding & Churn Account termination, GDPR/CCPA data purge requests, auto-renew cancellations.
naderu_identity Identity & Provenance "Who made you?", "Is this Naderu by BytesBrains?", base-model and lineage questions.

naderu_identity fires on questions about the model. Operational queries that merely name Naderu or BytesBrains ("cancel our Naderu subscription") route to their operational label; the evaluation gate checks this. Breaking change from v0.1.0: logits went from 4 to 5 columns — update any hard-coded label map.


Quickstart

1. Standard Hugging Face Transformers

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model_id = "bytesbrains/naderu-laya-150m"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)

device = "mps" if torch.backends.mps.is_available() else ("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)

prompt = "System event: disk threshold warning above 92% capacity."
inputs = tokenizer(prompt, return_tensors="pt").to(device)

with torch.no_grad():
    logits = model(**inputs).logits
    decision = torch.argmax(logits, dim=-1).item()

print(f"Routing Decision: {model.config.id2label[decision]}")

2. Upstream Laya Pipeline

import laya

# Load the Laya agent or router
router = laya.LayaRouter(
    criteria={
        "account_billing": "Billing inquiries, invoice adjustments, and payment questions",
        "technical_support": "Technical defects, outage warnings, and infrastructure errors",
        "sales_inquiry": "Enterprise sales, demo requests, and contract quotes",
        "cancellation": "Subscription cancellations and account offboarding",
        "naderu_identity": "Questions about the model's identity, creator, and provenance",
    },
    model="bytesbrains/naderu-laya-150m",
)

result = router.invoke("System event: disk threshold warning above 92% capacity.")
print(f"Laya Routed to: {result}")

Performance & Evaluation Gate

Evaluated by eval/suites/laya_decision_v1/ on a template-disjoint holdout: 140 records from 32 independent units (templates or edge-case texts), none of which appears in training. Full output: eval/suites/laya_decision_v1/RESULTS.md.

Metric Target Gate Measured
Top-1 Accuracy (unseen templates) — (reported) 93.57% (131/140)
Macro-F1 Lift over Base $\ge +12.0%$ +83.31% (base 10.24% → 93.55%)
Expected Calibration Error (ECE) $< 0.06$ 0.0356 ($T = 1.1306$)
Apple Silicon MPS Latency (p95) $< 50$ ms 11.42 ms (M4 Mac mini)
CPU Latency (p95), reported $< 45$ ms SLA 15.84 ms PyTorch, 9.93 ms ONNX Runtime (M4, arm64)

Where it misses. The 9 errors come from 4 test templates, all near a class boundary: asking to email invoices to two addresses (billing → sales), a pro-forma invoice for renewal (billing → sales, 1 of 4 renders), "we're staying, but move us to the smaller plan" (billing → cancellation), and an RFP asking for pricing (sales → billing). Confidence on misses is 0.63–0.97, lower than on hits. With 32 test units, one template moves accuracy by ~3 points; treat 93.57% as roughly ±5 points. x86_64 CPU latency has not been measured.

How v0.2.0 got here: a template-disjoint split exposed that the earlier 20-template-per-class bank did not generalise (72.86% top-1, ECE 0.2151); the bank was tripled with a labeling guide and boundary cases, the splits were hash-recorded before training, and the recipe was trained and gated once. Every run, and the gate's identity-routing checks, are recorded in versions.md.


Edge Runbook & Quantization

ONNX Runtime Edge Serving

# Run ONNX export with dynamic batching
python training/recipes/laya_decision/export_onnx.py \
  --model-dir output/laya-decision-base-v0.2.0 \
  --output-onnx output/laya-decision-base-v0.2.0/model.onnx

Run edge inference using ONNX Runtime with CoreML / CPU providers:

import onnxruntime as ort
from transformers import AutoTokenizer
import numpy as np

session = ort.InferenceSession("model.onnx", providers=["CoreMLExecutionProvider", "CPUExecutionProvider"])
tokenizer = AutoTokenizer.from_pretrained("bytesbrains/naderu-laya-150m")

inputs = tokenizer("Payment failed on renewal invoice.", return_tensors="np")
outputs = session.run(None, {"input_ids": inputs["input_ids"], "attention_mask": inputs["attention_mask"]})
pred_idx = np.argmax(outputs[0], axis=-1)[0]

Apple CoreML Conversion (via coremltools)

import coremltools as ct
import torch
from transformers import AutoModelForSequenceClassification

model = AutoModelForSequenceClassification.from_pretrained("output/laya-decision-base-v0.2.0")
model.train(False)
example_input = torch.randint(0, 1000, (1, 128))
example_mask = torch.ones((1, 128), dtype=torch.int64)

traced = torch.jit.trace(model, (example_input, example_mask))
mlmodel = ct.convert(
    traced,
    inputs=[
        ct.TensorType(name="input_ids", shape=(1, ct.RangeDim(1, 512)), dtype=np.int64),
        ct.TensorType(name="attention_mask", shape=(1, ct.RangeDim(1, 512)), dtype=np.int64),
    ],
    compute_precision=ct.precision.FLOAT16,
)
mlmodel.save("LayaDecision.mlpackage")

Provenance & Attribution

  • Foundation Backbone: answerdotai/ModernBERT-base (Benjamin Warner, Antoine Chaffin, Benjamin Clavié, et al. Apache-2.0).
  • Upstream Framework: laya (Apache-2.0).
  • Fine-Tuning & Gate: Naderu, a venture of BytesBrains Pte. Ltd. (naderu.com).
Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bytesbrains/naderu-laya-150m

Finetuned
(1504)
this model