Laya Decision-Plugin β€” combined Core AI decision model (r38)

A typed decision classifier for coding agents: given an agent state and a typed question, it answers which tool, which skill, allow/ask/block, language, reply-or-act, triage, mail-sort, supervise, choose, compact, rerank β€” with a calibrated confidence. Zero token generation. One .aimodel asset (Apple Core AI, pure f16, static shape) ships a shared frozen encoder plus eleven LoRA chains selected per request by a trained router with fixed decision rails. Fully ANE-supported: the export contains zero fp32 islands, so the Apple Neural Engine compiler accepts the whole graph β€” it loads and runs cleanly pinned to neural_engine (NE p50 20 ms, matching GPU; ~2Γ— faster than the previous fp32-island export), with GPU and CPU as drop-in fallbacks via coreai-core β€” p50 20–340 ms per decision on Apple Silicon (macOS 27+), depending on question complexity.

This is an open-source alternative in the "decision-head agent plugin" category: it does not replace the main LLM β€” it is the fast, on-device decision layer in front of it (route, guard, escalate).

Architecture

Laya combined model architecture

Invariants, all mechanically checked: one encoder, N chains (chain switch = LoRA adapter swap, not a model load); closed-set router (unknown or foreign route_task strings clamp to the refusal-safe base chain); rails are code, not weights (the same thresholds the RL trainers optimized against); train == serve string (each family's binary answer renders in exactly the dialect it was trained on); temperatures ship per option-bucket, matching each RL head's optimization point (builder aborts on stale configs).

Scores β€” wire battery, r38 (plugin repo bench/reports/r38-final/)

Accuracy is measured through the deployed daemon over MCP (train==serve strings), on held-out oracles. "wire = head" means the serving path adds no loss versus offline evaluation.

# Use case Chain Wire accuracy (oracle, n) Wire p50/p95 ms Status
1 guardrail guard_command ft_gr14_ppo 0.315 exact-disposition, adversarial n=184; red-line holds 11/91 never-trained reds 34/38 ms shipped
2 lang lang_route ft_lang 0.829 (open-massive, n=2606; base 0.159) 36/39 ms shipped
3 skill skill_select ft_skill5_ppo 0.348 wire (n=161); stage-2 fit-judgment heads added: needs_skill 0.887, skill_needed 0.881 (dual-blind ho, n=801; previous chain was at chance there) 68/101 ms shipped
4 triage triage_message ft_triage5 0.575 wire = 0.625 head (n=40; teacher's own pack 0.700 caps end-to-end) 215/219 ms shipped
5 mail_sort mail_sort ft_mail8 0.625–0.708 wire (two harnesses, n=24) = 0.70 head champion (teacher ceiling 0.95) 35/37 ms shipped
6 supervise_run ft_sup3 0.679 wire (n=28; head 0.643, base 0.607) 178/182 ms shipped
7 rerank_noul ft_rrk3 0.100 exact = head (n=10; jaccard 0.765; base 0.000; teacher ceiling 8/10) 215/281 ms shipped
8 compact_context ft_cmp2 0.300 wire = head (n=10 exact-invariant golds; base 0.000) 324/340 ms shipped
9 choose_action ppo_cho2 0.800 wire = head (n=20; base 0.550, r37 chain 0.650) 21/22 ms shipped
10 ladder_plan / plan_request / ack_gate none by design code-only use cases (arithmetic, rails, local rule) β€” zero model calls 0 by design

Three further use cases (ladder planning, request planning, ack gate) are deliberately code, not weights β€” verified zero model calls.

Merit with vs without the model

Model level β€” base chain (no fine-tuning, same encoder) vs shipped chain, identical held-out golds:

Use case base chain shipped chain Ξ”
guardrail 0.194 (raw red 10/35 unseen) 0.465 routed; red-line holds 11/91 never-trained reds at floor 2.4Γ—
lang 0.159 0.829 wire (0.958 head) 5.2Γ—
tool_route (20-slate) 0.000 (acted 2/28, all wrong) 0.750 @ 80 % coverage (frozen oc2) ∞
skill_select (107-slate) 0.242 0.615 @ 100 % coverage stage-1 (dual-blind ho); stage-2 fit-judgment 0.887 / 0.881 added r38 2.5Γ—
triage 0.250 0.625 head / 0.575 wire 2.5Γ—
mail_sort 0.250 0.70 champion (wire 0.625–0.708) 2.8Γ—
supervise 0.607 0.679 wire (head 0.643) +12 %
choose 0.550 0.800 wire = head (ppo_cho2; r37 chain 0.650) +45 %
compact 0.000 0.300 wire = head ∞
rerank 0.000 0.100 exact / 0.765 jaccard ∞

Plugin level β€” same agent, same scenarios, rig without vs with the plugin (decision + guidance), correctness-gated:

model base (no plugin) guided (plugin as deterministic gate) warm (tool exposed) laya (tool+prompt)
q38 (75 GB MoE) 5/5 red, 2/2 gray 5/5, 2/2 β€” engine engaged 7/7, 54 ms mean 5/5, 2/2 3/5, 2/2 β€” engaged 1/12; the misses are attempted-then-blocked or environment-failed, not refusals
q36 (35B MoE) 3/5, 0/2 5/5, 2/2 β€” restores compliance, 58 ms mean 4/5, 1/2 4/5, 1/2 β€” engaged 1/12
gemma-4-26b 3/5, 1/2 5/5, 1/2 β€” engine 6 calls, 141 ms mean; one gray confirm was overridden by the model (macOS SIP contained it) 4/5, 1/2 3/5, 1/2 β€” engaged 0/12

Reading: without the plugin, two of three models execute never-negotiable red commands (q36 and gemma: 3/5 refused, and 0/2–1/2 gray held). With the plugin as a deterministic pre-command gate (guided arm), every model reaches 5/5 red, the decision engine is consulted on every guardrail- relevant command (~55–140 ms per call), and the engine itself blocked or escalated 4–5 of those 7 decisions. Gray compliance remains model-dependent even guided. Tool-exposed arms depend on the model choosing to call the tool: engagement is ≀1/12, which is the honest open problem (grace mode + read-only allowlist planned), not a claim.

What ships here

  • laya-combined-f16.aimodel/ β€” the single combined asset (B=1, L=1024 static, K=128), 11 chains, pure f16 with zero fp32 islands β€” fully ANE-supported (loads and runs on a neural_engine pin; the runtime auto-pads every call to L_max), sha-pinned per chain in combined_provenance.json.
  • run.py β€” one-file runner (prompt in, JSON verdict out); src/laya_port/ carries the torch-free runtime it imports.
  • configs/ β€” per-chain fitted deployment temperatures (option-bucketed; the PPO chains ship at the temperature their RL reward was optimized at), plus the tokenizer/ needed to build prompts.
  • combined_provenance.json β€” sha256s of pinned source + every chain, torch parity numbers, shapes.

Training corpora, per-round eval metrics, and the fine-tune ledger are NOT redistributed here β€” they live in the source repo and are reproduced by its Makefile (make model; see docs/REPRODUCE.md).

How to run

# pip install coreai-core transformers numpy  (no torch, no Xcode needed)
from laya_port.combined_agent import CombinedAgent
ag = CombinedAgent("laya-combined-f16.aimodel", "configs", unit="gpu")
d = ag.decide("guardrail", state="rm -rf /home/user/projects",
              question={"disposition": {"type": "choice", "instructions": "...",
                                        "criteria": {"allow": "...", "block": "..."}}})
# -> {'choice': 'block', 'confidence': 0.97, 'acted': True, ...}

Pin the compute unit (gpu default; ne runs the ANE β€” this asset is fully ANE-supported, both pin to a working specialization; unpinned loads can SIGABRT on ANE type-inference). One CombinedAgent per process; reuse it. For agent use (20+ use cases over MCP, auto-pull of this repo, one-line install) see the Laya Decision Plugin β€” github.com/Andrei-cloud/laya-plugin:

curl -fsSL https://raw.githubusercontent.com/Andrei-cloud/laya-plugin/master/install.sh | sh

Training in one paragraph

Head-only fine-tuning (encoder frozen — verified bit-identical across heads, which is what makes the combined asset legal), warm-start continuation, PPO over the acted decision (reward = wire behaviour, not teacher agreement), dual-blind teacher verification, session-disjoint splits with machine- re-asserted leak flags. 37 rounds; every VOID round and incident documented in the source repo's docs/FINETUNE.md — including a test-leak caught by the pipeline's own assertions, and the r36 class of train≠serve bugs (a dialect rewrite and a stale inherited temperature) that made wire numbers lie while the weights were fine. The r36 fixes are why every row above now reads wire = head.

Limitations (honest)

  • Golds are stronger-teacher agreement, not human consensus. Teacher self-agreement ceilings (0.38–0.95 per corpus) are measured and shipped.
  • Guardrail unseen-red band is thin (11/91 at floor on never-trained reds) β€” treat it as a confident gate with a regex advisory backstop, not as frictionless autonomy for destructive classes.
  • Skill-route real in-harness traffic is sparse; the verdict rests on a dual-blind-graded public holdout.
  • rerank/compact are early-loop chains: exact-match accuracy is low but each beats the base chain (0.000). Public-corpus expansion was attempted r38 and honestly reverted: the harvested faces taught a generic prior that contradicted the product register (sharegpt turns score "summarize" under the dual-blind rubric while the product keeps most turns), so the wire pack regressed despite winning offline shootouts. Face construction, not face volume, is the lever for these two chains.
  • Trained for the 20-tool / 107-slate coding-agent harness; foreign tool slates degrade gracefully to escalation, not to correct guesses.
  • Base checkpoint convaiinnovations/laya-multilingual is Apache-2.0; derivative weights here inherit Apache-2.0. Per-corpus licenses are in the source repo's eval/fetch_public_corpora.py headers.

Built on macOS 27.2 / Apple Silicon (M5 Max) with coreai-torch 0.4.2 + torch 2.13.0. Questions β†’ docs/ first; everything measurable is in there.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for AndyInQtr/laya-decision-plugin

Finetuned
(64)
this model