laya-typed-decisions-multilingual

A 322M "System One" decision head: hand it an agent trace (the state) and typed questions, and it returns calibrated decisions in millisecond-class time on CPU — what to do with the trace, whether a human should review it, how the task ended. Built by modusensus as a community fine-tune combining what #320 asked for: the mmBERT-base multilingual encoder with decision heads trained on the typed-decisions workflows, using the official public recipe — the Kaggle 2xT4 notebook, unmodified.

The three question types

type answers returns
choice pick one option from criteria argmax label + full probability vector
noul yes/no — does the statement hold noul = P(yes), calibrated
score rate on an ordered criteria scale expected score + per-level probabilities

Results

Official test split, 400 cases / 2,000 decisions, argmax vs gold label:

model choice noul score total
laya-multilingual (zero-shot) 0.295 0.497 0.286 0.352
this checkpoint 0.752 0.862 0.759 0.7875
laya-typed-decisions (published, English encoder) — — — 0.766
meraGPT Decider 1 (closed, zero-shot, leaderboard) — — — 0.768
TypeSafe Jev 1.13.0 (closed, measured by the dataset card) — — — 0.727

All five rows are the same benchmark (all/test, argmax vs gold), so they are directly comparable. Numbers quoted from other cards' own harnesses or from vendor sites (including Jev's own site claims) use different rulers and are deliberately omitted — see the round-4 task book discipline (kaggle_eval/HANDOFF_NLI_V4.1.md, 存档 A/C).

ECE as shipped (max-prob confidence, 10 bins): 0.159 — same split and same binning as the official laya-typed-decisions 0.213, so those two ECE numbers are comparable; other cards' ECE figures come from different harnesses and are not. Temperatures were fitted on the notebook's held-out-from-training calibration slice (choice 1.11, score 1.05, noul 1.20); that slice is still in-distribution for the benchmark, so treat the calibration number as optimistic.

Option-name robustness (round-4 diagnostic). Renaming the criteria keys (option names) while keeping the rubric texts and order fixed — neutral A/B/… or per-question random strings, the axis arXiv:2609.26758 shows typed decision heads latch onto — flips 11.8% (neutral, McNemar p=0.0007) / 17.5% (random, p=0.014) of the 600 choice decisions and costs ~4.5pp accuracy (0.752 → 0.707 / 0.710): probability mass migrates toward the still-salient semantic names (stop, success). The head leans on the option-name words, not only the rubrics bound to them; name-system augmentation is queued for the next training round. The baseline for this check reproduces the numbers above (0.7880 total, GPU). Reproduce: kaggle_eval/typed_polarity_kernel.py (Kaggle GPU) or typed_polarity_check.py (CPU).

Full write-up and early non-English (zh/es/ja) spot-check observations: laya Discussions #482. Short version: choice routing survives fine-tuning in all three languages tested; the score head's non-English calibration moved (mostly directional improvements, one regression in ja). Per-language held-out evaluation is still pending — no multilingual accuracy claims are made here.

Usage

Run verbatim (output below is from this exact snippet):

import laya  # pip install laya

agent = laya.load("Modusnsus/laya-typed-decisions-multilingual")
state = {"task": "Rotate the expired TLS certificate on the staging load balancer.",
         "constraints": ["Do not exceed a $50 spend on cloud resources"],
         "trace_summary": {"steps": 11, "duration_s": 32.5, "tool_errors": 0,
                            "constraint_violations": 0, "irreversible_actions": 0}}
questions = {"action": {"type": "choice",
                        "instructions": "What should the observability system do with this trace?",
                        "criteria": {"continue": "Let the agent proceed without interruption.",
                                     "human_review": "Queue this trace for a human to review.",
                                     "observe": "Keep running, but flag the trace for later sampling.",
                                     "stop": "Halt the agent now."}},
             "needs_review": {"type": "noul",
                              "instructions": "This trace requires human review."}}
print(agent.predict(state, questions)["answers"])

A clean trace sits near the continue/observe boundary, and review is confidently declined:

{"action": {"type": "choice", "choice": "continue",
            "probabilities": {"continue": 0.4461, "human_review": 0.1126,
                              "observe": 0.4196, "stop": 0.0217},
            "confidence": 0.24, "answer_confidence": 0.4461},
 "needs_review": {"type": "noul", "noul": 0.1119, "confidence": 0.8881}}

answer_confidence is the calibrated per-answer confidence (one number you can gate on across all three types); confidence keeps the per-type legacy meaning.

Training

  • Base: convaiinnovations/laya-multilingual (mmBERT-base, 322M)
  • Data: LocalLLaMA/typed-decisions train split (~30k items), full-encoder RLCD (proper-scoring-rule reward + noisy-logit policy gradient + soft CE), 4 epochs
  • Hardware: Kaggle 2x T4 (DDP), ~1.5 h wall clock
  • Post-training: per-type temperature calibration, temperature_by_options removed from the config

If this is your first small-model fine-tune: the point of this repo is that the whole thing — 30k rows, free 2×T4, unmodified public notebook — reproduces end to end, and the result beats the published English-encoder checkpoint on the official split.

Sibling head

Modusnsus/laya-nli-memory-conflict — a memory-conflict noul head (v4) trained to detect when new information contradicts stored memory; the two models share the Laya base and API but answer different questions.

Reproduce

Run notebooks/laya_finetune_typed_decisions_2xT4_kaggle.ipynb with convaiinnovations/laya-multilingual as the base checkpoint, then evaluate on the all/test split of LocalLLaMA/typed-decisions.

model.safetensors SHA256: 8bb8cdd4875313baae2cdea52a4f682a2fbc86fd174c9973165b6b04a1aa2718

License

Apache 2.0, matching the base checkpoint. Original model by Convai Innovations.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Modusnsus/laya-typed-decisions-multilingual

Finetuned
(55)
this model

Dataset used to train Modusnsus/laya-typed-decisions-multilingual

Paper for Modusnsus/laya-typed-decisions-multilingual

Evaluation results