Sev — typed-decisions, CE-only, 1024/256

A ModernBERT-large decision model fine-tuned on LocalLLaMA/typed-decisions with plain cross-entropy against the teacher's soft targets, at the documented 1024-token context / 256-token option budget.

It is the best-performing checkpoint from the Sev study (Dissecting RLCD), where I took apart the training method behind TypeSafe's Jev and its open reproduction, Laya. The short version of that study: Laya's RL term is a noise-smoothed cross-entropy gradient, and on this benchmark plain CE matches or beats it on every proper score.

Test-set results (2,000 decisions, 400 cases):

Metric Value
Accuracy (vs hard label) 0.7885
Brier (vs soft targets) 0.0495
NLL (vs soft targets) 0.8581
Soft accuracy 0.5236
Mean confidence 0.6444
Mean target max 0.6589
Score MAE (ordinal questions) 0.2115
Fitted temperature (choice / score / noul) 1.116 / 1.070 / 1.120

For reference, the laya-typed-decisions checkpoint reports 0.766 accuracy and 0.062 Brier; TypeSafe's published Jev 1.13.0 figure is 0.727.

Model description

A bidirectional encoder plus a small transformer head that reads one logit per masked option marker:

[CLS] <type> question: <instr> [SEP] [MASK] opt0 [MASK] opt1 ... [SEP] <state> [SEP]
                                                        |
                                  ModernBERT-large encoder (bidirectional)
                                                        |
                                  h += type_emb(qtype)
                                                        |
                                  2 x TransformerEncoderLayer (pre-norm)
                                                        |
                                  gather h at [MASK] positions -> scorer MLP -> 1 logit/option

One input row per typed question. Question types are choice (pick one of K named options), score (ordinal), and noul (binary yes/no). Option count is per-row, not fixed.

  • Encoder: answerdotai/ModernBERT-large (395M)
  • Head: 2 pre-norm transformer layers, dropout 0.1
  • Scorer: LayerNorm, Linear, GELU, Linear(->1)
  • Total: ~421M parameters
  • Sequence budget: 1024 tokens total, 256 for the question + option block

The output is a probability distribution over the row's options. For the score type, the expected level is sum i * p_i.

Intended use

Structured, closed-set probabilistic decisions over short English text: routing, triage, classification with calibrated confidence, and ordinal rating tasks where you control the option set.

Out of scope: open-ended generation (this model does not generate text), option sets larger than ~255, languages other than English (the base is English; a multilingual Laya variant exists), and any use where the option wording and the state distribution differ sharply from the training workflows.

How to use

This is a custom architecture, not a transformers AutoModel. Load it with the code from the reproduction repo:

git clone https://github.com/LakoreAI/sev
cd rlcd-reverse-engineering && uv sync
from pathlib import Path
import torch
from huggingface_hub import snapshot_download

import sys
sys.path.insert(0, "rlcd-reverse-engineering")  # or pip install -e it

from src.pipelines.infer import load_model, infer

d = snapshot_download("LakoreAI/sev")
model, cfg, tokenizer = load_model(Path(d) / "model.safetensors", torch.device("cuda"))

result = infer(
    Path(d) / "model.safetensors",
    state='{"task": "triage", "ticket": "customer cannot log in after password reset"}',
    question='{"type": "choice", "instructions": "Which queue should this go to?",'
             ' "criteria": {"billing": "payment or invoice", "auth": "login or account access",'
             ' "network": "connectivity"}}',
)
print(result["predicted_key"], result["probabilities"])

load_model reads model_config.json to rebuild the architecture and temperatures.json to apply the fitted per-(type, K-bucket) temperature at inference, so the reported probabilities are already calibration-scaled.

Training data

LocalLLaMA/typed-decisions (Apache-2.0): 1,200 train cases and 400 test cases across four workflows (agent-trace observability, customer service, invoice processing, security incidents), flattened to ~5,400 training decisions. Targets are the dataset's soft teacher distributions, not the hard labels.

Training procedure

  • Initialised from convaiinnovations/laya.
  • Loss: soft cross-entropy against the teacher distributions only (no RL term).
  • 4 epochs, effective batch 64 (micro-batch 8 x 4 accumulation on a single 24 GB GPU), AdamW, lr 2.5e-5 encoder / 1e-4 head, weight decay 0.01, cosine decay to 1e-6, gradient clipping 1.0, bf16.
  • No early stopping; the final-epoch weights are evaluated.
  • Post-hoc temperature fitted per (question type, option-count bucket) by NLL on a held-out 10% slice of the training cases.

Evaluation

Full test-set breakdown by question type (2,000 decisions):

Type n Accuracy Brier NLL Mean conf. Mean target max
choice 600 0.7550 0.0611 0.9818 0.6194 0.6345
score 800 0.7588 0.0511 1.0331 0.5817 0.5929
noul 600 0.8617 0.0356 0.5009 0.7532 0.7714
overall 2000 0.7885 0.0495 0.8581 0.6444 0.6589

A calibration caveat that matters here. On this benchmark the gold hard label is the teacher's argmax on 98.4% of rows, so hard-label ECE is not a valid calibration measure: a model that reproduces the teacher perfectly still scores 0.326 hard-label ECE, and over-sharpening lowers it. The raw hard-label ECE of this checkpoint is 0.144 (0.162 after temperature). Judge calibration against the soft targets instead, with Brier and NLL above. The fitted temperatures (all > 1) are the model's own signal that the raw logits were slightly over-sharp before scaling.

Relation to RLCD

This checkpoint is deliberately RL-free. In the accompanying study, adding Laya's RL term (a score-function estimate of a noise-smoothed proper score) did not improve accuracy and made the fitted temperature rise with the training noise scale. CE-only was at least as good on every proper score. See the linked repo for the full ablation, the estimator derivation, and the per-run result files.

Limitations

  • Single benchmark; the four workflows are synthetic and teacher-generated.
  • English only.
  • Option sets above ~20 options degrade, because options share a fixed token budget.
  • Accuracy differences of ~1 point are within seed noise on a dataset this size; three seeds gave 0.782 +/- 0.004 for the 512-token variant.
  • The base checkpoint was already fine-tuned by the Laya authors; this model is a further fine-tune on the same benchmark's train split.

Links

Citation

@misc{leduc2026dissectingrlcd,
  title  = {Dissecting RLCD: What Reinforcement Learning Does (and Doesn't) Do for Calibrated Typed Decisions},
  author = {Le Duc Minh},
  year   = {2026},
  url    = {https://github.com/LakoreAI/sev}
}
Downloads last month
12
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LakoreAI/sev

Finetuned
(84)
this model

Dataset used to train LakoreAI/sev

Evaluation results