--- license: apache-2.0 base_model: convaiinnovations/laya library_name: pytorch pipeline_tag: text-classification tags: - calibration - decision-making - typed-decisions - non-autoregressive - modernbert - proper-scoring-rules datasets: - LocalLLaMA/typed-decisions language: - en metrics: - accuracy - brier_score - nll model-index: - name: sev results: - task: type: text-classification name: Typed decisions dataset: type: LocalLLaMA/typed-decisions name: LocalLLaMA/typed-decisions split: test metrics: - type: accuracy value: 0.7885 name: Accuracy - type: brier_score value: 0.0495 name: Brier (vs soft targets) - type: nll value: 0.8581 name: NLL (vs soft targets) --- # Sev — typed-decisions, CE-only, 1024/256 A ModernBERT-large decision model fine-tuned on `LocalLLaMA/typed-decisions` with **plain cross-entropy** against the teacher's soft targets, at the documented 1024-token context / 256-token option budget. It is the best-performing checkpoint from the **Sev** study (*Dissecting RLCD*), where I took apart the training method behind TypeSafe's Jev and its open reproduction, Laya. The short version of that study: Laya's RL term is a noise-smoothed cross-entropy gradient, and on this benchmark plain CE matches or beats it on every proper score. **Test-set results (2,000 decisions, 400 cases):** | Metric | Value | |---|---| | Accuracy (vs hard label) | **0.7885** | | Brier (vs soft targets) | **0.0495** | | NLL (vs soft targets) | **0.8581** | | Soft accuracy | 0.5236 | | Mean confidence | 0.6444 | | Mean target max | 0.6589 | | Score MAE (ordinal questions) | 0.2115 | | Fitted temperature (choice / score / noul) | 1.116 / 1.070 / 1.120 | For reference, the `laya-typed-decisions` checkpoint reports 0.766 accuracy and 0.062 Brier; TypeSafe's published Jev 1.13.0 figure is 0.727. ## Model description A bidirectional encoder plus a small transformer head that reads one logit per masked option marker: ``` [CLS] question: [SEP] [MASK] opt0 [MASK] opt1 ... [SEP] [SEP] | ModernBERT-large encoder (bidirectional) | h += type_emb(qtype) | 2 x TransformerEncoderLayer (pre-norm) | gather h at [MASK] positions -> scorer MLP -> 1 logit/option ``` One input row per typed question. Question types are `choice` (pick one of K named options), `score` (ordinal), and `noul` (binary yes/no). Option count is per-row, not fixed. - Encoder: `answerdotai/ModernBERT-large` (395M) - Head: 2 pre-norm transformer layers, dropout 0.1 - Scorer: LayerNorm, Linear, GELU, Linear(->1) - Total: ~421M parameters - Sequence budget: 1024 tokens total, 256 for the question + option block The output is a probability distribution over the row's options. For the `score` type, the expected level is `sum i * p_i`. ## Intended use Structured, closed-set probabilistic decisions over short English text: routing, triage, classification with calibrated confidence, and ordinal rating tasks where you control the option set. **Out of scope:** open-ended generation (this model does not generate text), option sets larger than ~255, languages other than English (the base is English; a multilingual Laya variant exists), and any use where the option wording and the state distribution differ sharply from the training workflows. ## How to use This is a custom architecture, not a `transformers` `AutoModel`. Load it with the code from the reproduction repo: ```bash git clone https://github.com/LakoreAI/sev cd rlcd-reverse-engineering && uv sync ``` ```python from pathlib import Path import torch from huggingface_hub import snapshot_download import sys sys.path.insert(0, "rlcd-reverse-engineering") # or pip install -e it from src.pipelines.infer import load_model, infer d = snapshot_download("LakoreAI/sev") model, cfg, tokenizer = load_model(Path(d) / "model.safetensors", torch.device("cuda")) result = infer( Path(d) / "model.safetensors", state='{"task": "triage", "ticket": "customer cannot log in after password reset"}', question='{"type": "choice", "instructions": "Which queue should this go to?",' ' "criteria": {"billing": "payment or invoice", "auth": "login or account access",' ' "network": "connectivity"}}', ) print(result["predicted_key"], result["probabilities"]) ``` `load_model` reads `model_config.json` to rebuild the architecture and `temperatures.json` to apply the fitted per-(type, K-bucket) temperature at inference, so the reported probabilities are already calibration-scaled. ## Training data `LocalLLaMA/typed-decisions` (Apache-2.0): 1,200 train cases and 400 test cases across four workflows (agent-trace observability, customer service, invoice processing, security incidents), flattened to ~5,400 training decisions. Targets are the dataset's soft teacher distributions, not the hard labels. ## Training procedure - Initialised from `convaiinnovations/laya`. - Loss: soft cross-entropy against the teacher distributions only (no RL term). - 4 epochs, effective batch 64 (micro-batch 8 x 4 accumulation on a single 24 GB GPU), AdamW, lr 2.5e-5 encoder / 1e-4 head, weight decay 0.01, cosine decay to 1e-6, gradient clipping 1.0, bf16. - No early stopping; the final-epoch weights are evaluated. - Post-hoc temperature fitted per (question type, option-count bucket) by NLL on a held-out 10% slice of the training cases. ## Evaluation Full test-set breakdown by question type (2,000 decisions): | Type | n | Accuracy | Brier | NLL | Mean conf. | Mean target max | |---|---|---|---|---|---|---| | choice | 600 | 0.7550 | 0.0611 | 0.9818 | 0.6194 | 0.6345 | | score | 800 | 0.7588 | 0.0511 | 1.0331 | 0.5817 | 0.5929 | | noul | 600 | 0.8617 | 0.0356 | 0.5009 | 0.7532 | 0.7714 | | overall | 2000 | 0.7885 | 0.0495 | 0.8581 | 0.6444 | 0.6589 | **A calibration caveat that matters here.** On this benchmark the gold hard label is the teacher's argmax on 98.4% of rows, so hard-label ECE is not a valid calibration measure: a model that reproduces the teacher perfectly still scores 0.326 hard-label ECE, and over-sharpening *lowers* it. The raw hard-label ECE of this checkpoint is 0.144 (0.162 after temperature). Judge calibration against the soft targets instead, with Brier and NLL above. The fitted temperatures (all > 1) are the model's own signal that the raw logits were slightly over-sharp before scaling. ## Relation to RLCD This checkpoint is deliberately RL-free. In the accompanying study, adding Laya's RL term (a score-function estimate of a noise-smoothed proper score) did not improve accuracy and made the fitted temperature rise with the training noise scale. CE-only was at least as good on every proper score. See the linked repo for the full ablation, the estimator derivation, and the per-run result files. ## Limitations - Single benchmark; the four workflows are synthetic and teacher-generated. - English only. - Option sets above ~20 options degrade, because options share a fixed token budget. - Accuracy differences of ~1 point are within seed noise on a dataset this size; three seeds gave 0.782 +/- 0.004 for the 512-token variant. - The base checkpoint was already fine-tuned by the Laya authors; this model is a further fine-tune on the same benchmark's train split. ## Links - Code, configs, and per-run results: [LakoreAI/sev](https://github.com/LakoreAI/sev) - All ablation checkpoints and metrics: [minhleduc/rlcd-e2-checkpoints](https://huggingface.co/minhleduc/rlcd-e2-checkpoints) - Base model: [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya) - Dataset: [LocalLLaMA/typed-decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions) ## Citation ```bibtex @misc{leduc2026dissectingrlcd, title = {Dissecting RLCD: What Reinforcement Learning Does (and Doesn't) Do for Calibrated Typed Decisions}, author = {Le Duc Minh}, year = {2026}, url = {https://github.com/LakoreAI/sev} } ```