sev / README.md
minhleduc's picture
Rename model card to Sev
08c6ae0 verified
|
Raw History Blame Contribute Delete
8.36 kB
---
license: apache-2.0
base_model: convaiinnovations/laya
library_name: pytorch
pipeline_tag: text-classification
tags:
- calibration
- decision-making
- typed-decisions
- non-autoregressive
- modernbert
- proper-scoring-rules
datasets:
- LocalLLaMA/typed-decisions
language:
- en
metrics:
- accuracy
- brier_score
- nll
model-index:
- name: sev
results:
- task:
type: text-classification
name: Typed decisions
dataset:
type: LocalLLaMA/typed-decisions
name: LocalLLaMA/typed-decisions
split: test
metrics:
- type: accuracy
value: 0.7885
name: Accuracy
- type: brier_score
value: 0.0495
name: Brier (vs soft targets)
- type: nll
value: 0.8581
name: NLL (vs soft targets)
---
# Sev — typed-decisions, CE-only, 1024/256
A ModernBERT-large decision model fine-tuned on `LocalLLaMA/typed-decisions` with **plain cross-entropy** against the teacher's soft targets, at the documented 1024-token context / 256-token option budget.
It is the best-performing checkpoint from the **Sev** study (*Dissecting RLCD*), where I took apart the training method behind TypeSafe's Jev and its open reproduction, Laya. The short version of that study: Laya's RL term is a noise-smoothed cross-entropy gradient, and on this benchmark plain CE matches or beats it on every proper score.
**Test-set results (2,000 decisions, 400 cases):**
| Metric | Value |
|---|---|
| Accuracy (vs hard label) | **0.7885** |
| Brier (vs soft targets) | **0.0495** |
| NLL (vs soft targets) | **0.8581** |
| Soft accuracy | 0.5236 |
| Mean confidence | 0.6444 |
| Mean target max | 0.6589 |
| Score MAE (ordinal questions) | 0.2115 |
| Fitted temperature (choice / score / noul) | 1.116 / 1.070 / 1.120 |
For reference, the `laya-typed-decisions` checkpoint reports 0.766 accuracy and 0.062 Brier; TypeSafe's published Jev 1.13.0 figure is 0.727.
## Model description
A bidirectional encoder plus a small transformer head that reads one logit per masked option marker:
```
[CLS] <type> question: <instr> [SEP] [MASK] opt0 [MASK] opt1 ... [SEP] <state> [SEP]
|
ModernBERT-large encoder (bidirectional)
|
h += type_emb(qtype)
|
2 x TransformerEncoderLayer (pre-norm)
|
gather h at [MASK] positions -> scorer MLP -> 1 logit/option
```
One input row per typed question. Question types are `choice` (pick one of K named options), `score` (ordinal), and `noul` (binary yes/no). Option count is per-row, not fixed.
- Encoder: `answerdotai/ModernBERT-large` (395M)
- Head: 2 pre-norm transformer layers, dropout 0.1
- Scorer: LayerNorm, Linear, GELU, Linear(->1)
- Total: ~421M parameters
- Sequence budget: 1024 tokens total, 256 for the question + option block
The output is a probability distribution over the row's options. For the `score` type, the expected level is `sum i * p_i`.
## Intended use
Structured, closed-set probabilistic decisions over short English text: routing, triage, classification with calibrated confidence, and ordinal rating tasks where you control the option set.
**Out of scope:** open-ended generation (this model does not generate text), option sets larger than ~255, languages other than English (the base is English; a multilingual Laya variant exists), and any use where the option wording and the state distribution differ sharply from the training workflows.
## How to use
This is a custom architecture, not a `transformers` `AutoModel`. Load it with the code from the reproduction repo:
```bash
git clone https://github.com/LakoreAI/sev
cd rlcd-reverse-engineering && uv sync
```
```python
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
import sys
sys.path.insert(0, "rlcd-reverse-engineering") # or pip install -e it
from src.pipelines.infer import load_model, infer
d = snapshot_download("LakoreAI/sev")
model, cfg, tokenizer = load_model(Path(d) / "model.safetensors", torch.device("cuda"))
result = infer(
Path(d) / "model.safetensors",
state='{"task": "triage", "ticket": "customer cannot log in after password reset"}',
question='{"type": "choice", "instructions": "Which queue should this go to?",'
' "criteria": {"billing": "payment or invoice", "auth": "login or account access",'
' "network": "connectivity"}}',
)
print(result["predicted_key"], result["probabilities"])
```
`load_model` reads `model_config.json` to rebuild the architecture and `temperatures.json` to apply the fitted per-(type, K-bucket) temperature at inference, so the reported probabilities are already calibration-scaled.
## Training data
`LocalLLaMA/typed-decisions` (Apache-2.0): 1,200 train cases and 400 test cases across four workflows (agent-trace observability, customer service, invoice processing, security incidents), flattened to ~5,400 training decisions. Targets are the dataset's soft teacher distributions, not the hard labels.
## Training procedure
- Initialised from `convaiinnovations/laya`.
- Loss: soft cross-entropy against the teacher distributions only (no RL term).
- 4 epochs, effective batch 64 (micro-batch 8 x 4 accumulation on a single 24 GB GPU), AdamW, lr 2.5e-5 encoder / 1e-4 head, weight decay 0.01, cosine decay to 1e-6, gradient clipping 1.0, bf16.
- No early stopping; the final-epoch weights are evaluated.
- Post-hoc temperature fitted per (question type, option-count bucket) by NLL on a held-out 10% slice of the training cases.
## Evaluation
Full test-set breakdown by question type (2,000 decisions):
| Type | n | Accuracy | Brier | NLL | Mean conf. | Mean target max |
|---|---|---|---|---|---|---|
| choice | 600 | 0.7550 | 0.0611 | 0.9818 | 0.6194 | 0.6345 |
| score | 800 | 0.7588 | 0.0511 | 1.0331 | 0.5817 | 0.5929 |
| noul | 600 | 0.8617 | 0.0356 | 0.5009 | 0.7532 | 0.7714 |
| overall | 2000 | 0.7885 | 0.0495 | 0.8581 | 0.6444 | 0.6589 |
**A calibration caveat that matters here.** On this benchmark the gold hard label is the teacher's argmax on 98.4% of rows, so hard-label ECE is not a valid calibration measure: a model that reproduces the teacher perfectly still scores 0.326 hard-label ECE, and over-sharpening *lowers* it. The raw hard-label ECE of this checkpoint is 0.144 (0.162 after temperature). Judge calibration against the soft targets instead, with Brier and NLL above. The fitted temperatures (all > 1) are the model's own signal that the raw logits were slightly over-sharp before scaling.
## Relation to RLCD
This checkpoint is deliberately RL-free. In the accompanying study, adding Laya's RL term (a score-function estimate of a noise-smoothed proper score) did not improve accuracy and made the fitted temperature rise with the training noise scale. CE-only was at least as good on every proper score. See the linked repo for the full ablation, the estimator derivation, and the per-run result files.
## Limitations
- Single benchmark; the four workflows are synthetic and teacher-generated.
- English only.
- Option sets above ~20 options degrade, because options share a fixed token budget.
- Accuracy differences of ~1 point are within seed noise on a dataset this size; three seeds gave 0.782 +/- 0.004 for the 512-token variant.
- The base checkpoint was already fine-tuned by the Laya authors; this model is a further fine-tune on the same benchmark's train split.
## Links
- Code, configs, and per-run results: [LakoreAI/sev](https://github.com/LakoreAI/sev)
- All ablation checkpoints and metrics: [minhleduc/rlcd-e2-checkpoints](https://huggingface.co/minhleduc/rlcd-e2-checkpoints)
- Base model: [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya)
- Dataset: [LocalLLaMA/typed-decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions)
## Citation
```bibtex
@misc{leduc2026dissectingrlcd,
title = {Dissecting RLCD: What Reinforcement Learning Does (and Doesn't) Do for Calibrated Typed Decisions},
author = {Le Duc Minh},
year = {2026},
url = {https://github.com/LakoreAI/sev}
}
```