|
Download README.md from LakoreAI/sev: direct link, hf CLI and curl.
- Browser
- Download file 8.36 kB
-
https://huggingface.co/LakoreAI/sev/resolve/main/README.md
- Command line
-
hf download hf://LakoreAI/sev/README.md
-
curl -L -o README.md https://huggingface.co/LakoreAI/sev/resolve/main/README.md
8.36 kB
| license: apache-2.0 | |
| base_model: convaiinnovations/laya | |
| library_name: pytorch | |
| pipeline_tag: text-classification | |
| tags: | |
| - calibration | |
| - decision-making | |
| - typed-decisions | |
| - non-autoregressive | |
| - modernbert | |
| - proper-scoring-rules | |
| datasets: | |
| - LocalLLaMA/typed-decisions | |
| language: | |
| - en | |
| metrics: | |
| - accuracy | |
| - brier_score | |
| - nll | |
| model-index: | |
| - name: sev | |
| results: | |
| - task: | |
| type: text-classification | |
| name: Typed decisions | |
| dataset: | |
| type: LocalLLaMA/typed-decisions | |
| name: LocalLLaMA/typed-decisions | |
| split: test | |
| metrics: | |
| - type: accuracy | |
| value: 0.7885 | |
| name: Accuracy | |
| - type: brier_score | |
| value: 0.0495 | |
| name: Brier (vs soft targets) | |
| - type: nll | |
| value: 0.8581 | |
| name: NLL (vs soft targets) | |
| # Sev — typed-decisions, CE-only, 1024/256 | |
| A ModernBERT-large decision model fine-tuned on `LocalLLaMA/typed-decisions` with **plain cross-entropy** against the teacher's soft targets, at the documented 1024-token context / 256-token option budget. | |
| It is the best-performing checkpoint from the **Sev** study (*Dissecting RLCD*), where I took apart the training method behind TypeSafe's Jev and its open reproduction, Laya. The short version of that study: Laya's RL term is a noise-smoothed cross-entropy gradient, and on this benchmark plain CE matches or beats it on every proper score. | |
| **Test-set results (2,000 decisions, 400 cases):** | |
| | Metric | Value | | |
| |---|---| | |
| | Accuracy (vs hard label) | **0.7885** | | |
| | Brier (vs soft targets) | **0.0495** | | |
| | NLL (vs soft targets) | **0.8581** | | |
| | Soft accuracy | 0.5236 | | |
| | Mean confidence | 0.6444 | | |
| | Mean target max | 0.6589 | | |
| | Score MAE (ordinal questions) | 0.2115 | | |
| | Fitted temperature (choice / score / noul) | 1.116 / 1.070 / 1.120 | | |
| For reference, the `laya-typed-decisions` checkpoint reports 0.766 accuracy and 0.062 Brier; TypeSafe's published Jev 1.13.0 figure is 0.727. | |
| ## Model description | |
| A bidirectional encoder plus a small transformer head that reads one logit per masked option marker: | |
| ``` | |
| [CLS] <type> question: <instr> [SEP] [MASK] opt0 [MASK] opt1 ... [SEP] <state> [SEP] | |
| | | |
| ModernBERT-large encoder (bidirectional) | |
| | | |
| h += type_emb(qtype) | |
| | | |
| 2 x TransformerEncoderLayer (pre-norm) | |
| | | |
| gather h at [MASK] positions -> scorer MLP -> 1 logit/option | |
| ``` | |
| One input row per typed question. Question types are `choice` (pick one of K named options), `score` (ordinal), and `noul` (binary yes/no). Option count is per-row, not fixed. | |
| - Encoder: `answerdotai/ModernBERT-large` (395M) | |
| - Head: 2 pre-norm transformer layers, dropout 0.1 | |
| - Scorer: LayerNorm, Linear, GELU, Linear(->1) | |
| - Total: ~421M parameters | |
| - Sequence budget: 1024 tokens total, 256 for the question + option block | |
| The output is a probability distribution over the row's options. For the `score` type, the expected level is `sum i * p_i`. | |
| ## Intended use | |
| Structured, closed-set probabilistic decisions over short English text: routing, triage, classification with calibrated confidence, and ordinal rating tasks where you control the option set. | |
| **Out of scope:** open-ended generation (this model does not generate text), option sets larger than ~255, languages other than English (the base is English; a multilingual Laya variant exists), and any use where the option wording and the state distribution differ sharply from the training workflows. | |
| ## How to use | |
| This is a custom architecture, not a `transformers` `AutoModel`. Load it with the code from the reproduction repo: | |
| ```bash | |
| git clone https://github.com/LakoreAI/sev | |
| cd rlcd-reverse-engineering && uv sync | |
| ``` | |
| ```python | |
| from pathlib import Path | |
| import torch | |
| from huggingface_hub import snapshot_download | |
| import sys | |
| sys.path.insert(0, "rlcd-reverse-engineering") # or pip install -e it | |
| from src.pipelines.infer import load_model, infer | |
| d = snapshot_download("LakoreAI/sev") | |
| model, cfg, tokenizer = load_model(Path(d) / "model.safetensors", torch.device("cuda")) | |
| result = infer( | |
| Path(d) / "model.safetensors", | |
| state='{"task": "triage", "ticket": "customer cannot log in after password reset"}', | |
| question='{"type": "choice", "instructions": "Which queue should this go to?",' | |
| ' "criteria": {"billing": "payment or invoice", "auth": "login or account access",' | |
| ' "network": "connectivity"}}', | |
| ) | |
| print(result["predicted_key"], result["probabilities"]) | |
| ``` | |
| `load_model` reads `model_config.json` to rebuild the architecture and `temperatures.json` to apply the fitted per-(type, K-bucket) temperature at inference, so the reported probabilities are already calibration-scaled. | |
| ## Training data | |
| `LocalLLaMA/typed-decisions` (Apache-2.0): 1,200 train cases and 400 test cases across four workflows (agent-trace observability, customer service, invoice processing, security incidents), flattened to ~5,400 training decisions. Targets are the dataset's soft teacher distributions, not the hard labels. | |
| ## Training procedure | |
| - Initialised from `convaiinnovations/laya`. | |
| - Loss: soft cross-entropy against the teacher distributions only (no RL term). | |
| - 4 epochs, effective batch 64 (micro-batch 8 x 4 accumulation on a single 24 GB GPU), AdamW, lr 2.5e-5 encoder / 1e-4 head, weight decay 0.01, cosine decay to 1e-6, gradient clipping 1.0, bf16. | |
| - No early stopping; the final-epoch weights are evaluated. | |
| - Post-hoc temperature fitted per (question type, option-count bucket) by NLL on a held-out 10% slice of the training cases. | |
| ## Evaluation | |
| Full test-set breakdown by question type (2,000 decisions): | |
| | Type | n | Accuracy | Brier | NLL | Mean conf. | Mean target max | | |
| |---|---|---|---|---|---|---| | |
| | choice | 600 | 0.7550 | 0.0611 | 0.9818 | 0.6194 | 0.6345 | | |
| | score | 800 | 0.7588 | 0.0511 | 1.0331 | 0.5817 | 0.5929 | | |
| | noul | 600 | 0.8617 | 0.0356 | 0.5009 | 0.7532 | 0.7714 | | |
| | overall | 2000 | 0.7885 | 0.0495 | 0.8581 | 0.6444 | 0.6589 | | |
| **A calibration caveat that matters here.** On this benchmark the gold hard label is the teacher's argmax on 98.4% of rows, so hard-label ECE is not a valid calibration measure: a model that reproduces the teacher perfectly still scores 0.326 hard-label ECE, and over-sharpening *lowers* it. The raw hard-label ECE of this checkpoint is 0.144 (0.162 after temperature). Judge calibration against the soft targets instead, with Brier and NLL above. The fitted temperatures (all > 1) are the model's own signal that the raw logits were slightly over-sharp before scaling. | |
| ## Relation to RLCD | |
| This checkpoint is deliberately RL-free. In the accompanying study, adding Laya's RL term (a score-function estimate of a noise-smoothed proper score) did not improve accuracy and made the fitted temperature rise with the training noise scale. CE-only was at least as good on every proper score. See the linked repo for the full ablation, the estimator derivation, and the per-run result files. | |
| ## Limitations | |
| - Single benchmark; the four workflows are synthetic and teacher-generated. | |
| - English only. | |
| - Option sets above ~20 options degrade, because options share a fixed token budget. | |
| - Accuracy differences of ~1 point are within seed noise on a dataset this size; three seeds gave 0.782 +/- 0.004 for the 512-token variant. | |
| - The base checkpoint was already fine-tuned by the Laya authors; this model is a further fine-tune on the same benchmark's train split. | |
| ## Links | |
| - Code, configs, and per-run results: [LakoreAI/sev](https://github.com/LakoreAI/sev) | |
| - All ablation checkpoints and metrics: [minhleduc/rlcd-e2-checkpoints](https://huggingface.co/minhleduc/rlcd-e2-checkpoints) | |
| - Base model: [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya) | |
| - Dataset: [LocalLLaMA/typed-decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions) | |
| ## Citation | |
| ```bibtex | |
| @misc{leduc2026dissectingrlcd, | |
| title = {Dissecting RLCD: What Reinforcement Learning Does (and Doesn't) Do for Calibrated Typed Decisions}, | |
| author = {Le Duc Minh}, | |
| year = {2026}, | |
| url = {https://github.com/LakoreAI/sev} | |
| } | |
| ``` | |