size-decision-general
A calibrated multiple-choice decision model with an mmBERT-base (MIT) encoder, fine-tuned with reinforcement learning against strictly proper scoring rules (RLCD) and then calibrated per option count. It answers one forward pass at a time and returns a full probability distribution over the options.
License: Gemma. Commercial use is permitted, subject to the Gemma Terms of Use — see
LICENSE,NOTICEandGEMMA_TERMS_OF_USE.md.
Measured results
Every row is a full run on the published test split, not a sample.
| Benchmark | Options | Accuracy | ECE | Brier | n | p50 latency |
|---|---|---|---|---|---|---|
| Banking77 (intent) | 77 | 0.9114 | 0.0293 | 0.1416 | 3080 | 13.3 ms |
| AG News (topic) | 4 | 0.9208 | 0.0226 | 0.1299 | 7600 | 3.8 ms |
| DAIR (emotion) | 6 | 0.8555 | 0.0338 | 0.2197 | 2000 | 3.9 ms |
| CLINC150 (intent + OOS) | 151 | 0.8387 | 0.0532 | 0.2517 | 5500 | 23.1 ms |
| GoEmotions (emotion) | 28 | 0.5646 | 0.0177 | 0.5954 | 5427 | 8.5 ms |
CLINC150 and GoEmotions are internal probes: they are reported here because they exercise large and near-synonymous option sets, but unlike the three sets above they have no widely published reference figure to compare against. See Limitations.
score and noul work on their trained tasks
The model answers three question types. This checkpoint supports all three; the two
non-choice types are graded and binary judgments over an ordered level list and
a yes/no question respectively.
An earlier revision of this card reported that score was unavailable and that
noul worked only in Chinese. That was wrong, and the error was in the probe,
not the model. The probe asked a banking intent-match question built from
CLINC150 while these heads are trained on contract-clause judgments, so it was
measuring cross-task transfer and calling the result a capability gap; a second
probe bug read a nonexistent probability key for noul and returned a constant
0.5 AUC. Probing the tasks the heads were actually trained on — reusing the
training questions verbatim — gives:
| Type | Task | Levels | Train fit (r / acc / AUC) | Holdout (r / acc / AUC) |
|---|---|---|---|---|
score |
clause fairness | 4 | 0.959 | 0.940 (en) / 0.942 (zh) |
score |
document maturity | 4 | 0.999 | 0.999 |
score |
urgency level | 4 | 0.941 | 1.000 |
noul |
risky / one-sided clause | 2 | 0.999 | 0.977 (en) / 0.979 (zh) |
noul |
needs legal review | 2 | 0.998 | 0.984 |
"Holdout" withholds whole contracts the checkpoint never saw (20% of CUAD
contract titles, plus 15% of synthetic records), moved out of the replay pool so
it cannot be sampled back in. Clause fairness lands at r ≈ 0.94 on unseen
contracts versus 0.96 in-sample, and the binary noul tasks at AUC 0.98 on
unseen contracts versus 0.999 in-sample: the heads generalise, they do not just
memorise. English and Chinese asks score within a point of each other, so the
heads are not Chinese-only.
Per-task holdout sample sizes are small for the synthetic tasks (n = 49-56 for
document maturity and urgency), so treat those two as directional; clause
fairness and the noul tasks have n = 500 held-out records each.
Cross-task, transfer is weak and should not be expected. The heads encode the task they were trained on (clause fairness, urgency, risky-clause detection), not "intent match" or any other unlabelled notion. If you need a graded judgment of a different quantity, fine-tune on that quantity rather than pointing these heads at it.
Scaling across option counts
| Options | 4 | 6 | 28 | 77 | 151 |
|---|---|---|---|---|---|
| Accuracy | 0.9208 | 0.8555 | 0.5646 | 0.9114 | 0.8387 |
| Random baseline | 0.2500 | 0.1667 | 0.0357 | 0.0130 | 0.0066 |
It is commonly assumed that choice questions degrade beyond ~20 options, that
Banking77's ~0.43 is an architectural budget wall, and that choice questions
should be kept under ~20 options. None of that holds for this model: every
configuration stays far above its random baseline, and 151 options reaches
0.8387 — 127× random. The head_max_len budget was raised from 512 to 1024
during training, and accuracy is no longer limited by tokens-per-option.
Accuracy is not monotonic in option count, which is worth noting: 28 options
(GoEmotions, 0.5646) is the weakest cell while 151 options scores 0.8387. The
driver is label granularity, not option count. GoEmotions asks for one of 28
overlapping emotions (annoyance/anger, sadness/grief, desire/love)
from comments where 30% are neutral; CLINC asks for one of 151 lexically
distinct intents. If your option list contains near-synonyms, expect the 28-option
row, not the 151-option row, to be the relevant one.
Out-of-scope rejection
CLINC150 splits into two different questions, reported separately because one blended number hides both:
| Group | n | Accuracy | ECE | Mean confidence |
|---|---|---|---|---|
| In-scope intents | 5470 | 0.8386 | 0.0532 | 0.8837 |
| Out-of-scope | 30 | 0.8667 | 0.1227 | 0.8993 |
Treat the 0.8667 OOS row as a signal, not a measurement. 30 samples give a
95% Wilson interval of roughly 0.66–0.96, and the 100 out-of-scope training
records were all retained in the mix, so this row is not fully out-of-sample. For
dependable rejection, either express "none of these" as an explicit option in a
choice question, or use a noul question on a task the head was trained for
(see the noul results above) rather than as a general answerability gate.
⚠️ Load it with DecisionModel, NOT AutoModel
This is the single most important thing to know about this checkpoint.
model.safetensors contains 170 tensors, of which only 134 are the encoder.
The remaining 36 are decision heads:
head.layers.{0,1}.* 2-layer decision head (linear1, linear2, self_attn)
scorer.{0,1,2}.* scoring heads
act_head.{0,2}.* action head
type_emb.weight question-type embedding
temperature calibration tensor
Loading this with transformers will fail or mislead you:
AutoModelForMaskedLM.from_pretrained()raisesValueError: Unrecognized model in ... Should have a 'model_type' key in its config.json— there is no top-levelconfig.json, onlyencoder/config.json.- If you work around that by pointing at
encoder/instead, you get the masked language model with all 36 decision heads silently discarded: the tensor keys areencoder.*whereastransformersexpectsmodel.*, so nothing maps, and you get a stub that assigns confidently wrong labels with no error raised.
Use DecisionModel.load() instead — it selects the correct runtime for you.
Usage
from decision_model import DecisionModel
# hub="hf" downloads from this repo, hub="modelscope" from ModelScope, and a
# local directory path with hub="local" runs fully offline.
model = DecisionModel.load(hub="hf", device="cuda")
answer = model.predict(
"Please check my account: I was charged twice for order #4471 last month.",
{
"pick": {
"type": "choice",
"instructions": "Choose the issue category",
"criteria": {
"billing": "billing",
"shipping": "shipping",
"technical": "technical",
"other": "other",
},
}
},
)
print(answer["answers"]["pick"]["choice"]) # e.g. "billing"
print(answer["answers"]["pick"]["probabilities"]) # calibrated distribution
decision_model.py ships alongside the weights, so run this from the directory
you downloaded (or add that directory to PYTHONPATH).
Install: pip install -r requirements.txt.
Scope and limitations
What this model is. A domain-adapted decision model for single-label multiple-choice classification with calibrated probabilities. It is non- autoregressive: one forward pass returns the full distribution, so confidence is available at no extra cost.
What this model is not. Despite the general in the repository name, this
checkpoint has been fine-tuned on five single-label text classification
datasets plus a fixed set of contract-clause score/noul tasks, and has been
evaluated only on those. Its measured scope is narrow; specifically:
- The reported numbers cover Banking77, AG News, DAIR, CLINC150 and GoEmotions. Performance on tasks outside these is not measured and unknown. In particular the application workflows that matter in production — spam and phishing filtering, jailbreak detection, toxicity moderation, RAG relevance, ticket triage — are not evaluated here, so treat accuracy on those as unestablished. The five sets above are all single-label text classification.
score/noulgeneralisation is measured, not assumed. The table under "scoreandnoulwork on their trained tasks" reports both the training fit and a holdout over contracts the checkpoint never saw. Generalisation is close to the fit (clause fairness r 0.94 vs 0.96) but the two synthetic tasks (document maturity, urgency) rest on small holdouts (n = 49-56) and should be treated as directional.- The
score/noulheads are task-specific. They were trained on contract clause judgments (fairness, maturity, urgency, risky clause, legal review). Pointing them at a different graded or binary question is cross-task transfer and is expected to be weak; fine-tune on the target task instead. - A cross-task
score/noulprobe built from CLINC150 fails, and an earlier version of this card wrongly presented that as the capability being broken. The failure is the task mismatch, not the head; see the section above. - The OOS figure of 0.8667 comes from 30 samples (95% Wilson interval ≈ 0.70–0.95), and the 100 out-of-scope training records were retained in the mix, so it is not fully out-of-sample. Validate rejection on your own traffic.
- GoEmotions is the weakest cell at 0.5646 and is the realistic reference for any option list built from near-synonymous labels. Note that the source data is multi-label while this evaluation set projects to single-label by taking the first label, which makes the figure a lower bound.
- CLINC150 and GoEmotions are internal probes with no widely published reference figure to compare against, so they are not presented as comparisons.
- No claim is made about generative ability. This is a discriminator over a fixed option list; it cannot produce free-form text.
- Option lists beyond 151 are untested. Accuracy varies non-monotonically with option count (see the table above), so do not extrapolate upward.
- Temperatures were fitted post-hoc on these same evaluation sets. The
per-option-count temperatures were selected against the full evaluation
distributions, so the accuracy figures are unaffected (a
temperature-scaled softmax preserves argmax at every positive temperature) but
the ECE figures are not from independent held-out calibration data.
Calibration quality on genuinely unseen data is unverified. The
scoreandnoultypes carry no fitted temperature at all — the calibration set is Banking77-only, so both fell back to 1.00. - Multilingual capability is inherited from the mmBERT encoder and is not
preserved or verified by this fine-tune, whose
choicedata is English and whosescore/nouldata is mixed Chinese/English. Thescore/noulheads do work when asked in either language on their trained tasks (see above), but that is a narrow, task-specific result. Treat broad cross-language transfer as unestablished.
The general label refers to the project's intent and to the model's capability
profile (any option list, typed decisions, calibrated output), not to an
evaluation breadth that has not been demonstrated. Treat the five benchmarks
above as the actual measured scope, and everything else as unknown.
Calibration
Probability quality is a first-class property of this release, not an afterthought. ECE is under 0.06 on the three public classification benchmarks, and confidence is fitted per option count rather than globally:
| Bucket | Options | Temperature | Benchmark | Resulting ECE |
|---|---|---|---|---|
choice:3-5 |
4 | 1.2 | AG News | 0.0226 |
choice:6-10 |
6 | 1.2 | DAIR | 0.0338 |
choice:11+ |
77 | 1.60 | Banking77 | 0.0152 |
choice:11+ |
28 | 1.40 | GoEmotions | 0.0330 |
choice:11+ |
151 | 1.4724 | CLINC150 (bucket mean) | 0.0532 |
A single shared temperature cannot fit all buckets: the per-bucket ECE optima differ, and forcing one value onto all buckets pushed AG News ECE from 0.022 to 0.135 in testing.
Because softmax(z / T) preserves the argmax, this step is calibration-only —
accuracy is provably unaffected. This was verified across T ∈ [0.7, 3.0].
Training recipe
| Item | Value |
|---|---|
| Base checkpoint | convaiinnovations/laya-multilingual |
| Encoder | jhu-clsp/mmBERT-base (307M params, 22 layers, hidden 768) |
| Objective | RL against strictly proper scoring rules (RLCD) + cross-entropy |
| Epochs | 2 (avg loss 0.4611 → 0.5065) |
| Training records | 62,446 tokenized items from 35,977 records |
| Source records | 30,581 new + 5,396 replay |
| New data mix | Banking77 10,003 · GoEmotions 8,000 · CLINC150 6,000 · AG News 2,000 · DAIR 1,200 · noul/score 3,378 records (10,458 + 4,868 questions, holdout excluded) |
| Holdout | 809 records (20% of CUAD contract titles + 15% of synthetic records), never trained on |
sigma_start → sigma_end |
0.4 → 0.05 |
| Replay ratio | 0.15 |
| Micro-batch / grad-accum | 16 / 2 |
| Max sequence length | 1,024 tokens |
| Checkpoint interval | 900 s |
| Wall clock | ~8.3 h on one RTX 3090 (2 epochs) |
| Hardware note | Sustained 95 °C under bf16; thermal cap, not a driver fault |
Mixing in AG News and DAIR was the decisive choice: training on Banking77 alone left AG News at 0.9093 and DAIR at 0.5350, whereas the three-way mix raised all three benchmarks, including Banking77 itself. Cross-domain data was a gain, not a distraction.
Provenance and licensing
This checkpoint is a Model Derivative in the sense defined by the Gemma Terms of Use, because its tokenizer derives from Gemma 2. Distribution is therefore subject to those terms.
| Layer | Component | License |
|---|---|---|
| Framework | Laya | Apache-2.0 |
| Encoder | jhu-clsp/mmBERT-base |
MIT |
| Tokenizer | Gemma 2 (256K vocab) | Gemma Terms of Use |
Before redistributing this model or any derivative of it, read
GEMMA_TERMS_OF_USE.md. In short, Section 3.1 requires that recipients receive
a copy of the terms, that the Section 3.2 use restrictions travel with the
license as enforceable provisions, that modified files be identified, and that
this NOTICE file accompany the distribution. All four are satisfied here:
LICENSE— Apache-2.0 with the Gemma Section 3.2 restrictions incorporated as binding termsGEMMA_TERMS_OF_USE.md— full terms and the required noticeNOTICE— attributions plus an explicit list of modified files- Modified files:
model.safetensors,rl_agent_config.json,encoder/config.json(tokenizer files are unmodified)
The sizeai-authored portions (fine-tuning, calibration, training and release packaging) are offered under Apache-2.0; that does not narrow or replace the Gemma terms, which continue to govern the tokenizer.
Commercial use. Nothing here makes the model non-commercial. The Gemma Terms permit commercial use and redistribution; they attach conditions (the four obligations above) and prohibit only the uses listed in the Gemma Prohibited Use Policy. In practice this checkpoint can be used and redistributed commercially, provided recipients accept the Gemma Terms. Organisations whose procurement policy excludes Gemma-derived artefacts should review that policy before adoption, since the restriction is contractual rather than technical.
Intended use
Appropriate for: routing and triage over a fixed label set, intent classification, support-ticket categorisation, and any decision where a calibrated confidence is required — for example gating on confidence rather than only on the argmax.
Not appropriate for: high-stakes automated action on a single low-confidence prediction, tasks outside the evaluated domain, or as a general-purpose language model.
Citation
Technical report: https://github.com/beidald/size-decision-general
@techreport{sizeai2026calibrated,
title = {A calibrated encoder classifier as the cost floor for production decisions},
author = {{sizeai}},
year = {2026},
institution = {sizeai},
url = {https://github.com/beidald/size-decision-general},
note = {Technical report for the size-decision-general model.}
}
@misc{sizeai2026sizedecisiongeneral,
title = {size-decision-general: a calibrated decision model},
author = {{sizeai}},
year = {2026},
howpublished = {\url{https://huggingface.co/sizeai/size-decision-general}},
note = {Based on Laya (Apache-2.0) and mmBERT-base (MIT);
tokenizer derived from Gemma 2, subject to the Gemma Terms of Use.}
}
References
Model tree for sizeai/size-decision-general
Base model
convaiinnovations/laya-multilingual