size-decision-general

A calibrated multiple-choice decision model with an mmBERT-base (MIT) encoder, fine-tuned with reinforcement learning against strictly proper scoring rules (RLCD) and then calibrated per option count. It answers one forward pass at a time and returns a full probability distribution over the options.

License: Gemma. Commercial use is permitted, subject to the Gemma Terms of Use — see LICENSE, NOTICE and GEMMA_TERMS_OF_USE.md.

Measured results

Every row is a full run on the published test split, not a sample.

Benchmark Options Accuracy ECE Brier n p50 latency
Banking77 (intent) 77 0.9114 0.0293 0.1416 3080 13.3 ms
AG News (topic) 4 0.9208 0.0226 0.1299 7600 3.8 ms
DAIR (emotion) 6 0.8555 0.0338 0.2197 2000 3.9 ms
CLINC150 (intent + OOS) 151 0.8387 0.0532 0.2517 5500 23.1 ms
GoEmotions (emotion) 28 0.5646 0.0177 0.5954 5427 8.5 ms

CLINC150 and GoEmotions are internal probes: they are reported here because they exercise large and near-synonymous option sets, but unlike the three sets above they have no widely published reference figure to compare against. See Limitations.

score and noul work on their trained tasks

The model answers three question types. This checkpoint supports all three; the two non-choice types are graded and binary judgments over an ordered level list and a yes/no question respectively.

An earlier revision of this card reported that score was unavailable and that noul worked only in Chinese. That was wrong, and the error was in the probe, not the model. The probe asked a banking intent-match question built from CLINC150 while these heads are trained on contract-clause judgments, so it was measuring cross-task transfer and calling the result a capability gap; a second probe bug read a nonexistent probability key for noul and returned a constant 0.5 AUC. Probing the tasks the heads were actually trained on — reusing the training questions verbatim — gives:

Type Task Levels Train fit (r / acc / AUC) Holdout (r / acc / AUC)
score clause fairness 4 0.959 0.940 (en) / 0.942 (zh)
score document maturity 4 0.999 0.999
score urgency level 4 0.941 1.000
noul risky / one-sided clause 2 0.999 0.977 (en) / 0.979 (zh)
noul needs legal review 2 0.998 0.984

"Holdout" withholds whole contracts the checkpoint never saw (20% of CUAD contract titles, plus 15% of synthetic records), moved out of the replay pool so it cannot be sampled back in. Clause fairness lands at r ≈ 0.94 on unseen contracts versus 0.96 in-sample, and the binary noul tasks at AUC 0.98 on unseen contracts versus 0.999 in-sample: the heads generalise, they do not just memorise. English and Chinese asks score within a point of each other, so the heads are not Chinese-only.

Per-task holdout sample sizes are small for the synthetic tasks (n = 49-56 for document maturity and urgency), so treat those two as directional; clause fairness and the noul tasks have n = 500 held-out records each.

Cross-task, transfer is weak and should not be expected. The heads encode the task they were trained on (clause fairness, urgency, risky-clause detection), not "intent match" or any other unlabelled notion. If you need a graded judgment of a different quantity, fine-tune on that quantity rather than pointing these heads at it.

Scaling across option counts

Options 4 6 28 77 151
Accuracy 0.9208 0.8555 0.5646 0.9114 0.8387
Random baseline 0.2500 0.1667 0.0357 0.0130 0.0066

It is commonly assumed that choice questions degrade beyond ~20 options, that Banking77's ~0.43 is an architectural budget wall, and that choice questions should be kept under ~20 options. None of that holds for this model: every configuration stays far above its random baseline, and 151 options reaches 0.8387 — 127× random. The head_max_len budget was raised from 512 to 1024 during training, and accuracy is no longer limited by tokens-per-option.

Accuracy is not monotonic in option count, which is worth noting: 28 options (GoEmotions, 0.5646) is the weakest cell while 151 options scores 0.8387. The driver is label granularity, not option count. GoEmotions asks for one of 28 overlapping emotions (annoyance/anger, sadness/grief, desire/love) from comments where 30% are neutral; CLINC asks for one of 151 lexically distinct intents. If your option list contains near-synonyms, expect the 28-option row, not the 151-option row, to be the relevant one.

Out-of-scope rejection

CLINC150 splits into two different questions, reported separately because one blended number hides both:

Group n Accuracy ECE Mean confidence
In-scope intents 5470 0.8386 0.0532 0.8837
Out-of-scope 30 0.8667 0.1227 0.8993

Treat the 0.8667 OOS row as a signal, not a measurement. 30 samples give a 95% Wilson interval of roughly 0.66–0.96, and the 100 out-of-scope training records were all retained in the mix, so this row is not fully out-of-sample. For dependable rejection, either express "none of these" as an explicit option in a choice question, or use a noul question on a task the head was trained for (see the noul results above) rather than as a general answerability gate.


⚠️ Load it with DecisionModel, NOT AutoModel

This is the single most important thing to know about this checkpoint.

model.safetensors contains 170 tensors, of which only 134 are the encoder. The remaining 36 are decision heads:

head.layers.{0,1}.*      2-layer decision head (linear1, linear2, self_attn)
scorer.{0,1,2}.*         scoring heads
act_head.{0,2}.*         action head
type_emb.weight          question-type embedding
temperature              calibration tensor

Loading this with transformers will fail or mislead you:

  • AutoModelForMaskedLM.from_pretrained() raises ValueError: Unrecognized model in ... Should have a 'model_type' key in its config.json — there is no top-level config.json, only encoder/config.json.
  • If you work around that by pointing at encoder/ instead, you get the masked language model with all 36 decision heads silently discarded: the tensor keys are encoder.* whereas transformers expects model.*, so nothing maps, and you get a stub that assigns confidently wrong labels with no error raised.

Use DecisionModel.load() instead — it selects the correct runtime for you.

Usage

from decision_model import DecisionModel

# hub="hf" downloads from this repo, hub="modelscope" from ModelScope, and a
# local directory path with hub="local" runs fully offline.
model = DecisionModel.load(hub="hf", device="cuda")

answer = model.predict(
    "Please check my account: I was charged twice for order #4471 last month.",
    {
        "pick": {
            "type": "choice",
            "instructions": "Choose the issue category",
            "criteria": {
                "billing": "billing",
                "shipping": "shipping",
                "technical": "technical",
                "other": "other",
            },
        }
    },
)

print(answer["answers"]["pick"]["choice"])          # e.g. "billing"
print(answer["answers"]["pick"]["probabilities"])    # calibrated distribution

decision_model.py ships alongside the weights, so run this from the directory you downloaded (or add that directory to PYTHONPATH).

Install: pip install -r requirements.txt.


Scope and limitations

What this model is. A domain-adapted decision model for single-label multiple-choice classification with calibrated probabilities. It is non- autoregressive: one forward pass returns the full distribution, so confidence is available at no extra cost.

What this model is not. Despite the general in the repository name, this checkpoint has been fine-tuned on five single-label text classification datasets plus a fixed set of contract-clause score/noul tasks, and has been evaluated only on those. Its measured scope is narrow; specifically:

  • The reported numbers cover Banking77, AG News, DAIR, CLINC150 and GoEmotions. Performance on tasks outside these is not measured and unknown. In particular the application workflows that matter in production — spam and phishing filtering, jailbreak detection, toxicity moderation, RAG relevance, ticket triage — are not evaluated here, so treat accuracy on those as unestablished. The five sets above are all single-label text classification.
  • score/noul generalisation is measured, not assumed. The table under "score and noul work on their trained tasks" reports both the training fit and a holdout over contracts the checkpoint never saw. Generalisation is close to the fit (clause fairness r 0.94 vs 0.96) but the two synthetic tasks (document maturity, urgency) rest on small holdouts (n = 49-56) and should be treated as directional.
  • The score/noul heads are task-specific. They were trained on contract clause judgments (fairness, maturity, urgency, risky clause, legal review). Pointing them at a different graded or binary question is cross-task transfer and is expected to be weak; fine-tune on the target task instead.
  • A cross-task score/noul probe built from CLINC150 fails, and an earlier version of this card wrongly presented that as the capability being broken. The failure is the task mismatch, not the head; see the section above.
  • The OOS figure of 0.8667 comes from 30 samples (95% Wilson interval ≈ 0.70–0.95), and the 100 out-of-scope training records were retained in the mix, so it is not fully out-of-sample. Validate rejection on your own traffic.
  • GoEmotions is the weakest cell at 0.5646 and is the realistic reference for any option list built from near-synonymous labels. Note that the source data is multi-label while this evaluation set projects to single-label by taking the first label, which makes the figure a lower bound.
  • CLINC150 and GoEmotions are internal probes with no widely published reference figure to compare against, so they are not presented as comparisons.
  • No claim is made about generative ability. This is a discriminator over a fixed option list; it cannot produce free-form text.
  • Option lists beyond 151 are untested. Accuracy varies non-monotonically with option count (see the table above), so do not extrapolate upward.
  • Temperatures were fitted post-hoc on these same evaluation sets. The per-option-count temperatures were selected against the full evaluation distributions, so the accuracy figures are unaffected (a temperature-scaled softmax preserves argmax at every positive temperature) but the ECE figures are not from independent held-out calibration data. Calibration quality on genuinely unseen data is unverified. The score and noul types carry no fitted temperature at all — the calibration set is Banking77-only, so both fell back to 1.00.
  • Multilingual capability is inherited from the mmBERT encoder and is not preserved or verified by this fine-tune, whose choice data is English and whose score/noul data is mixed Chinese/English. The score/noul heads do work when asked in either language on their trained tasks (see above), but that is a narrow, task-specific result. Treat broad cross-language transfer as unestablished.

The general label refers to the project's intent and to the model's capability profile (any option list, typed decisions, calibrated output), not to an evaluation breadth that has not been demonstrated. Treat the five benchmarks above as the actual measured scope, and everything else as unknown.


Calibration

Probability quality is a first-class property of this release, not an afterthought. ECE is under 0.06 on the three public classification benchmarks, and confidence is fitted per option count rather than globally:

Bucket Options Temperature Benchmark Resulting ECE
choice:3-5 4 1.2 AG News 0.0226
choice:6-10 6 1.2 DAIR 0.0338
choice:11+ 77 1.60 Banking77 0.0152
choice:11+ 28 1.40 GoEmotions 0.0330
choice:11+ 151 1.4724 CLINC150 (bucket mean) 0.0532

A single shared temperature cannot fit all buckets: the per-bucket ECE optima differ, and forcing one value onto all buckets pushed AG News ECE from 0.022 to 0.135 in testing.

Because softmax(z / T) preserves the argmax, this step is calibration-only — accuracy is provably unaffected. This was verified across T ∈ [0.7, 3.0].


Training recipe

Item Value
Base checkpoint convaiinnovations/laya-multilingual
Encoder jhu-clsp/mmBERT-base (307M params, 22 layers, hidden 768)
Objective RL against strictly proper scoring rules (RLCD) + cross-entropy
Epochs 2 (avg loss 0.4611 → 0.5065)
Training records 62,446 tokenized items from 35,977 records
Source records 30,581 new + 5,396 replay
New data mix Banking77 10,003 · GoEmotions 8,000 · CLINC150 6,000 · AG News 2,000 · DAIR 1,200 · noul/score 3,378 records (10,458 + 4,868 questions, holdout excluded)
Holdout 809 records (20% of CUAD contract titles + 15% of synthetic records), never trained on
sigma_start → sigma_end 0.4 → 0.05
Replay ratio 0.15
Micro-batch / grad-accum 16 / 2
Max sequence length 1,024 tokens
Checkpoint interval 900 s
Wall clock ~8.3 h on one RTX 3090 (2 epochs)
Hardware note Sustained 95 °C under bf16; thermal cap, not a driver fault

Mixing in AG News and DAIR was the decisive choice: training on Banking77 alone left AG News at 0.9093 and DAIR at 0.5350, whereas the three-way mix raised all three benchmarks, including Banking77 itself. Cross-domain data was a gain, not a distraction.


Provenance and licensing

This checkpoint is a Model Derivative in the sense defined by the Gemma Terms of Use, because its tokenizer derives from Gemma 2. Distribution is therefore subject to those terms.

Layer Component License
Framework Laya Apache-2.0
Encoder jhu-clsp/mmBERT-base MIT
Tokenizer Gemma 2 (256K vocab) Gemma Terms of Use

Before redistributing this model or any derivative of it, read GEMMA_TERMS_OF_USE.md. In short, Section 3.1 requires that recipients receive a copy of the terms, that the Section 3.2 use restrictions travel with the license as enforceable provisions, that modified files be identified, and that this NOTICE file accompany the distribution. All four are satisfied here:

  • LICENSE — Apache-2.0 with the Gemma Section 3.2 restrictions incorporated as binding terms
  • GEMMA_TERMS_OF_USE.md — full terms and the required notice
  • NOTICE — attributions plus an explicit list of modified files
  • Modified files: model.safetensors, rl_agent_config.json, encoder/config.json (tokenizer files are unmodified)

The sizeai-authored portions (fine-tuning, calibration, training and release packaging) are offered under Apache-2.0; that does not narrow or replace the Gemma terms, which continue to govern the tokenizer.

Commercial use. Nothing here makes the model non-commercial. The Gemma Terms permit commercial use and redistribution; they attach conditions (the four obligations above) and prohibit only the uses listed in the Gemma Prohibited Use Policy. In practice this checkpoint can be used and redistributed commercially, provided recipients accept the Gemma Terms. Organisations whose procurement policy excludes Gemma-derived artefacts should review that policy before adoption, since the restriction is contractual rather than technical.


Intended use

Appropriate for: routing and triage over a fixed label set, intent classification, support-ticket categorisation, and any decision where a calibrated confidence is required — for example gating on confidence rather than only on the argmax.

Not appropriate for: high-stakes automated action on a single low-confidence prediction, tasks outside the evaluated domain, or as a general-purpose language model.


Citation

Technical report: https://github.com/beidald/size-decision-general

@techreport{sizeai2026calibrated,
  title       = {A calibrated encoder classifier as the cost floor for production decisions},
  author      = {{sizeai}},
  year        = {2026},
  institution = {sizeai},
  url         = {https://github.com/beidald/size-decision-general},
  note        = {Technical report for the size-decision-general model.}
}

@misc{sizeai2026sizedecisiongeneral,
  title         = {size-decision-general: a calibrated decision model},
  author        = {{sizeai}},
  year          = {2026},
  howpublished  = {\url{https://huggingface.co/sizeai/size-decision-general}},
  note          = {Based on Laya (Apache-2.0) and mmBERT-base (MIT);
                  tokenizer derived from Gemma 2, subject to the Gemma Terms of Use.}
}

References

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sizeai/size-decision-general

Finetuned
(66)
this model

Paper for sizeai/size-decision-general