Mapika commited on
Commit
eb5fbdf
·
verified ·
1 Parent(s): 49564dd

decider-4b v2.1

Browse files

decider-4b v1 + merged LoRA (rank 64, attention and MLP) with replay trained toward v1's distribution; temperature 1.099, temperature_by_type choice 1.110 / noul 1.560 / score 1.287 (decider-ai 1.4.0). v2 is under the tag v2, v1 under v1.

README.md CHANGED
@@ -17,93 +17,126 @@ Base model: [Qwen/Qwen3.5-4B-Base](https://huggingface.co/Qwen/Qwen3.5-4B-Base):
17
  attention and 24 with gated delta-net linear attention, hidden size 2,560. Two supervised stages. Stage 1 (decider-4b v1): one
18
  pass of cross-entropy on the slot readout over mixture v2, the public decision mixture of
19
  [decider-2b](https://huggingface.co/Mapika/decider-2b) plus 26 further public decision datasets and ten programmatically
20
- generated families with verifiable gold (742M tokens). Stage 2 (v2): a LoRA of rank 64 on the attention and MLP weights,
21
- trained for 2 epochs on 29,356 rows of harder decisions and replay, then merged into the weights. There is no
22
- reinforcement-learning stage. **This repository holds v2**, the bf16 weights (8.4 GB). v1 stays available under the Hub tag
23
- `v1` (see [Changes from v1](#changes-from-v1) for who should keep using it). The other sizes are listed under The decider
24
- family. `decider/` in this repository is the inference subset of the GitHub package.
25
-
26
- Against decider-2b v10 on the same rows: accuracy is higher on 66 of the 95 regression tasks (in-task 0.824 against 0.805,
27
- held-out 0.779 against 0.755), +1.8 points on the 847 validation rows, +3.5 on OpenJev, +4.8 on Mind2Web, +5.9 on the TypeSafe
28
- workflow rows (interval includes zero), JevBench hard tier 0.676 against 0.459, Bespoke's public suite 0.773 against 0.704 macro.
29
- Against decider-35b-a3b it is 3.1 to 3.2 points lower on the regression set, 1.6 to 5.0 points lower on three fixtures, level on
30
- TypeSafe and level on the JevBench hard tier (0.676 both), at 8.4 GB instead of 65 GB. On live browser tasks it plays at the level
31
- of the 2B in greedy mode (92.6% against 90.9%) and below it on the six held-out tasks (81.2% against 91.7%), because it has no RL
32
- stage. Details under Evaluation.
33
-
34
- **Contents:** [Changes from v1](#changes-from-v1) · [The decider family](#the-decider-family) · [Usage](#usage) · [How it works](#how-it-works) · [Training](#training) · [Evaluation](#evaluation) · [Calibration](#calibration) · [Speed](#speed) · [Limitations](#limitations) · [Changelog](#changelog) · [Reproduction](#reproduction)
35
-
36
- ## Changes from v1
37
-
38
- v2 is v1 plus the stage-2 LoRA described under Training. It is better than v1 on hard and document-based judgments and on
39
- TypeSafe, OpenJev and Bespoke's suite, and worse on sampled game and browser play, on some text games and by about 1 point on
40
- everyday tasks. All rows below are on identical inputs, v1 at its stored temperature 1.05 and v2 at 1.935, through decider-ai
41
- 1.2.1; intervals are 95% paired bootstrap intervals of v2 minus v1 (tasks, rows, items, boards or task-seed pairs).
42
-
43
- | set | v1 | v2 | v2 minus v1 |
44
- |---|---|---|---|
45
- | regression set, 67 in-task tasks, accuracy / NLL / ECE | 0.834 / 0.404 / 0.027 | 0.824 / 0.441 / 0.041 | accuracy −1.0 (−1.2 to −0.7), NLL +0.037 (+0.029 to +0.046), ECE +0.014 (+0.008 to +0.021) |
46
- | regression set, 28 held-out tasks | 0.788 / 0.558 / 0.071 | 0.779 / 0.566 / 0.080 | accuracy −0.9 (−2.0 to +0.1), NLL +0.008 (−0.024 to +0.035), ECE +0.009 (−0.003 to +0.021) |
47
- | 847 in-task validation rows, accuracy / NLL | 86.1% / 0.417 | 85.0% / 0.419 | −1.1 (−2.6 to +0.5); NLL +0.002 (−0.019 to +0.021) |
48
- | OpenJev, 5,252 rows | 63.9% / 0.893 | 66.7% / 0.789 | +2.8 (+1.8 to +3.9); NLL −0.104 (−0.121 to −0.087) |
49
- | Mind2Web, 1,770 rows | 88.4% / 0.366 | 87.4% / 0.393 | −1.0 (−2.0 to +0.1); NLL +0.027 (+0.011 to +0.044) |
50
- | TypeSafe workflow decisions, 102 rows | 81.4% / 0.611 | 86.3% / 0.410 | +4.9 (−1.0 to +10.8); NLL −0.201 (−0.348 to −0.072) |
51
- | Bespoke's public suite, macro / micro | 0.757 / 0.765 | 0.773 / 0.781 | macro +1.6 |
52
- | JevBench public items, easy / standard / hard | 1.000 / 0.958 / 0.550 | 1.000 / 0.986 / 0.676 | hard +12.6 (+4.5 to +20.7) |
53
- | JevBench hard tier, top-label ECE (v2: recomputed at T 1.935 from stored probabilities, the 6 Score items at T 1.719) | 0.288 | 0.071 | |
54
- | Decision Index 4,000-request sample, calibration error (site definition; v2 recomputed at T 1.935 from stored probabilities) | 0.086 | 0.074 | |
55
- | live MiniWoB++, 22 tasks x 8 seeds, greedy (all / 6 held-out tasks) | 91.5% / 75.0% | 92.6% / 81.2% | +1.1 (−1.7 to +4.0) / +6.2 (−2.1 to +16.7) |
56
- | live MiniWoB++, sampled (all / 16 rewarded / 6 held-out) | 90.9% / 96.1% / 77.1% | 88.1% / 93.0% / 75.0% | −2.8 (−7.4 to +1.7) / −3.1 (−7.8 to +0.8) / −2.1 (−14.6 to +10.4) |
57
- | zero-shot games, win rate, sampled (234 boards) | 27.8% | 22.4% | −5.3 (−7.6 to −3.1) |
58
- | zero-shot games, greedy | 29.1% | 27.8% | −1.3 (−5.6 to +2.6) |
59
- | bag-draw games alone, sampled / greedy | 56.6% / 62.5% | 37.9% / 48.4% | −18.8 (−24.6 to −12.9) / −14.1 (−23.4 to −6.2) |
60
- | ten text games, greedy: CliffWalking / BabyAI-GoTo / Breakout | −13 / 0.54 / 14 | −60 / 0.35 / 12 | |
61
- | behaviour probes: model-router tier / needs-live-data / touches-outside-project | 0.968 / 0.871 / 0.956 | 0.935 / 0.839 / 0.933 | |
62
- | behaviour probes: abstention battery / catch-all / command risk / browser element and action | 7 of 8 / 0.90 / 0.889 / 0.875 and 0.875 | 8 of 8 / 0.95 / 0.911 / 0.938 and 0.938 | |
63
-
64
- The v1 column is v1 measured again on 2026-09-24 through decider-ai 1.2.1, so that both versions are read by the same code on
65
- the same day. It differs from the numbers first published for v1 by at most 1.0 point on the regression set, the fixtures and
66
- JevBench, by up to 2.1 points on the live browser and game rows (sampled held-out browser tasks 77.1% against 79.2%, bag-draw
67
- greedy 62.5% against 60.9%), and on Breakout (14 against 18). The first-published values are in the v1 card under the tag `v1`.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
68
 
69
  **Regressions, stated plainly.**
70
- * Sampled play in games: the bag-draw games (choose the bag with the highest expected value) fall from 56.6% to 37.9% wins in
71
- sampled play, and zero-shot games overall from 27.8% to 22.4%. An earlier candidate (the Qwen3.5-4B instruct model plus a LoRA on
72
- 15,768 of these 29,356 rows) showed a drop of 12.5 points on the same bag-draw games; this points to the new data, but it was not tested
73
- directly.
74
- * Sampled browser play: −2.8 points over all 22 tasks and −3.1 on the 16 rewarded tasks (intervals include zero); the losses
75
- are on click-dialog-2, click-tab, click-checkboxes-large, click-checkboxes-soft, click-checkboxes-transfer and click-option,
76
- the gains on click-tab-2, click-tab-2-hard and focus-text-2.
77
- * Text games (greedy): CliffWalking −60 against −13, BabyAI-GoTo 0.35 against 0.54, Breakout 12 against 14.
78
- * Behaviour probes: model-router tier 0.935 against 0.968, needs-live-data 0.839 against 0.871, touches-outside-project 0.933
79
- against 0.956, generic bucket choice 0.95 against 1.00.
80
- * Form filling: on the two form-filling cases reported in issue #9 (choose the document entity for a form field, or skip /
81
- click / check), v1 answers both correctly (0.79 and 0.98 on the gold option). v2 answers one wrongly ("skip" at 0.68, gold
82
- 0.16) and the other correctly by 0.02 (0.28 against "click" at 0.26). Two cases do not measure a rate, but for form filling
83
- we suggest v1.
84
- * Everyday tasks: regression in-task accuracy −1.0 point and held-out −0.9 (the held-out interval includes zero), lower on 75
85
- of 95 tasks, mostly by under 3 points; the largest losses are CommitmentBank −8.9, PubMedQA −6.2, hate-speech tweets −4.7,
86
- emotion −4.3, LIAR −3.9, QuALITY (full) −3.3, MedQA −3.0. Regression-set calibration is worse (in-task ECE 0.041 against
87
- 0.027), and rating (Score) tasks are the least calibrated in the probe (see Calibration).
88
-
89
- If you rely on sampled play (games, browser agents that sample actions) or on the probes above, pin revision `v1`. The
90
- package loads a local folder, so download the revision first:
 
 
 
 
 
 
 
 
 
 
 
 
91
 
92
  ```python
93
  from huggingface_hub import snapshot_download
94
  from decider.infer import Decider
95
- d = Decider(snapshot_download("Mapika/decider-4b", revision="v1"))
96
  ```
97
 
98
  For the HTTP server, set `DECIDER_MODEL` to the same downloaded folder.
99
 
100
- **How v2 was chosen.** The training run had a pre-registered rule for replacing v1: the candidate had to beat an earlier
101
- candidate (the Qwen3.5-4B instruct model plus a LoRA on 15,768 of the same hard-decision rows) on our own held-out hard sets. v2 did not meet it:
102
- it was 0.8 points short on the held-out teacher-written set and 0.5 points short on the held-out generated families, that is, it
103
- tied the earlier candidate there. The earlier candidate had lost v1's everyday skills and v2 keeps them within about 1 point, so
104
- v2 was then measured on every set of this card, and the release was decided on that full comparison with v1. The JevBench public
105
- items were read once for v2, after selection; they were not used for training, selection or the temperature.
106
-
107
  ## The decider family
108
 
109
  All six repositories share one interface (`decider.infer.Decider`, `POST /v1/systemone` in TypeSafe's format) and one
@@ -111,8 +144,8 @@ readout: the letter logits at an answer slot, softmaxed over the options. Pick b
111
 
112
  | model | base | weights | use it for | numbers |
113
  |---|---|---|---|---|
114
- | [decider-2b](https://huggingface.co/Mapika/decider-2b) v10 | Qwen3.5-2B-Base | 3.5 GB bf16 | the default: routing, classification, judgments, browser agents; 4 ms per request with CUDA graphs on one GPU | regression set 0.805 in-task / 0.755 held-out; live browser 93%; Bespoke suite 0.704 |
115
- | [decider-4b](https://huggingface.co/Mapika/decider-4b) v2 | Qwen3.5-4B-Base | 8.4 GB bf16 | the middle point: knowledge questions and hard judgments above the 2B in a dense 8.4 GB model; no RL stage; v1 under the tag `v1` for sampled play | 0.824 / 0.779, above the 2B on 66 of 95 tasks; JevBench hard 0.676; Bespoke 0.773 |
116
  | [decider-35b-a3b](https://huggingface.co/Mapika/decider-35b-a3b) v1 | Qwen3.5-35B-A3B-Base (3B active) | 65 GB bf16 | when accuracy is worth 3 to 4 times the cost per decision: knowledge and multi-step questions, long policies | 0.855 / 0.810, above the 2B on 93 of 95 tasks; JevBench hard 0.676; Bespoke 0.774; no RL stage |
117
  | [decider-35b-a3b-nvfp4](https://huggingface.co/Mapika/decider-35b-a3b-nvfp4) | the 35B in NVFP4 | 19.6 GB | the 35B on Blackwell through vLLM or TensorRT-LLM | 1.0 to 1.5 points under bf16 on the measured fixtures |
118
  | [decider-0.8b](https://huggingface.co/Mapika/decider-0.8b) | Qwen3.5-0.8B-Base | 1.4 GB bf16 | the smallest: routing, yes/no and short-state lookups within 1 to 4 points of the 2B, 1.5x faster | 0.776 / 0.707 on the single-run protocol (2B: 0.809 / 0.739) |
@@ -133,19 +166,22 @@ d.decide("My card was charged twice for the same purchase.",
133
 
134
  The API is the same as decider-2b's: `decide_batch` scores many states with many questions in one call, `abstain_below=t`
135
  returns `None` under a confidence threshold, a question can have 2 to 255 options, and `system_one` / `decider.serve` accept
136
- TypeSafe's `POST /v1/systemone` request shape (the official `typesafe-sdk` works with `TYPESAFE_BASE_URL` pointing at the
137
- server; checked with decider-2b, not again with this model). Every question and every Score level is scored in its own row. The state may be a string, object or array of up to
138
- 32k tokens. See the decider-2b card for the full description of the request shape, field types and the schema cache.
139
 
140
  Requirements: `torch`, `transformers>=5`, and `flash-linear-attention` (Triton kernels for the Qwen3.5 linear-attention layers;
141
- the model runs without it but several times slower). The weights take 8.4 GB in bf16. v2 uses the plain prompt layout, as v1
142
- does, so it needs no new package version. decider-ai 1.0.2, 1.1.4, 1.2.1 and 1.2.2 and the `decider/` subset in this repository
143
- (taken from 1.2.2) were checked on CUDA: each loads v2 with the name `decider-4b-v2` and temperature 1.935 and gives the same
144
- probabilities to four decimals on the usage example above, on the eager path and on the CUDA-graph path (the two paths differ in
145
- the third decimal, as for every model). The 1.2.2 HTTP server was checked with one `/v1/systemone` request. Other versions,
146
- devices and the FP8 path were not checked on v2. On Blackwell GPUs use 1.0.2 or later (1.0.0 and 1.0.1 have the cuDNN attention fault fixed
147
- in 1.0.2). The model is dense, so the CUDA-graph engine, `torch.compile` and the FP8 path of the helper package apply to it as
148
- to decider-2b (`use_graphs=False` selects eager PyTorch). The measurements below were taken with the eager path unless stated.
 
 
 
149
 
150
  Without the helper package, the same computation in plain `transformers`:
151
 
@@ -159,7 +195,7 @@ ids = tok(prompt, return_tensors="pt").to("cuda")
159
  with torch.no_grad():
160
  logits = m(**ids).logits[0, -1]
161
  letters = [tok.encode(L, add_special_tokens=False)[0] for L in "ABC"]
162
- probs = torch.softmax(logits[letters].float() / 1.935, -1) # 1.935 is the stored temperature (v1: 1.05)
163
  ```
164
 
165
  ## How it works
@@ -167,8 +203,13 @@ probs = torch.softmax(logits[letters].float() / 1.935, -1) # 1.935 is the st
167
  The prompt is `Context: ...` followed by, for each question, the question text, the lettered options `(A) ... (B) ...` and an
168
  answer slot `Answer k: (`. The hidden state at each slot is projected with the option-letter rows of the LM head and softmaxed
169
  over the valid letters, divided by the temperature in `decider_config.json`. Letters are never generated, so all slots are read
170
- from one pass. Large label sets were sub-sampled to at most 10 options per training example (gold always kept, order shuffled),
171
- so the model conditions on the supplied candidates rather than on a fixed head.
 
 
 
 
 
172
 
173
  ## Training
174
 
@@ -201,250 +242,171 @@ schedule, only the optimizer changed), AdamW on the bf16 parameters beat AdamW w
201
  through and moves the weights further from the base model; without it, updates below the bf16 resolution round away and more of
202
  the base model's knowledge is kept.
203
 
204
- **Stage 2 (v2).** A LoRA of rank 64 (alpha 128) on the attention and MLP weights of v1, trained with cross-entropy on the slot
205
- readout for 2 epochs over 29,356 rows in v1's plain state-first layout with isolated Score levels, then merged into the bf16
206
- weights:
207
 
208
- | source | rows | content |
209
- |---|---|---|
210
- | generated decision families | 8,000 | ten families (temporal and numeric decisions, subtle answer judgment, long policies, multi-hop lookup, abstention, probability, constrained trade-offs, safety judgment, paraphrase sensitivity, adversarial traps); the answers are computed by the generating code |
211
- | questions over business documents, written by Qwen3.6-27B with thinking on | 11,356 | in two rounds, one realistic business document (up to 34 domains and 22 document kinds) plus three or four typed questions per writer call; each question was answered twice more by the same model in fresh contexts, with shuffled options and without the writer's answer, and kept only when both answers agreed with the writer's (89% and 91% kept) |
212
- | human-labelled public sets (training halves) | 3,300 | MMLU, ARC, CommonsenseQA, BoolQ, MNLI, SNLI, Banking77, RACE, OpenBookQA, LogiQA 2, MedQA, Winogrande |
213
- | replay of mixture v2 | 6,700 | 100 rows from the training half of each of the 67 in-task regression tasks |
 
 
 
214
 
215
  | stage 2 | |
216
  |---|---|
217
  | trainable parameters | LoRA rank 64, alpha 128, on the attention and MLP projections; merged after training |
218
- | schedule | learning rate 1e-4, 5% warm-up then cosine, 1,517 steps of 65,536 tokens (2 epochs), seed 0 |
219
- | hardware | one NVIDIA B300, 87 minutes |
 
220
 
221
  No JevBench item and no Decision Index item was used for training, for writing the generators or the document questions, for
222
- selecting the checkpoint, or for the temperature. The ten generated families and the skill list of the document questions were
223
  written from the family names that JevBench publishes for its sealed set, not from its items. Every training row was checked
224
- against every evaluation file used for selection and against the evaluation half of all 95 regression tasks: there is no exact
225
- state-and-question overlap. The held-out sets used for selection are generated families from held-out templates and document
226
- questions from business domains that are not in the training data.
227
-
228
- **Temperature.** 1.935, fitted by NLL (all rows pooled) on the in-task half of the public regression set without Banking77,
229
- CLINC-OOS, MMLU, ARC, Winogrande and HellaSwag: 61 tasks, 102,804 rows. Fitted on all 67 in-task tasks the value is 1.942. v1's
230
- temperature (1.05) was fitted on the 67 in-task tasks, the same set and method as decider-2b and decider-35b-a3b. No
231
- reinforcement-learning stage was run on this model; the RL recipe of decider-2b v10 is documented in `docs/RL.md` of the GitHub
232
- repository.
 
 
 
 
233
 
234
  ## Evaluation
235
 
236
- All v2 numbers are at the stored temperature 1.935 through decider-ai 1.2.1, except where a paragraph says otherwise. Greedy
237
- play and accuracy do not depend on the temperature; sampled play and calibration do, and those were measured again at 1.935.
238
 
239
  **Public regression set**, rebuilt on this machine (95 tasks: 67 in-task, 28 held-out; large label sets sub-sampled to 10
240
- options; one temperature per model fitted on in-task data). All four rows are the same rows. ECE is the expected calibration
241
- error with 15 bins; intervals are 95% bootstrap intervals over tasks, paired.
242
 
243
  | model | in-task acc / NLL / ECE (67 tasks) | held-out acc / NLL / ECE (28 tasks) |
244
  |---|---|---|
245
- | decider-2b v10, T=1.30 | 0.805 / 0.474 / 0.037 | 0.755 / 0.622 / 0.084 |
 
246
  | decider-4b v1, T=1.05 | 0.834 / 0.404 / 0.027 | 0.788 / 0.558 / 0.071 |
247
- | **decider-4b v2 (this repository), T=1.935** | **0.824 / 0.441 / 0.041** | **0.779 / 0.566 / 0.080** |
 
248
  | decider-35b-a3b v1, T=1.08 | 0.855 / 0.357 / 0.026 | 0.810 / 0.497 / 0.069 |
249
 
250
- Against v10: +1.9 in-task points (+1.0 to +2.8) and +2.4 held-out points (+1.1 to +3.7); NLL −0.032 (−0.054 to −0.011) and
251
- −0.056 (−0.094 to −0.024); accuracy higher on 66 of the 95 tasks, equal on two, lower on 27 (hate-speech tweets −4.9 points,
252
- emotion −3.6, ADE −2.9, abstention probe −2.3, SST-5 −2.1, HelpSteer3 preference −2.1, LIAR −2.1, the others under 2). The
253
- largest gains are on knowledge and reasoning tasks: MedQA +13.4, MedMCQA +12.3, Winogrande +11.7, TruthfulQA +11.6, MMLU +11.3,
254
- OpenBookQA +8.8, StrategyQA +8.0. Mean accuracy over nine knowledge tasks (MMLU, ARC, HellaSwag, MedQA, MedMCQA, Social IQa,
255
- COPA, TruthfulQA, TREC): 0.797 (v1 0.801) against v10's 0.707 and the 35B's 0.862. Against the 35B: −3.1 in-task (−4.0 to −2.3)
256
- and −3.2 held-out points (−4.5 to −1.8), lower on 85 of 95 tasks and higher on 7; the largest gaps are MedQA −17.7, HelpSteer3
257
- preference −16.6, MedMCQA −11.9, StrategyQA −10.8, TruthfulQA −10.0.
258
 
259
- <details>
260
- <summary><b>Per-task accuracy / ECE on the 28 held-out datasets, decider-2b v10, this model, decider-35b-a3b</b></summary>
 
261
 
262
- | task | decider-2b v10 | decider-4b v2 | decider-35b-a3b |
263
- |---|---|---|---|
264
- | abstain_probe | 0.606 / 0.134 | 0.582 / 0.133 | 0.622 / 0.085 |
265
- | ade | 0.817 / 0.038 | 0.789 / 0.107 | 0.837 / 0.035 |
266
- | arena_pref | 0.483 / 0.189 | 0.503 / 0.218 | 0.521 / 0.121 |
267
- | bbc_news | 0.927 / 0.013 | 0.951 / 0.068 | 0.944 / 0.027 |
268
- | cb | 0.857 / 0.093 | 0.839 / 0.118 | 0.893 / 0.084 |
269
- | cr_reviews | 0.903 / 0.031 | 0.898 / 0.046 | 0.914 / 0.033 |
270
- | dbpedia_l2 | 0.950 / 0.018 | 0.951 / 0.017 | 0.961 / 0.010 |
271
- | dbpedia_l3 | 0.987 / 0.005 | 0.991 / 0.012 | 0.992 / 0.004 |
272
- | dolly_category | 0.299 / 0.203 | 0.374 / 0.142 | 0.354 / 0.098 |
273
- | fin_phrasebank | 0.694 / 0.042 | 0.699 / 0.077 | 0.759 / 0.110 |
274
- | fin_sentiment | 0.793 / 0.058 | 0.826 / 0.051 | 0.839 / 0.136 |
275
- | hermes_tools | 0.723 / 0.208 | 0.737 / 0.124 | 0.799 / 0.085 |
276
- | hwu64 | 0.961 / 0.030 | 0.963 / 0.037 | 0.975 / 0.022 |
277
- | massive_scenario | 0.756 / 0.041 | 0.793 / 0.052 | 0.799 / 0.027 |
278
- | offtopic_probe | 0.841 / 0.027 | 0.823 / 0.043 | 0.870 / 0.038 |
279
- | paws | 0.724 / 0.145 | 0.794 / 0.124 | 0.729 / 0.169 |
280
- | pubmedqa | 0.756 / 0.085 | 0.758 / 0.055 | 0.820 / 0.078 |
281
- | quality | 0.494 / 0.233 | 0.561 / 0.179 | 0.632 / 0.096 |
282
- | quality_full | 0.508 / 0.198 | 0.505 / 0.188 | 0.565 / 0.112 |
283
- | reward_bench | 0.819 / 0.045 | 0.853 / 0.032 | 0.919 / 0.024 |
284
- | sciq | 0.982 / 0.024 | 0.991 / 0.021 | 0.993 / 0.011 |
285
- | social_iqa | 0.708 / 0.077 | 0.759 / 0.062 | 0.823 / 0.025 |
286
- | strategyqa | 0.552 / 0.138 | 0.632 / 0.157 | 0.739 / 0.036 |
287
- | student_questions | 0.925 / 0.045 | 0.940 / 0.052 | 0.954 / 0.090 |
288
- | trec | 0.784 / 0.066 | 0.842 / 0.027 | 0.832 / 0.160 |
289
- | truthfulqa | 0.537 / 0.090 | 0.654 / 0.063 | 0.754 / 0.068 |
290
- | tweet_irony | 0.795 / 0.052 | 0.827 / 0.027 | 0.861 / 0.129 |
291
- | xstory_cloze | 0.962 / 0.017 | 0.971 / 0.017 | 0.995 / 0.016 |
292
-
293
- </details>
294
-
295
- **On the same rows as decider-2b v10 and decider-35b-a3b.** Every row below is scored by all three models on identical inputs
296
- and seeds. Intervals are 95% paired bootstrap intervals (rows for the fixtures, boards for the games, task-seed pairs for the
297
- browser). The table under [Changes from v1](#changes-from-v1) has v1 on the same rows.
298
-
299
- | | decider-2b v10 | **decider-4b v2** | decider-35b-a3b | 4B minus 2B | 4B minus 35B |
300
- |---|---|---|---|---|---|
301
- | 847 in-task validation rows, accuracy / NLL | 83.2% / 0.444 | 85.0% / 0.419 | 90.0% / 0.329 | +1.8 (−0.5 to +3.9) | −5.0 (−7.1 to −2.8) |
302
- | OpenJev, 5,252 rows, accuracy / NLL | 63.3% / 0.916 | 66.7% / 0.789 | 68.3% / 0.752 | +3.5 (+2.2 to +4.7) | −1.6 (−2.8 to −0.3) |
303
- | Mind2Web element and action choice, 1,770 rows | 82.7% / 0.543 | 87.4% / 0.393 | 89.6% / 0.316 | +4.8 (+3.1 to +6.4) | −2.2 (−3.7 to −0.6) |
304
- | TypeSafe workflow decisions, 102 rows, accuracy / NLL | 80.4% / 0.585 | 86.3% / 0.410 | 86.3% / 0.342 | +5.9 (−2.9 to +14.7) | 0.0 (−5.9 to +4.9) |
305
- | Bespoke's public suite, 13 subsets, macro / micro | 0.704 / 0.711 | 0.773 / 0.781 | 0.774 / 0.787 | | |
306
- | JevBench public items, easy / standard / hard accuracy | 1.000 / 0.889 / 0.459 | 1.000 / 0.986 / 0.676 | 1.000 / 0.972 / 0.676 | | |
307
- | live MiniWoB++ click tasks, 22 tasks x 8 seeds, greedy play | 90.9% | 92.6% | 97.2% | +1.7 (−4.0 to +7.4) | −4.5 (−8.0 to −1.7) |
308
- | the same, 6 tasks v10 never used for reward, greedy | 91.7% | 81.2% | 97.9% | −10.4 (−22.9 to +2.1) | −16.7 (−27.1 to −6.2) |
309
- | live MiniWoB++ click tasks, sampled play | 93.2% | 88.1% | 86.4% | −5.1 (−10.8 to +0.6) | +1.7 (−3.4 to +6.8) |
310
- | the same, 6 held-out tasks, sampled | 91.7% | 75.0% | 79.2% | −16.7 (−31.2 to −4.2) | −4.2 (−16.7 to +8.3) |
311
- | zero-shot games, win rate, sampled play (234 boards) | 23.7% | 22.4% | 24.1% | −1.3 (−4.2 to +1.5) | −1.7 (−3.7 to +0.2) |
312
- | zero-shot games, greedy play | 26.5% | 27.8% | 37.2% | +1.3 (−4.3 to +6.8) | −9.4 (−15.0 to −3.8) |
313
- | bag-draw games alone, sampled | 41.4% | 37.9% | 41.8% | −3.5 (−9.0 to +2.0) | −3.9 (−8.6 to +0.8) |
314
-
315
- On the two fixtures whose rows are in neither model's training data (TypeSafe, OpenJev) v2 is above the 2B (OpenJev +3.5, interval
316
- excludes zero; TypeSafe +5.9, interval includes zero) and on TypeSafe level with the 35B. Stage 2 contains teacher-written and
317
- human-labelled judgment rows, not TypeSafe or OpenJev rows. The browser rows show the missing RL stage: greedy play is level
318
- with v10 over all tasks and 10 points below it on the six tasks that v10's RL never rewarded (interval includes zero), where the
319
- 4B's weakest tasks are focus-text-2 (0.375) and click-collapsible-2 (0.5). On the bag-draw games v1 won about 15 points more
320
- often than either other model in sampled play; v2 does not (37.9% against 41.4% and 41.8%). On tic-tac-toe, the slippery grid
321
- and minesweeper every model is near the random floor.
322
-
323
- **Ten text games, zero-shot** (`decider.games.play`, five episodes per game, greedy, eager; the 2B was trained on the first
324
- four, the 4B and the 35B on none): Pong −21 (teacher 8, 2B 8, 35B −21), Breakout 12 (22, 22, 6), CliffWalking −60 (teacher −13,
325
- 2B −13, 35B −1,248), MiniGrid-Empty 0 (teacher 0.96, both others 0), Freeway 1 (teacher 5, 2B 0, 35B 1), FrozenLake 0 (teacher 1,
326
- others 0), Blackjack −0.6 (teacher −0.6, 2B −1, 35B −0.6), MiniGrid-LavaGap 0, MiniGrid-DoorKey 0 (teacher 0), BabyAI-GoTo 0.35
327
- (teacher 0.34, 2B 0, 35B 0.19). v1 reached the teacher on CliffWalking (−13) and read 0.54 on BabyAI-GoTo; v2 keeps neither.
328
- Neither version learns Pong from the text state. Greedy play takes the most probable option, so these results do not depend on
329
- the temperature.
330
 
331
  **JevBench public items** (231 items of [Benchmark Heaven](https://benchmarkheaven.com/jev-models); argmax over the exact
332
- label set with the request the harness's TypeSafe adapter builds). The items were read once for v2, after selection, at the
333
- candidate temperature 1.719; the accuracies do not depend on the temperature, and the calibration numbers below are recomputed
334
- at 1.935 from the stored probabilities. Jev 1.13.0 is at 1.000 / 0.986 / 0.730, SemIf (also a Qwen3.5-4B) at 1.000 / 0.986 /
335
- 0.613, decider-35b-a3b at 1.000 / 0.972 / 0.676 on the same items. The one standard-tier family with misses is adequacy (0.92).
336
- Hard-tier families (v1 as first published in brackets): trap 1.00 (1.00), hard routing 1.00 (1.00), adversarial 1.00 (0.67), probability 0.90 (0.60),
337
- ambiguous 0.86 (0.29), multi-hop 0.72 (0.78), trade-off 0.67 (0.33), long policy 0.63 (0.47), judge-hard 0.47 (0.47),
338
- temporal-numeric 0.27 (0.13). Top-label ECE is 0.008 / 0.085 / 0.071 by tier (v1, read again on 2026-09-24: 0.001 / 0.047 / 0.288;
339
- the 35B's hard tier 0.151). On the hard tier the mean confidence is 0.74 at accuracy 0.68; 31% of hard items are answered at confidence 0.9 or more,
340
- with accuracy 0.94 on them (v1: 50% at 0.73). The 18 Score items were answered with isolated levels, whose per-level values are
341
- not stored, so they keep their T 1.719 probabilities inside these ECE values. The public hard tier is 111 items (95% interval
342
- about ±9 points), and the gain over v1 there (+12.6) is larger than the gain on our own held-out hard sets.
343
 
344
  **Bespoke's public suite** (13 human-labelled subsets, 3,880 records in Jev's wire format, answered through `system_one` as
345
- shipped). Nimble-9B and Jev 1.13.0 numbers are copied from Bespoke's report.
 
346
 
347
- | subset (type) | decider-2b v10 | decider-4b v2 | decider-35b-a3b | Nimble-9B | Jev 1.13.0 |
348
- |---|---|---|---|---|---|
349
- | vitaminc-dev (choice) | 0.639 | 0.778 | 0.795 | 0.766 | 0.801 |
350
- | massive-en-US (choice; trained) | 0.823 | 0.869 | 0.880 | 0.869 | 0.874 |
351
- | massive-de-DE (choice, German) | 0.797 | 0.837 | 0.869 | 0.834 | 0.869 |
352
- | boolq (noul; trained) | 0.803 | 0.860 | 0.887 | 0.860 | 0.897 |
353
- | squad2 (noul) | 0.776 | 0.793 | 0.749 | 0.806 | 0.829 |
354
- | paws (noul; trained) | 0.720 | 0.832 | 0.768 | 0.828 | 0.892 |
355
- | multinli (choice; trained) | 0.856 | 0.926 | 0.910 | 0.853 | 0.829 |
356
- | civil_comments (noul; trained) | 0.840 | 0.857 | 0.907 | 0.703 | 0.810 |
357
- | aegis2 (noul) | 0.728 | 0.812 | 0.808 | 0.812 | 0.804 |
358
- | helpsteer2 (score; trained) | 0.426 | 0.466 | 0.478 | 0.390 | 0.341 |
359
- | summeval-relevance (score) | 0.354 | 0.463 | 0.483 | 0.492 | 0.350 |
360
- | summeval-consistency (score) | 0.660 | 0.826 | 0.757 | 0.757 | 0.812 |
361
- | pubmedqa (choice; trained) | 0.724 | 0.724 | 0.768 | 0.756 | 0.772 |
362
- | **macro / micro** | 0.704 / 0.711 | **0.773 / 0.781** | 0.774 / 0.787 | 0.748 / 0.759 | 0.760 / 0.773 |
363
-
364
- On the six subsets whose training split is not in the mixture the macro accuracy is 0.751 (v1 0.731, 2B 0.659, 35B 0.744). v2 is
365
- 1.3 points above Jev 1.13.0 and 2.5 above Nimble-9B on the average (v1: 0.757). It is still behind Jev where a claim has to be
366
- checked against evidence that nearly matches it (PAWS 0.832 against 0.892, VitaminC 0.778 against 0.801) and on SQuAD2
367
- answerability (0.793 against 0.829), where it is now 1.7 points above the 2B (v1 was 7 points under it).
368
-
369
- **Behaviour probes** (teacher-labelled, same probes as the other releases; v1 in brackets): generic-versus-specific bucket choice
370
- 0.95 / 1.00 (1.00 / 1.00), catch-all when nothing fits 0.95 (0.90; 35B 0.95), abstention battery 8 of 8 (7 of 8); model-router
371
- tier 0.935 (0.968) and needs-live-data 0.839 (0.871); command-risk classification 0.911 (0.889) with no destructive command
372
- called safe (35B 0.933), touches-outside-project 0.933 (0.956); browser-agent element and action choice 0.938 / 0.938 (0.875 /
373
- 0.875; 35B 0.938). Scoring a Score level alone against scoring it with its neighbours changes accuracy by at most 2.0 points on
374
- five rating datasets, and the per-level fits sum to between 0.92 and 1.00 (this probe reads the raw logits, T=1).
375
 
376
  ## Calibration
377
 
378
- The stored temperature 1.935 was fitted on 61 of the 67 in-task regression tasks (see Training). On the regression set v2 is less
379
- well calibrated than v1 and the 35B (ECE 0.041 / 0.080 against v1's 0.027 / 0.071 and the 35B's 0.026 / 0.069), with per-task
380
- exceptions: hate-speech tweets 0.349 (in-task), Arena preferences 0.218, QuALITY (full) 0.188, QuALITY 0.179, StrategyQA 0.157,
381
- Dolly categories 0.142, abstention probe 0.133; the in-task preference tasks HH-RLHF (0.113) and SHP (0.111) are also above 0.1.
382
- On the five rating datasets of the Score-level probe, which reads the raw logits (T=1) and scores all levels together, the ECE
383
- is 0.13 to 0.27 (v1 0.05 to 0.09 in the same probe); these sets were not measured at the served temperatures. One temperature
384
- does not fit every task type. Outside the regression set v2 is better calibrated than v1 on most sets:
385
 
386
- | set | measure | decider-2b v10 | decider-4b v1 (read again 2026-09-24) | decider-4b v2 | decider-35b-a3b |
387
- |---|---|---|---|---|---|
388
- | JevBench hard tier, 111 items | top-label ECE | 0.304 | 0.288 | 0.071 | 0.151 |
389
- | TypeSafe, 102 rows | 10-bin ECE / mean total variation to the frontier reference | 0.091 / 0.242 | 0.124 / 0.232 | 0.072 / 0.196 | 0.065 / 0.165 |
390
- | OpenJev, 5,252 rows | 10-bin ECE | 0.150 | 0.158 | 0.131 | 0.049 |
391
- | 847 in-task validation rows | 10-bin ECE | | 0.037 | 0.031 | |
392
- | Mind2Web, 1,770 rows | 10-bin ECE | | 0.016 | 0.051 | |
393
- | Decision Index 4,000-request sample, 33 benchmarks | benchmark-weighted ECE (the site's definition) | 0.093 | 0.086 | 0.074 | 0.027 |
394
- | the same | share of all answers given at 95% or more confidence and wrong | 0.010 | 0.022 | 0.012 | 0.005 |
395
- | the same | mean confidence against accuracy | 0.658 / 0.566 | 0.722 / 0.637 | 0.690 / 0.637 | 0.684 / 0.690 |
396
- | the same | sample index (about 4 points under a full run) | 42.3 | 48.5 | 48.3 | 50.4 |
397
-
398
- The Decision Index rows are a readout of the 4,000-request sample (about 88 cases per benchmark) through the public server,
399
- computed with the site's definition (benchmark-weighted pooled bins, ten bins; the same computation gives 0.027 on the 35B
400
- against the site's published 0.031). v2 was read once at the candidate temperature 1.719 (ECE 0.090, sample index 48.3); the
401
- chosen answers do not change with the temperature, and the calibration rows above are recomputed at 1.935 from the stored
402
- probabilities. The JevBench hard-tier value is recomputed the same way, except its 6 Score items, which keep their T 1.719
403
- probabilities. The sample index was not recomputed. Nothing was fitted on index rows. v2's confidence exceeds its accuracy by
404
- 0.053 on that sample (v1 0.085, the 2B 0.092; the 35B has no gap). The gap is per-benchmark heterogeneity, not a global scale: a
405
- single temperature that removes it on the knowledge benchmarks would make the wide label sets underconfident. If you route on
406
- confidence, calibrate on your own labels.
407
 
408
  ## Speed
409
 
410
- v2 has the same architecture and size as v1, so it runs at v1's speed. These timings were taken with the same weights in the
411
- candidate readout session of 2026-09-24 (the server was configured with the candidate temperature 1.719; the temperature does
412
- not change the amount of computation). In one session on one unshared NVIDIA B300 (bf16), both models measured one after the
413
- other: 34.6 ms (v1 32.4 ms) median per decision over 200 game-state decisions of 156 tokens median,
414
- batch of one, eager PyTorch without CUDA graphs or `torch.compile`, timed around the forward pass with `torch.cuda.synchronize()`.
415
- The eager path is launch-bound, so host load changes it: v1's first measurement was 24.7 ms (10th to 90th percentile 24.6 to
416
- 43.5 ms), with decider-2b at 17.9 ms and decider-35b-a3b at 41.4 ms on the same decisions and method. With the helper's CUDA
417
- graphs and `torch.compile`, one support-ticket request (228 tokens, 3 questions) takes 5.2 ms (FP8 5.0 ms); a batch of 32 such
418
- states takes 81.5 ms, 1,178 decisions per second (FP8 71.2 ms, 1,349 per second). The HTTP server (`/decide`, bf16) answers 72.8
419
- requests per second at a median of 13.3 ms with one client and 190 requests per second with 64 clients.
420
 
421
  ## Limitations
422
 
423
- * No reinforcement-learning stage: stated beliefs about action outcomes were not trained against exact laws, and on live
424
- browser tasks the model is 10 points below decider-2b v10 on the six held-out tasks in greedy play (81% against 92%, interval
425
- includes zero) and 17 points below it in sampled play (75% against 92%).
426
- * Sampled play is worse than v1's: bag-draw games 37.9% against 56.6% wins, zero-shot games 22.4% against 27.8%, sampled browser
427
- play −2.8 points; CliffWalking −60 against −13. Pin revision `v1` if you depend on these (see Changes from v1).
428
- * Less calibrated than v1 on the regression set (in-task ECE 0.041 against 0.027), on Mind2Web (0.051 against 0.016) and on
429
- rating tasks in the Score-level probe (ECE 0.13 to 0.27 on the five rating sets at T=1); still overconfident outside the regression set: OpenJev ECE 0.13,
430
- Decision Index sample ECE 0.074 where the 35B reads 0.027. On the Decision Index sample 1.2% of all answers are given at 95% or
431
- more confidence and are wrong, against 0.5% for the 35B. Calibration is measured on public datasets, not on your traffic.
432
- * About 1 point under v1 on everyday tasks (regression set, validation rows, Mind2Web) and lower than v1 on the model-router and
433
- command-scope probes.
434
- * Below the 35B on 85 of 95 regression tasks by 3.1 / 3.2 points and on three fixtures by 1.6 to 5.0 points; level on TypeSafe
435
- and on the JevBench hard tier. The largest gaps are on MedQA, MedMCQA, StrategyQA, TruthfulQA and preference judgments.
436
- * The abstention probe is 2.4 points under the 2B (0.582 against 0.606), with ECE 0.133.
437
- * The JevBench public hard tier is 111 items; v2's gain there (+12.6 points over v1) is larger than its gain on our own held-out
438
- hard sets, where it tied an earlier candidate. Do not read it as a gain of that size on hard items in general.
439
  * The stage-2 LoRA was trained only in the plain state-first layout. The schema-first layout (the server's opt-in schema cache)
440
- was not measured on v2, and `decider_config.json` does not mark v2 as trained for it (`schema_first_trained: false`), so
441
- `DECIDER_SCHEMA_CACHE=1` does not turn the schema cache on for this model.
442
  * Mixture v2's 26 additional public datasets and ten programmatic families, and stage 2's generators and document questions, are
443
  described above but their builders are not in the public package; `scripts/train.sh full` reproduces the 60% of stage 1's data
444
- that is the public mixture. On the held-out variants of mixture v2's programmatic families v1 was at 0.635 mean accuracy (logs
445
- 0.28, plans 0.44, probability 0.43, code 0.88, policy 0.99); this was not measured on v2.
446
  * English is the main language; the multilingual rows (XNLI, PAWS-X, MASSIVE, Belebele, XCOPA) are a small share of the data
447
- and were not measured beyond the mixture-v2 evaluation set.
448
  * Everything else in the decider-2b card's limitations (packed questions see each other, long JSON arrays by position, full
449
  label sets against sampled options, abstention wording, rules in the question) applies; those shapes were not re-measured at
450
  this size.
@@ -453,7 +415,8 @@ requests per second at a median of 13.3 ms with one client and 190 requests per
453
 
454
  | version | what changed |
455
  |---|---|
456
- | **v2** (2026-09-24, these weights) | v1 + a merged LoRA (rank 64, attention and MLP, 2 epochs, 29,356 rows: generated decision families, document questions written by Qwen3.6-27B and kept when two independent answers agreed, human-labelled public sets, replay of mixture v2); temperature 1.935; plain layout, checked with decider-ai 1.0.2, 1.1.4, 1.2.1 and 1.2.2. Better on hard judgments, TypeSafe, OpenJev and Bespoke's suite; worse on sampled play, some text games and by about 1 point on everyday tasks |
 
457
  | v1 (2026-09-22, Hub tag `v1`) | first release: one pass over mixture v2 on Qwen3.5-4B-Base with AdamW on bf16 parameters, no RL stage; temperature 1.05 |
458
 
459
  The GitHub repository's [docs/CHANGELOG.md](https://github.com/Mapika/decider/blob/main/docs/CHANGELOG.md) lists every
@@ -464,10 +427,10 @@ decider release.
464
  Code, data registry, training and evaluation scripts and the per-version history: https://github.com/Mapika/decider
465
  (`docs/HISTORY.md`, section "decider-4b"). Stage 1 was trained with the data-parallel trainer of the architecture A/B study
466
  (`arch_ab/train_dp_optvar.py` in the research repository, optimizer variant `bf16`); stage 2 with a LoRA trainer in the research
467
- repository. Both were evaluated with the public `decider.evaluate` and the head-to-head tools, and uploaded with
468
- `scripts/upload_hf.py`. `eval_results.json` in this repository has the per-task regression metrics at T 1.935, the fixtures, JevBench, Bespoke's
469
- suite, the games, the browser, the text games, the Decision Index calibration, the behaviour probes and speed, and the paired
470
- comparisons with v1 (regression set, fixtures, JevBench, games, browser).
471
 
472
  **Independence.** This is an independent project. It is not affiliated with or endorsed by TypeSafe AI. It is an open
473
  reproduction of the "System One" model class (TypeSafe AI's Jev); nothing was distilled from Jev. The training data is public
 
17
  attention and 24 with gated delta-net linear attention, hidden size 2,560. Two supervised stages. Stage 1 (decider-4b v1): one
18
  pass of cross-entropy on the slot readout over mixture v2, the public decision mixture of
19
  [decider-2b](https://huggingface.co/Mapika/decider-2b) plus 26 further public decision datasets and ten programmatically
20
+ generated families with verifiable gold (742M tokens). Stage 2 (v2.1): a LoRA of rank 64 on the attention and MLP weights,
21
+ trained for 2 epochs on 29,325 rows of harder decisions and replay, with the replay rows trained toward v1's own answer
22
+ distribution, then merged into the weights. There is no reinforcement-learning stage. **This repository holds v2.1**, the bf16
23
+ weights (8.4 GB), with one temperature per answer type. v2 stays available under the Hub tag `v2` and v1 under the tag `v1`
24
+ (see [Changes from v2](#changes-from-v2) for who should keep using them). The other sizes are listed under The decider family.
25
+ `decider/` in this repository is the inference subset of the GitHub package.
26
+
27
+ v2.1 keeps most of v2's gain on hard decisions and gets back most of what v2 lost in sampled play. Against v2 on the same rows:
28
+ bag-draw games in sampled play 52.0% against 37.9% wins (v1 56.6%), zero-shot games sampled 26.9% against 22.4% (v1 27.8%), live
29
+ browser tasks sampled 93.2% against 88.1% (v1 90.9%), CliffWalking −13 against −60, the regression set 0.831 / 0.784 against 0.824
30
+ / 0.779 (v1 0.834 / 0.788); held-out generated decision families 0.556 against 0.560 (v1 0.469) and the JevBench public hard tier
31
+ 0.649 against 0.676 (v1 0.550). It is less well calibrated than v2 on hard items (details under
32
+ [Calibration](#calibration)), still answers one of the two form-filling cases of issue #9 wrongly, and is worse than v1 on
33
+ BabyAI-GoTo and on greedy bag-draw play. Against decider-35b-a3b it is 2.4 and 2.6 points lower on the regression set.
34
+
35
+ **Contents:** [Changes from v2](#changes-from-v2) · [The decider family](#the-decider-family) · [Usage](#usage) · [How it works](#how-it-works) · [Training](#training) · [Evaluation](#evaluation) · [Calibration](#calibration) · [Speed](#speed) · [Limitations](#limitations) · [Changelog](#changelog) · [Reproduction](#reproduction)
36
+
37
+ ## Changes from v2
38
+
39
+ v2.1 differs from v2 in two things:
40
+
41
+ * **Replay rows are trained toward v1's own distribution.** Stage 2 of v2 trained every row, including 6,700 replay rows from v1's
42
+ training mixture, on its hard label. v2.1 trains the replay rows toward v1's answer distribution instead, with the loss
43
+ KL(p_v1 ‖ p_model) over the options of each answer, and keeps hard labels on every other row. Everything else is v2's recipe:
44
+ the same rows (29,325 after removing 31 rows that shared a context with an evaluation file; v2 had 29,356), LoRA rank 64, learning
45
+ rate 1e-4, 2 epochs. Training the replay rows on hard labels had sharpened the logits everywhere; the fitted temperature then
46
+ rose to 1.935 and flattened every served answer, which is what cost v2 its sampled play. With the replay rows trained toward
47
+ v1, the fitted temperature is 1.099 (v1 1.05). On 1,000 held-out replay rows the mean KL(p_v1 ‖ p_model) at temperature 1 is 0.021
48
+ nats (v2: 0.112), and 96.3% of the argmax answers equal v1's (v2: 94.1%).
49
+ * **One temperature per answer type.** `decider_config.json` has `temperature` 1.099 and `temperature_by_type`
50
+ `{"choice": 1.110, "noul": 1.560, "score": 1.287}`, fitted by NLL with `decider.calibrate` (decider-ai 1.4.0). decider-ai 1.4.0
51
+ and later use the map. **decider-ai 1.3.0 and earlier ignore the map and serve every answer at 1.099**; a temperature does not
52
+ change which option is most probable, so the answers are the same either way (except exact ties between the level rows of an
53
+ isolated Score answer, where rounding can break the tie differently: 1 of 5,000 held-out generated-family rows and 2 of 447
54
+ document-question validation rows changed); the
55
+ probabilities differ, mostly on yes/no and Score answers. The map changes little in practice: sampled
56
+ play, the fixtures and the probes are the same within noise with and without it, and it lowers the calibration error on
57
+ held-out yes/no answers (0.146 to 0.101) and Score answers (0.140 to 0.108). It does not fix the overconfidence on hard
58
+ Choice answers (see Calibration).
59
+
60
+ All rows below are on identical inputs and seeds: v2.1 through decider-ai 1.4.0 with its map, v1 and v2 through decider-ai 1.3.0
61
+ at their stored temperatures (1.05 and 1.935), measured in the same session on 2026-09-24. Intervals are 95% paired bootstrap
62
+ intervals (rows for the fixtures, boards for the games, task-seed pairs for the browser). The regression set and the two held-out
63
+ sets were read from stored temperature-1 logits at each model's served temperatures. The JevBench files of v1 and v2 were read
64
+ earlier through decider-ai 1.2.1, v2's at the candidate temperature 1.719; accuracy does not depend on the temperature.
65
+
66
+ | set | v1 (T 1.05) | v2 (T 1.935) | v2.1 (map) | v2.1 minus v1 | v2.1 minus v2 |
67
+ |---|---|---|---|---|---|
68
+ | regression set, 67 in-task tasks, accuracy / NLL / ECE | 0.834 / 0.404 / 0.027 | 0.824 / 0.441 / 0.041 | 0.831 / 0.414 / 0.031 | −0.3 | +0.7 |
69
+ | regression set, 28 held-out tasks | 0.788 / 0.558 / 0.071 | 0.779 / 0.566 / 0.080 | 0.784 / 0.569 / 0.077 | −0.4 | +0.5 |
70
+ | 847 in-task validation rows, accuracy / NLL | 86.1% / 0.417 | 85.0% / 0.419 | 85.0% / 0.432 | −1.1 (−2.1 to +0.0); NLL +0.015 (+0.003 to +0.028) | 0.0 (−1.5 to +1.5); NLL +0.014 (−0.006 to +0.035) |
71
+ | OpenJev, 5,252 rows, accuracy / NLL | 63.9% / 0.893 | 66.7% / 0.789 | 66.0% / 0.846 | +2.1 (+1.2 to +2.9); NLL −0.047 (−0.060 to −0.035) | −0.7 (−1.6 to +0.2); NLL +0.057 (+0.044 to +0.070) |
72
+ | Mind2Web, 1,770 rows, accuracy / NLL | 88.4% / 0.366 | 87.4% / 0.393 | 87.5% / 0.375 | −0.8 (−1.7 to +0.0); NLL +0.009 (−0.004 to +0.023) | +0.1 (−0.8 to +1.1); NLL −0.018 (−0.032 to −0.005) |
73
+ | TypeSafe workflow decisions, 102 rows, accuracy / NLL | 81.4% / 0.611 | 86.3% / 0.410 | 84.3% / 0.441 | +2.9 (−2.9 to +8.8); NLL −0.170 (−0.324 to −0.036) | −2.0 (−7.8 to +3.9); NLL +0.031 (−0.069 to +0.134) |
74
+ | held-out generated families (heldout_jb), 5,000 rows, accuracy / ECE | 0.469 / 0.260 | 0.560 / 0.046 | 0.556 / 0.147 | +8.7 | −0.4 |
75
+ | held-out document questions (test_teacher2), 449 rows, accuracy / ECE | 0.726 / 0.110 | 0.826 / 0.053 | 0.820 / 0.043 | +9.4 | −0.7 |
76
+ | JevBench public items, easy / standard / hard accuracy | 1.000 / 0.958 / 0.550 | 1.000 / 0.986 / 0.676 | 1.000 / 0.986 / 0.649 | hard +11 items | hard −3 items |
77
+ | JevBench hard tier, top-label ECE (v1, v2: files read through 1.2.1, v2 at T 1.719) | 0.288 | 0.104 | 0.184 | | |
78
+ | Bespoke's public suite, macro / micro | 0.757 / 0.765 | 0.773 / 0.781 | 0.756 / 0.765 | −0.1 macro | −1.6 macro |
79
+ | live MiniWoB++, sampled, all 22 tasks | 90.9% | 88.1% | 93.2% | +2.3 (−1.7 to +6.2) | +5.1 (+0.6 to +9.7) |
80
+ | live MiniWoB++, sampled, 16 rewarded tasks | 96.1% | 93.0% | 95.3% | −0.8 (−3.9 to +2.3) | +2.3 (−2.3 to +7.0) |
81
+ | live MiniWoB++, sampled, 6 held-out tasks | 77.1% | 75.0% | 87.5% | +10.4 (0.0 to +20.8) | +12.5 (0.0 to +25.0) |
82
+ | live MiniWoB++, greedy, all 22 tasks | 91.5% | 92.6% | 93.8% | +2.3 (−0.6 to +5.7) | +1.1 (−1.7 to +4.5) |
83
+ | live MiniWoB++, greedy, 16 rewarded tasks | 97.7% | 96.9% | 96.1% | −1.6 (−3.9 to +0.0) | −0.8 (−2.3 to +0.0) |
84
+ | live MiniWoB++, greedy, 6 held-out tasks | 75.0% | 81.2% | 87.5% | +12.5 (+4.2 to +22.9) | +6.2 (−4.2 to +16.7) |
85
+ | zero-shot games, 234 boards, sampled, win rate | 27.8% | 22.4% | 26.9% | −0.9 (−2.6 to +1.0) | +4.5 (+2.5 to +6.6) |
86
+ | bag-draw games, 64 boards, sampled, win rate | 56.6% | 37.9% | 52.0% | −4.7 (−9.0 to −0.4) | +14.1 (+8.2 to +19.9) |
87
+ | slippery-grid games, 64 boards, sampled, win rate | 16.4% | 16.0% | 16.8% | +0.4 (−2.7 to +3.5) | +0.8 (−2.3 to +3.9) |
88
+ | zero-shot games, 234 boards, greedy, win rate | 29.1% | 27.8% | 25.2% | −3.8 (−7.3 to −0.4) | −2.6 (−6.4 to +1.3) |
89
+ | bag-draw games, 64 boards, greedy, win rate | 62.5% | 48.4% | 53.1% | −9.4 (−17.2 to −3.1) | +4.7 (−3.1 to +12.5) |
90
+ | slippery-grid games, 64 boards, greedy, win rate | 12.5% | 18.8% | 10.9% | −1.6 (−7.9 to +4.7) | −7.8 (−15.6 to +0.0) |
91
+ | ten text games, greedy: Pong / Breakout / CliffWalking / BabyAI-GoTo / Freeway / Blackjack | −21 / 14 / −13 / 0.54 / 0 / −0.6 | −21 / 12 / −60 / 0.35 / 1 / −0.6 | −5 / 35 / −13 / 0.19 / 0 / −0.6 | | |
92
+ | behaviour probes: model-router tier / needs-live-data (31 items) | 0.968 / 0.871 | 0.935 / 0.839 | 0.968 / 0.839 | 0 / −1 item | +1 / 0 items |
93
+ | behaviour probes: command risk / touches-outside-project (45 items) | 0.889 / 0.956 | 0.911 / 0.933 | 0.889 / 0.933 | 0 / −1 item | −1 / 0 items |
94
+ | behaviour probes: generic bucket / catch-all / abstention battery / browser element and action | 1.00 / 0.90 / 7 of 8 / 0.875 and 0.875 | 0.95 / 0.95 / 8 of 8 / 0.938 and 0.938 | 1.00 / 0.95 / 8 of 8 / 0.938 and 0.750 | | |
95
+ | issue #9 form cases c_1 / c_2 (probability of the gold option) | right (0.79) / right (0.98) | wrong (0.16) / right (0.28) | wrong (0.20) / right (0.41) | | |
96
 
97
  **Regressions, stated plainly.**
98
+ * BabyAI-GoTo (greedy text game): 0.19 against v1's 0.54 and v2's 0.35.
99
+ * Greedy bag-draw play: 53.1% against v1's 62.5% wins (−9.4 points, interval −17.2 to −3.1); greedy zero-shot games overall 25.2%
100
+ against 29.1% (−3.8, interval −7.3 to −0.4). Sampled bag-draw play is 4.7 points under v1 (interval −9.0 to −0.4).
101
+ * Behaviour probes: needs-live-data 0.839 against v1's 0.871 and touches-outside-project 0.933 against 0.956, one item each (as
102
+ v2); the browser-agent action choice is 12 of 16 items (0.750), 2 fewer than v1 (0.875) and 3 fewer than v2
103
+ (0.938).
104
+ * Form filling: issue #9 case c_1 is still answered wrongly. The field is "Degree earned" and the document entity is "Studied:
105
+ Associate of Arts"; v2.1 chooses "skip" at 0.76 and gives the gold entity 0.20 (v1: gold 0.79; v2: gold 0.16). Case c_2 is
106
+ answered correctly at 0.41. For form filling we suggest v1.
107
+ * Calibration on hard items: on the held-out generated families the calibration error is 0.147 against v2's 0.046 (see
108
+ Calibration), and on the JevBench public hard tier 0.184 against v2's 0.104.
109
+ * JevBench public hard tier: 72 of 111 items (0.649) against v2's 75 (0.676).
110
+ * Against v2: TypeSafe −2.0 points and OpenJev −0.7 (intervals include zero), NLL higher on both; Bespoke's suite 0.756 against
111
+ 0.773 macro (v1 0.757); greedy slippery-grid play 10.9% against 18.8%.
112
+
113
+ **How v2.1 was chosen, and why it is released although it did not pass.** The run had pre-registered release rules. v2.1 passed
114
+ the numeric items (held-out document questions 0.820 against the required 0.816, held-out generated families 0.556 against 0.550,
115
+ regression accuracy and calibration not worse than v2's, sampled bag-draw above the midpoint of v1 and v2, sampled zero-shot
116
+ games and browser play not shown to be below v1's (upper end of the 95% interval at least 0), each probe at most one item below v1) and failed the last item, which required both
117
+ issue #9 form cases to be right: c_1 is wrong. By the rule it was not recommended. A second pre-registered rule for the
118
+ temperature map required the calibration error on the held-out generated families to be at most 0.08; the map reaches 0.147
119
+ (0.163 without it), so the map did not pass either. v2.1 is released with the map on a decision made after reading the full comparison above: it is better than v2 on sampled play, CliffWalking, the model-router probe and the regression set, at v2's
120
+ level on our held-out hard sets, and its failures are listed in this section. The JevBench public items were read once for the
121
+ candidate at its global temperature and once with the map, after the rule decisions; they were not used for training,
122
+ selection or the temperatures.
123
+
124
+ **Which version to use.**
125
+ * v2.1 (this revision): the default. Sampled play (games, browser agents that sample actions), hard decisions.
126
+ * v2 (`revision="v2"`): if you rely on confidence values on hard multi-step items, where v2 is better calibrated (held-out
127
+ generated families 0.046, JevBench hard tier 0.104), or on its slightly higher JevBench hard tier and TypeSafe accuracy.
128
+ * v1 (`revision="v1"`): form filling (issue #9), BabyAI-GoTo-like grid navigation, greedy bag-draw play.
129
+
130
+ The package loads a local folder, so download the revision first:
131
 
132
  ```python
133
  from huggingface_hub import snapshot_download
134
  from decider.infer import Decider
135
+ d = Decider(snapshot_download("Mapika/decider-4b", revision="v2")) # or revision="v1"
136
  ```
137
 
138
  For the HTTP server, set `DECIDER_MODEL` to the same downloaded folder.
139
 
 
 
 
 
 
 
 
140
  ## The decider family
141
 
142
  All six repositories share one interface (`decider.infer.Decider`, `POST /v1/systemone` in TypeSafe's format) and one
 
144
 
145
  | model | base | weights | use it for | numbers |
146
  |---|---|---|---|---|
147
+ | [decider-2b](https://huggingface.co/Mapika/decider-2b) v11 | Qwen3.5-2B-Base | 3.8 GB bf16 | the default: routing, classification, judgments, browser agents; 4 ms per request with CUDA graphs on one GPU; v10 under the tag `v10` | regression set 0.802 in-task / 0.752 held-out; JevBench hard 0.577; live browser 90% sampled; Bespoke suite 0.706 |
148
+ | [decider-4b](https://huggingface.co/Mapika/decider-4b) v2.1 | Qwen3.5-4B-Base | 8.4 GB bf16 | the middle point: knowledge questions and hard judgments above the 2B in a dense 8.4 GB model; no RL stage; v2 and v1 under the tags `v2` and `v1` | 0.831 / 0.784; JevBench hard 0.649; live browser 93% sampled; Bespoke 0.756 |
149
  | [decider-35b-a3b](https://huggingface.co/Mapika/decider-35b-a3b) v1 | Qwen3.5-35B-A3B-Base (3B active) | 65 GB bf16 | when accuracy is worth 3 to 4 times the cost per decision: knowledge and multi-step questions, long policies | 0.855 / 0.810, above the 2B on 93 of 95 tasks; JevBench hard 0.676; Bespoke 0.774; no RL stage |
150
  | [decider-35b-a3b-nvfp4](https://huggingface.co/Mapika/decider-35b-a3b-nvfp4) | the 35B in NVFP4 | 19.6 GB | the 35B on Blackwell through vLLM or TensorRT-LLM | 1.0 to 1.5 points under bf16 on the measured fixtures |
151
  | [decider-0.8b](https://huggingface.co/Mapika/decider-0.8b) | Qwen3.5-0.8B-Base | 1.4 GB bf16 | the smallest: routing, yes/no and short-state lookups within 1 to 4 points of the 2B, 1.5x faster | 0.776 / 0.707 on the single-run protocol (2B: 0.809 / 0.739) |
 
166
 
167
  The API is the same as decider-2b's: `decide_batch` scores many states with many questions in one call, `abstain_below=t`
168
  returns `None` under a confidence threshold, a question can have 2 to 255 options, and `system_one` / `decider.serve` accept
169
+ TypeSafe's `POST /v1/systemone` request shape. Every question and every Score level is scored in its own row. The state may be a
170
+ string, object or array of up to 32k tokens. See the decider-2b card for the full description of the request shape, field types
171
+ and the schema cache.
172
 
173
  Requirements: `torch`, `transformers>=5`, and `flash-linear-attention` (Triton kernels for the Qwen3.5 linear-attention layers;
174
+ the model runs without it but several times slower). The weights take 8.4 GB in bf16. v2.1 uses the plain prompt layout, as v1 and
175
+ v2 do. The per-type temperatures need decider-ai 1.4.0 or later (or the `decider/` subset in this repository, taken from 1.4.0).
176
+ Checked on CUDA, eager path: decider-ai 1.3.0 loads v2.1 with the name `decider-4b-v2.1` and temperature 1.099 and gives exactly
177
+ the probabilities that 1.4.0 gives with the map switched off; 1.4.0 and the `decider/` subset in this repository load the map
178
+ and give the same probabilities as each other. For Choice and yes/no answers the 1.4.0 output with the map equals
179
+ softmax(log(p) · 1.099 / T_type) of the 1.3.0 output p (to 1e-7 on `decide`, to 8e-5 on the four-decimal `system_one` values);
180
+ an isolated-level Score answer is rescaled per level row (each level's yes/no pair at 1.287 instead of 1.099) before the
181
+ levels are normalised, so it cannot be recomputed from the combined Score probabilities. The 1.4.0 HTTP server with CUDA graphs was checked with `/v1/systemone` and `/decide` requests. Other package
182
+ versions, MPS and the FP8 path were not checked on v2.1. On Blackwell GPUs use 1.0.2 or later (1.0.0 and 1.0.1 have the cuDNN
183
+ attention fault fixed in 1.0.2). The model is dense, so the CUDA-graph engine, `torch.compile` and the FP8 path of the helper
184
+ package apply to it as to decider-2b (`use_graphs=False` selects eager PyTorch).
185
 
186
  Without the helper package, the same computation in plain `transformers`:
187
 
 
195
  with torch.no_grad():
196
  logits = m(**ids).logits[0, -1]
197
  letters = [tok.encode(L, add_special_tokens=False)[0] for L in "ABC"]
198
+ probs = torch.softmax(logits[letters].float() / 1.110, -1) # 1.110 is the stored choice temperature (yes/no 1.560, Score 1.287)
199
  ```
200
 
201
  ## How it works
 
203
  The prompt is `Context: ...` followed by, for each question, the question text, the lettered options `(A) ... (B) ...` and an
204
  answer slot `Answer k: (`. The hidden state at each slot is projected with the option-letter rows of the LM head and softmaxed
205
  over the valid letters, divided by the temperature in `decider_config.json`. Letters are never generated, so all slots are read
206
+ from one pass. From decider-ai 1.4.0 the config may also hold `temperature_by_type`, one temperature per answer type
207
+ (`choice`, `noul`, `score`; a missing type uses `temperature`). This release's config has such a map: Choice answers use 1.110,
208
+ yes/no (noul) answers 1.560, Score answers 1.287 on each of their level rows. Package versions before 1.4.0 use `temperature`
209
+ (1.099) for every answer. `DECIDER_TEMPERATURE=T` or `Decider(path, temperature=T)` replaces `temperature` with T and switches the
210
+ map off, so every answer then uses T. Large label sets were sub-sampled to at most 10 options per
211
+ training example (gold always kept, order shuffled), so the model conditions on the supplied candidates rather than on a fixed
212
+ head.
213
 
214
  ## Training
215
 
 
242
  through and moves the weights further from the base model; without it, updates below the bf16 resolution round away and more of
243
  the base model's knowledge is kept.
244
 
245
+ **Stage 2 (v2.1).** A LoRA of rank 64 (alpha 128) on the attention and MLP weights of v1, trained for 2 epochs over 29,325 rows
246
+ in v1's plain state-first layout with isolated Score levels, then merged into the bf16 weights:
 
247
 
248
+ | source | rows | content | target |
249
+ |---|---|---|---|
250
+ | generated decision families | 8,000 | ten families (temporal and numeric decisions, subtle answer judgment, long policies, multi-hop lookup, abstention, probability, constrained trade-offs, safety judgment, paraphrase sensitivity, adversarial traps); the answers are computed by the generating code | the label |
251
+ | questions over business documents, written by Qwen3.6-27B with thinking on | 11,356 | in two rounds, one realistic business document (up to 34 domains and 22 document kinds) plus three or four typed questions per writer call; each question was answered twice more by the same model in fresh contexts, with shuffled options and without the writer's answer, and kept only when both answers agreed with the writer's (89% and 91% kept) | the label |
252
+ | human-labelled public sets (training halves) | 3,293 | MMLU, ARC, CommonsenseQA, BoolQ, MNLI, SNLI, Banking77, RACE, OpenBookQA, LogiQA 2, MedQA, Winogrande | the label |
253
+ | replay of mixture v2 | 6,676 | 100 rows from the training half of each of the 67 in-task regression tasks | v1's answer distribution: loss KL(p_v1 ‖ p_model) at temperature 1 |
254
+
255
+ These are v2's rows without 31 that the v2.1 checks found to share a context with an evaluation file (7 human-labelled rows and
256
+ 24 replay rows).
257
 
258
  | stage 2 | |
259
  |---|---|
260
  | trainable parameters | LoRA rank 64, alpha 128, on the attention and MLP projections; merged after training |
261
+ | loss | cross-entropy on the slot readout for labelled rows; KL(p_v1 ‖ p_model) over the options for replay rows, with p_v1 from the frozen v1 weights; the step loss is the sum over answers divided by the number of answers |
262
+ | schedule | learning rate 1e-4, 5% warm-up then cosine, 1,518 steps of 65,536 tokens (2 epochs), seed 0 |
263
+ | hardware | one NVIDIA B300, 94 minutes |
264
 
265
  No JevBench item and no Decision Index item was used for training, for writing the generators or the document questions, for
266
+ selecting the checkpoint, or for the temperatures. The ten generated families and the skill list of the document questions were
267
  written from the family names that JevBench publishes for its sealed set, not from its items. Every training row was checked
268
+ against every evaluation file used for selection, the canonical examples of the evaluation half of all 153 mixture-v2 evaluation
269
+ tasks, the four request fixtures and the issue #9 cases: there is no exact state-and-question overlap (a few contexts under 40
270
+ characters, such as a short utterance, occur under another task and question). The held-out sets used for selection are
271
+ generated families from held-out templates (heldout_jb) and document questions from business domains that are not in the
272
+ training data (test_teacher2). Two arms were trained to the end (B4 and C4, which differ in the size of the replay; two further arms were
273
+ stopped or cancelled at the user's request); the rule selected B4 after epoch 2.
274
+
275
+ **Temperatures.** `temperature` 1.099, fitted by NLL (all rows pooled) on the in-task half of the public regression set without
276
+ Banking77, CLINC-OOS, MMLU, ARC, Winogrande and HellaSwag: 61 tasks, 102,804 rows. `temperature_by_type` fitted by NLL per answer
277
+ type with `decider.calibrate.fit_by_type` on a pool of those regression rows and our own validation rows (the validation halves of
278
+ the generated families and document questions, human-labelled validation rows, held-out replay rows): 108,927 Choice answers,
279
+ 1,677 yes/no answers and 535 Score answers. The Choice pool is 94% regression rows, so the Choice temperature (1.110) stays close
280
+ to the global one. v2's temperature (1.935) and v1's (1.05) were one value each.
281
 
282
  ## Evaluation
283
 
284
+ All v2.1 numbers are with the per-type map through decider-ai 1.4.0, except where a paragraph says otherwise. Greedy play and
285
+ accuracy do not depend on the temperature; sampled play and calibration do.
286
 
287
  **Public regression set**, rebuilt on this machine (95 tasks: 67 in-task, 28 held-out; large label sets sub-sampled to 10
288
+ options; the temperatures fitted on in-task data). The rows are the same for every model. ECE is the expected calibration
289
+ error with 15 bins, per-task mean. The regression rows are all Choice answers, so v2.1 reads them at the Choice temperature 1.110.
290
 
291
  | model | in-task acc / NLL / ECE (67 tasks) | held-out acc / NLL / ECE (28 tasks) |
292
  |---|---|---|
293
+ | decider-2b v10, T=1.30 | 0.806 / 0.474 / 0.038 | 0.755 / 0.622 / 0.084 |
294
+ | decider-2b v11, per-type map | 0.802 / 0.481 / 0.038 | 0.752 / 0.626 / 0.083 |
295
  | decider-4b v1, T=1.05 | 0.834 / 0.404 / 0.027 | 0.788 / 0.558 / 0.071 |
296
+ | decider-4b v2, T=1.935 | 0.824 / 0.441 / 0.041 | 0.779 / 0.566 / 0.080 |
297
+ | **decider-4b v2.1 (this repository), per-type map** | **0.831 / 0.414 / 0.031** | **0.784 / 0.569 / 0.077** |
298
  | decider-35b-a3b v1, T=1.08 | 0.855 / 0.357 / 0.026 | 0.810 / 0.497 / 0.069 |
299
 
300
+ Per-task paired intervals were not recomputed for v2.1; the v2 card (tag `v2`) has them for v2 against v10 and the 35B, and v2.1
301
+ is within 0.7 points of v2 on both halves.
 
 
 
 
 
 
302
 
303
+ **The same rows as v1 and v2** are in the table under [Changes from v2](#changes-from-v2): the four fixtures (847 in-task
304
+ validation rows, OpenJev, Mind2Web, TypeSafe), our held-out hard sets, JevBench, Bespoke, live browser tasks, zero-shot games,
305
+ text games, behaviour probes and the issue #9 cases.
306
 
307
+ **Ten text games, zero-shot** (`decider.games.play`, five episodes per game, greedy, eager; the 2B was trained on the first four,
308
+ the 4B on none): Pong −5 (v1 −21, v2 −21; the 2B 8), Breakout 35 (v1 14, v2 12; the 2B 22), CliffWalking −13 (v1 −13, v2 −60;
309
+ teacher −13), BabyAI-GoTo 0.19 (v1 0.54, v2 0.35), Freeway 0 (v2 1), Blackjack −0.6, FrozenLake, MiniGrid-Empty, LavaGap and DoorKey
310
+ 0 as for v1 and v2. Greedy play takes the most probable option, so these results do not depend on the temperature.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
311
 
312
  **JevBench public items** (231 items of [Benchmark Heaven](https://benchmarkheaven.com/jev-models); argmax over the exact
313
+ label set with the request the harness's TypeSafe adapter builds). Read once for v2.1 at its global temperature (1.3.0) and once
314
+ with the map (1.4.0), both after the rule decisions: easy 1.000, standard 0.986, hard 0.649 (72 of 111) in both reads. Top-label
315
+ ECE on the hard tier is 0.184 with the map (0.210 at the global temperature); by answer type with the map, Choice 0.156 (67 items),
316
+ yes/no 0.259 (38), Score 0.354 (6). Mean confidence on the hard tier is 0.799 at accuracy 0.649. For comparison on the same items:
317
+ v2 0.676 (ECE 0.104 as read at T 1.719; 0.071 recomputed at its release temperature 1.935 with 6 Score items kept at 1.719), v1
318
+ 0.550 (ECE 0.288), decider-35b-a3b 0.676, Jev 1.13.0 0.730. The public hard tier is 111 items (95% interval about ±9 points).
 
 
 
 
 
319
 
320
  **Bespoke's public suite** (13 human-labelled subsets, 3,880 records in Jev's wire format, answered through `system_one` as
321
+ shipped). v1 and v2 through decider-ai 1.3.0, v2.1 through 1.4.0 with the map, same session. Nimble-9B (0.748 / 0.759 macro /
322
+ micro) and Jev 1.13.0 (0.760 / 0.773) are in Bespoke's report.
323
 
324
+ | subset (type) | decider-4b v1 | decider-4b v2 | decider-4b v2.1 |
325
+ |---|---|---|---|
326
+ | vitaminc-dev (choice) | 0.756 | 0.778 | 0.743 |
327
+ | massive-en-US (choice; trained) | 0.860 | 0.869 | 0.863 |
328
+ | massive-de-DE (choice) | 0.843 | 0.837 | 0.840 |
329
+ | boolq (noul; trained) | 0.873 | 0.860 | 0.867 |
330
+ | squad2 (noul) | 0.706 | 0.793 | 0.789 |
331
+ | paws (noul; trained) | 0.716 | 0.832 | 0.764 |
332
+ | multinli (choice; trained) | 0.933 | 0.926 | 0.933 |
333
+ | civil_comments (noul; trained) | 0.880 | 0.857 | 0.873 |
334
+ | aegis2 (noul) | 0.820 | 0.812 | 0.804 |
335
+ | helpsteer2 (score; trained) | 0.466 | 0.466 | 0.450 |
336
+ | summeval-relevance (score) | 0.425 | 0.463 | 0.396 |
337
+ | summeval-consistency (score) | 0.833 | 0.826 | 0.799 |
338
+ | pubmedqa (choice; trained) | 0.728 | 0.724 | 0.708 |
339
+ | **macro / micro** | 0.757 / 0.765 | 0.773 / 0.781 | 0.756 / 0.765 |
340
+ | macro over the subsets not trained on | 0.731 | 0.751 | 0.728 |
341
+
342
+ v2.1 is at v1's level on this suite and 1.6 points under v2 (macro). The largest differences to v2 are PAWS (0.764 against 0.832),
343
+ SummEval relevance (0.396 against 0.463) and VitaminC (0.743 against 0.778).
344
+
345
+ **Behaviour probes** (teacher-labelled, same probes as the other releases; v1 / v2 in brackets): generic-versus-specific bucket
346
+ choice 1.00 / 1.00 (1.00 / 1.00; 0.95 / 1.00), catch-all when nothing fits 0.95 (0.90; 0.95), abstention battery 8 of 8 (7 of 8;
347
+ 8 of 8); model-router tier 0.968 (0.968; 0.935) and needs-live-data 0.839 (0.871; 0.839); command-risk classification 0.889
348
+ (0.889; 0.911) with no destructive command called safe, touches-outside-project 0.933 (0.956; 0.933); browser-agent element and
349
+ action choice 0.938 / 0.750 (0.875 / 0.875; 0.938 / 0.938). The yes/no Brier score on the probe batteries is 0.010 with the map
350
+ and 0.006 at the global temperature: the yes/no temperature 1.560 flattens easy yes/no answers there.
 
351
 
352
  ## Calibration
353
 
354
+ The per-type map was fitted on everyday regression rows and our own validation rows. It helps where it can: yes/no and Score
355
+ answers on the held-out sets. It does not change Choice answers much, because the Choice temperature fitted on a pool that is
356
+ 94% everyday rows stays at 1.110. On the held-out generated families, whose answers are two-thirds Choice, v2.1 is
357
+ overconfident: calibration error 0.147 with the map, where our own release limit is 0.08. v2 is at 0.046 there, only because its
358
+ one temperature of 1.935 flattens every answer, which is also what cost it sampled play. ECE below is 10-bin top-label ECE on
359
+ stored temperature-1 logits read at the served temperatures.
 
360
 
361
+ | set | rows | accuracy | ECE at the global T / with the map | choice / noul / score ECE with the map |
362
+ |---|---|---|---|---|
363
+ | heldout_jb | 5,000 | 0.556 | 0.163 / 0.147 | 0.170 / 0.101 / 0.108 |
364
+ | test_teacher2 | 449 | 0.820 | 0.050 / 0.043 | 0.057 / 0.048 / 0.086 |
365
+ | guard | 2,994 | 0.819 | 0.019 / 0.017 | 0.017 / − / − |
366
+ | cal_human | 1,595 | 0.856 | 0.019 / 0.021 | 0.021 / − / − |
367
+
368
+ On the held-out generated families: v1 (T 1.05) 0.260, v2 (T 1.935) 0.046, v2.1 at T = 1 0.182. A map fitted without the
369
+ regression rows (Choice 1.259) would reach 0.127 there and would make the regression set less calibrated (in-task ECE 0.0375
370
+ against 0.0309); it was not used. On the fixtures (10-bin ECE, map): 847 validation rows 0.035 (v1 0.037, v2 0.031), TypeSafe
371
+ 0.063 (v1 0.124, v2 0.072), OpenJev 0.154 (v1 0.158, v2 0.131), Mind2Web 0.017 (v1 0.016, v2 0.051). The Decision Index
372
+ sample was not read for v2.1. If you route on confidence, calibrate on your own labels; `python -m decider.calibrate` fits a
373
+ per-type map from your own answers read at temperature 1.
 
 
 
 
 
 
 
 
374
 
375
  ## Speed
376
 
377
+ v2.1 has the same architecture and size as v1 and v2, and the temperature map is a division per answer; the speed was not
378
+ measured again. The measurements of v1 and v2 (the same weights layout): in one session on one unshared NVIDIA B300 (bf16), 34.6
379
+ ms (v2) and 32.4 ms (v1) median per decision over 200 game-state decisions of 156 tokens median, batch of one, eager PyTorch
380
+ without CUDA graphs or `torch.compile`, timed around the forward pass with `torch.cuda.synchronize()`. The eager path is
381
+ launch-bound, so host load changes it: v1's first measurement was 24.7 ms (10th to 90th percentile 24.6 to 43.5 ms), with
382
+ decider-2b at 17.9 ms and decider-35b-a3b at 41.4 ms on the same decisions and method. With the helper's CUDA graphs and
383
+ `torch.compile`, one support-ticket request (228 tokens, 3 questions) takes 5.2 ms (FP8 5.0 ms); a batch of 32 such states takes
384
+ 81.5 ms, 1,178 decisions per second (FP8 71.2 ms, 1,349 per second). The HTTP server (`/decide`, bf16) answers 72.8 requests per
385
+ second at a median of 13.3 ms with one client and 190 requests per second with 64 clients.
 
386
 
387
  ## Limitations
388
 
389
+ * No reinforcement-learning stage: stated beliefs about action outcomes were not trained against exact laws. On live browser
390
+ tasks v2.1 is at 93.2% sampled and 87.5% on the six held-out tasks; decider-2b v10, which has the RL stage, is at 93.2% and 91.7%.
391
+ * Overconfident on hard multi-step items: calibration error 0.147 on our held-out generated families and 0.184 on the JevBench
392
+ public hard tier, against v2's 0.046 and 0.104. On those items, do not read a confidence of 0.8 as an 80% chance of being right.
393
+ * Issue #9 form case c_1 is answered wrongly (see Changes from v2); BabyAI-GoTo 0.19 against v1's 0.54; greedy bag-draw play 9.4
394
+ points under v1.
395
+ * About 0.3 points under v1 on the regression set and 1.1 points under it on the 847 validation rows; needs-live-data and
396
+ touches-outside-project probes one item under v1 each.
397
+ * The per-type map needs decider-ai 1.4.0 or later. With 1.3.0 or earlier, or with `DECIDER_TEMPERATURE` set, every answer uses
398
+ 1.099: the answers are the same, yes/no answers are sharper (on the held-out generated families their ECE is 0.146 instead of
399
+ 0.101) and Score answers are sharper.
400
+ * Below the 35B by 2.4 / 2.6 points on the regression set and 2.7 on the JevBench public hard tier (3 items).
401
+ * The JevBench public hard tier is 111 items; differences of a few items between versions are within its noise.
 
 
 
402
  * The stage-2 LoRA was trained only in the plain state-first layout. The schema-first layout (the server's opt-in schema cache)
403
+ was not measured on v2.1, and `decider_config.json` does not mark it as trained for that layout (`schema_first_trained:
404
+ false`), so `DECIDER_SCHEMA_CACHE=1` does not turn the schema cache on for this model.
405
  * Mixture v2's 26 additional public datasets and ten programmatic families, and stage 2's generators and document questions, are
406
  described above but their builders are not in the public package; `scripts/train.sh full` reproduces the 60% of stage 1's data
407
+ that is the public mixture.
 
408
  * English is the main language; the multilingual rows (XNLI, PAWS-X, MASSIVE, Belebele, XCOPA) are a small share of the data
409
+ and were not measured beyond the mixture-v2 evaluation set and Bespoke's German MASSIVE subset.
410
  * Everything else in the decider-2b card's limitations (packed questions see each other, long JSON arrays by position, full
411
  label sets against sampled options, abstention wording, rules in the question) applies; those shapes were not re-measured at
412
  this size.
 
415
 
416
  | version | what changed |
417
  |---|---|
418
+ | **v2.1** (2026-09-24, these weights) | v1 + a merged LoRA on v2's rows (29,325 after removing 31 that overlapped evaluation files), with the replay rows trained toward v1's own answer distribution; temperature 1.099 and `temperature_by_type` {choice 1.110, noul 1.560, score 1.287} (decider-ai 1.4.0; older versions use 1.099). Against v2: sampled play recovered (bag-draw 52.0% against 37.9%, browser 93.2% against 88.1%), CliffWalking −13 against −60, regression set +0.7 / +0.5; hard sets at v2's level; less calibrated on hard items (0.147 against 0.046); issue #9 c_1 still wrong. Did not pass its pre-registered rules; released on the full comparison |
419
+ | v2 (2026-09-24, Hub tag `v2`) | v1 + a merged LoRA (rank 64, attention and MLP, 2 epochs, 29,356 rows: generated decision families, document questions written by Qwen3.6-27B and kept when two independent answers agreed, human-labelled public sets, replay of mixture v2); temperature 1.935. Better on hard judgments, TypeSafe, OpenJev and Bespoke's suite; worse on sampled play, some text games and by about 1 point on everyday tasks |
420
  | v1 (2026-09-22, Hub tag `v1`) | first release: one pass over mixture v2 on Qwen3.5-4B-Base with AdamW on bf16 parameters, no RL stage; temperature 1.05 |
421
 
422
  The GitHub repository's [docs/CHANGELOG.md](https://github.com/Mapika/decider/blob/main/docs/CHANGELOG.md) lists every
 
427
  Code, data registry, training and evaluation scripts and the per-version history: https://github.com/Mapika/decider
428
  (`docs/HISTORY.md`, section "decider-4b"). Stage 1 was trained with the data-parallel trainer of the architecture A/B study
429
  (`arch_ab/train_dp_optvar.py` in the research repository, optimizer variant `bf16`); stage 2 with a LoRA trainer in the research
430
+ repository. Both were evaluated with the public `decider.evaluate` and the head-to-head tools. `eval_results.json` in this
431
+ repository has the regression metrics with the map and at the global temperature, our held-out sets by answer type, the
432
+ fixtures, games, browser and Bespoke results with the paired comparisons against v1 and v2, the text games, the behaviour
433
+ probes, the issue #9 cases and the JevBench public items.
434
 
435
  **Independence.** This is an independent project. It is not affiliated with or endorsed by TypeSafe AI. It is an open
436
  reproduction of the "System One" model class (TypeSafe AI's Jev); nothing was distilled from Jev. The training data is public
decider/calibrate.py ADDED
@@ -0,0 +1,222 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Fit decider_config.json "temperature_by_type" by NLL, one temperature per answer type (decider.temperature).
2
+
3
+ python -m decider.calibrate records.jsonl [--min-rows 50]
4
+ -> {"temperature_by_type": {"choice": 1.48, "noul": 2.22, "score": 1.38}, "temperature": 1.6, "rows": {...}, "nll": {...}}
5
+
6
+ A record is one answer with its gold, read at temperature 1:
7
+ {"type": "choice" | "noul" | "score", "logits": [one value per option], "gold": option index}
8
+ {"type": "score", "level_logits": [[no, yes], one pair per level], "gold": level index} Score with isolated levels
9
+ "probs" / "level_probs" (probabilities at temperature 1) may stand in for "logits" / "level_logits": log p is the logit up to a
10
+ constant, which the softmax ignores. A noul gold is 1 for true, 0 for false. Every record is checked first (check_record):
11
+ a malformed one (gold out of range or not an integer, wrong shape, NaN, a row without a finite logit, an isolated Score whose
12
+ levels all have P(yes) = 0) raises ValueError naming its position.
13
+
14
+ An isolated-levels Score answer is fitted through the readout the server uses: every level row is softmax(row / T), and the
15
+ answer is P(yes) of each level divided by their sum (decider.systemone.combine_isolated). So the "score" temperature is fitted on
16
+ Score answers whichever readout the model serves them with; fit it on records collected with the same isolated_levels setting.
17
+
18
+ `collect(decider, examples)` produces records from a decider.infer.Decider and labelled /v1/systemone-shaped examples.
19
+ "temperature" in the output is one temperature fitted on all records together, for comparison; a type with fewer than
20
+ --min-rows records is left out of the map (it then uses "temperature" of the config).
21
+ """
22
+ import json, math, sys
23
+ import numpy as np
24
+
25
+ from decider.temperature import TYPES
26
+
27
+ GRID = np.exp(np.linspace(math.log(0.05), math.log(20.0), 801)) # 0.05 .. 20, about 0.75 % apart
28
+
29
+
30
+ def _log(x):
31
+ x = np.asarray(x, dtype=np.float64)
32
+ with np.errstate(divide="ignore"):
33
+ return np.where(x > 0, np.log(np.clip(x, 1e-300, None)), -np.inf)
34
+
35
+
36
+ def _logits(rec, key):
37
+ if key in rec:
38
+ return np.asarray(rec[key], dtype=np.float64)
39
+ alt = {"logits": "probs", "level_logits": "level_probs"}[key]
40
+ return _log(rec[alt])
41
+
42
+
43
+ def _lse(z, axis=-1):
44
+ m = np.max(z, axis=axis, keepdims=True)
45
+ m = np.where(np.isfinite(m), m, 0.0)
46
+ with np.errstate(divide="ignore"): # an all -inf row gives -inf, handled by the callers
47
+ return (m + np.log(np.sum(np.exp(z - m), axis=axis, keepdims=True))).squeeze(axis)
48
+
49
+
50
+ def nll(rec, T):
51
+ """Negative log-likelihood of one record's gold at temperature T."""
52
+ return float(_Batch([rec]).nll(T))
53
+
54
+
55
+ def _isolated(rec):
56
+ return "level_logits" in rec or "level_probs" in rec
57
+
58
+
59
+ def check_record(rec, i=None):
60
+ """Raise ValueError unless `rec` is a well-formed record (module docstring): an integral gold within the options or levels,
61
+ a 1-D row of at least two options or an [levels, 2] array of at least two levels, logits that are not NaN or +inf,
62
+ probabilities in [0, 1] with positive mass, and isolated levels only on a Score record."""
63
+ where = f"record {i}" if i is not None else "record"
64
+ if not isinstance(rec, dict):
65
+ raise ValueError(f"{where}: not a JSON object")
66
+ if rec.get("type") not in TYPES:
67
+ raise ValueError(f"{where}: type {rec.get('type')!r}, expected one of {', '.join(TYPES)}")
68
+ iso = _isolated(rec)
69
+ keys = ("level_logits", "level_probs") if iso else ("logits", "probs")
70
+ if sum(k in rec for k in keys) != 1 or (iso and ("logits" in rec or "probs" in rec)):
71
+ raise ValueError(f"{where}: give exactly one of logits, probs, level_logits, level_probs")
72
+ if iso and rec["type"] != "score":
73
+ raise ValueError(f"{where}: isolated levels (level_logits / level_probs) belong to a score record, not {rec['type']!r}")
74
+ key = keys[0] if keys[0] in rec else keys[1]
75
+ try:
76
+ a = np.asarray(rec[key], dtype=np.float64)
77
+ except (TypeError, ValueError):
78
+ raise ValueError(f"{where}: {key} is not a numeric array") from None
79
+ if iso and (a.ndim != 2 or a.shape[1] != 2 or a.shape[0] < 2):
80
+ raise ValueError(f"{where}: {key} must be one [no, yes] pair per level, at least two levels; got shape {a.shape}")
81
+ if not iso and (a.ndim != 1 or a.shape[0] < 2):
82
+ raise ValueError(f"{where}: {key} must be one value per option, at least two options; got shape {a.shape}")
83
+ if key.endswith("probs"):
84
+ if not np.all(np.isfinite(a)) or np.any(a < 0) or np.any(a > 1) or np.any(a.sum(-1) <= 0):
85
+ raise ValueError(f"{where}: {key} must be probabilities in [0, 1] with positive mass per row")
86
+ elif np.any(np.isnan(a)) or np.any(a == np.inf) or not np.all(np.any(np.isfinite(a), axis=-1)):
87
+ raise ValueError(f"{where}: {key} contains NaN or +inf, or a row without a finite value")
88
+ if rec["type"] == "noul" and a.shape[0] != 2:
89
+ raise ValueError(f"{where}: a noul record has exactly two options (false, true); got {a.shape[0]}")
90
+ if iso and not np.any(a[:, 1] > 0 if key == "level_probs" else np.isfinite(a[:, 1])):
91
+ raise ValueError(f"{where}: no level has a yes probability above 0, so the served Score answer has no distribution")
92
+ g = rec.get("gold")
93
+ if isinstance(g, bool) or not isinstance(g, (int, np.integer)) or not 0 <= g < a.shape[0]:
94
+ raise ValueError(f"{where}: gold must be an integer index in 0..{a.shape[0] - 1}, got {g!r}")
95
+ return rec
96
+
97
+
98
+ class _Batch:
99
+ """Records packed into padded arrays, so one temperature is scored over all of them at once."""
100
+ def __init__(self, records):
101
+ records = [check_record(r, i) for i, r in enumerate(records)]
102
+ lists = [r for r in records if not _isolated(r)]
103
+ isos = [r for r in records if _isolated(r)]
104
+ self.n = len(lists) + len(isos)
105
+ self.L = self.Lg = self.I = self.Im = self.Ig = None
106
+ if lists:
107
+ zs = [_logits(r, "logits") for r in lists]
108
+ K = max(len(z) for z in zs)
109
+ self.L = np.full((len(zs), K), -np.inf)
110
+ for i, z in enumerate(zs):
111
+ self.L[i, :len(z)] = z
112
+ self.Lg = np.array([int(r["gold"]) for r in lists])
113
+ if isos:
114
+ zs = [_logits(r, "level_logits").reshape(-1, 2) for r in isos]
115
+ n = max(len(z) for z in zs)
116
+ self.I = np.zeros((len(zs), n, 2)); self.Im = np.zeros((len(zs), n), dtype=bool)
117
+ for i, z in enumerate(zs):
118
+ self.I[i, :len(z)] = z; self.Im[i, :len(z)] = True
119
+ self.Ig = np.array([int(r["gold"]) for r in isos])
120
+
121
+ def nll(self, T):
122
+ """Sum of the records' NLL at temperature T."""
123
+ tot = 0.0
124
+ if self.L is not None:
125
+ z = self.L / T
126
+ zg = z[np.arange(len(z)), self.Lg]
127
+ tot += float(np.sum(np.where(np.isfinite(zg), _lse(z) - zg, 690.0)))
128
+ if self.I is not None: # in log space: P(yes) of very confident rows underflows otherwise
129
+ z = self.I / T
130
+ with np.errstate(invalid="ignore"):
131
+ lp = np.where(self.Im, z[..., 1] - _lse(z), -np.inf) # log P(yes) per level, -inf on padding
132
+ norm = _lse(lp) # log of the summed P(yes)
133
+ lg = lp[np.arange(len(z)), self.Ig]
134
+ with np.errstate(invalid="ignore"):
135
+ per = np.where(np.isfinite(lg), norm - lg, 690.0) # check_record guarantees a finite norm
136
+ tot += float(np.sum(per))
137
+ return tot
138
+
139
+
140
+ def mean_nll(records, T):
141
+ b = records if isinstance(records, _Batch) else _Batch(records)
142
+ return b.nll(T) / b.n
143
+
144
+
145
+ def fit(records, grid=GRID):
146
+ """The temperature with the lowest mean NLL on the grid, refined by a golden-section search between its neighbours."""
147
+ records = records if isinstance(records, _Batch) else _Batch(list(records))
148
+ if not records.n:
149
+ raise ValueError("no records to fit")
150
+ vals = [mean_nll(records, T) for T in grid]
151
+ i = int(np.argmin(vals))
152
+ lo, hi = math.log(grid[max(i - 1, 0)]), math.log(grid[min(i + 1, len(grid) - 1)])
153
+ r = (math.sqrt(5) - 1) / 2
154
+ a, b = hi - r * (hi - lo), lo + r * (hi - lo)
155
+ fa, fb = mean_nll(records, math.exp(a)), mean_nll(records, math.exp(b))
156
+ for _ in range(40):
157
+ if fa < fb:
158
+ hi, b, fb = b, a, fa; a = hi - r * (hi - lo); fa = mean_nll(records, math.exp(a))
159
+ else:
160
+ lo, a, fa = a, b, fb; b = lo + r * (hi - lo); fb = mean_nll(records, math.exp(b))
161
+ T = math.exp((lo + hi) / 2)
162
+ return T if mean_nll(records, T) <= vals[i] else float(grid[i])
163
+
164
+
165
+ def fit_by_type(records, min_rows=50):
166
+ """-> {"temperature_by_type": {type: T}, "temperature": pooled T, "rows": {type: n}, "nll": {type: {"T=1": .., "fitted": ..}}}"""
167
+ records = [check_record(r, i) for i, r in enumerate(records)]
168
+ groups = {t: [] for t in TYPES}
169
+ for r in records:
170
+ groups[r["type"]].append(r)
171
+ out = {"temperature_by_type": {}, "temperature": round(fit(records), 3) if records else None,
172
+ "rows": {t: len(g) for t, g in groups.items()}, "nll": {}}
173
+ for t, g in groups.items():
174
+ if len(g) < max(1, min_rows):
175
+ continue
176
+ b = _Batch(g); T = fit(b)
177
+ out["temperature_by_type"][t] = round(T, 3)
178
+ out["nll"][t] = {"T=1": round(mean_nll(b, 1.0), 4), "fitted": round(mean_nll(b, T), 4)}
179
+ return out
180
+
181
+
182
+ def collect(decider, examples, independent=True):
183
+ """Records for fit_by_type from a decider.infer.Decider, read at temperature 1 on the uncached state-first path.
184
+ examples: iterable of (state, questions, golds) with questions as /v1/systemone takes them and golds {question id: gold},
185
+ the gold being a Choice option name, a Noul true/false, or a Score level index. Questions without a gold are skipped.
186
+ Score questions are read with the Decider's isolated_levels setting, as system_one serves them."""
187
+ recs = []
188
+ for state, questions, golds in examples:
189
+ rqs, index, items = decider._system_one_items(state, questions, independent, layout="state_first")
190
+ rows = decider._system_one_probs(items, "state_first", temperature=1.0)
191
+ for k, kind, s, n in index:
192
+ if k not in golds:
193
+ continue
194
+ rq, g = rqs[k], golds[k]
195
+ if rq["type"] == "noul" and not (isinstance(g, bool) or g in (0, 1)):
196
+ raise ValueError(f"gold of noul question {k!r} must be true or false, got {g!r}")
197
+ if rq["type"] == "score" and (isinstance(g, bool) or not isinstance(g, int) or not 0 <= g < len(rq["names"])):
198
+ raise ValueError(f"gold of score question {k!r} must be a level index in 0..{len(rq['names']) - 1}, got {g!r}")
199
+ key = bool(g) if rq["type"] == "noul" else g
200
+ if key not in rq["names"]:
201
+ raise ValueError(f"gold of choice question {k!r} is not one of its options: {g!r}")
202
+ gold = rq["names"].index(key)
203
+ if kind == "iso":
204
+ recs.append({"type": rq["type"], "level_probs": [rows[s + j][:2] for j in range(n)], "gold": gold})
205
+ else:
206
+ recs.append({"type": rq["type"], "probs": rows[s][:len(rq["options"])], "gold": gold})
207
+ return recs
208
+
209
+
210
+ def main(argv=None):
211
+ import argparse
212
+ ap = argparse.ArgumentParser(description=__doc__.split("\n\n")[0])
213
+ ap.add_argument("records", help="JSON lines, one record per line (see the module docstring)")
214
+ ap.add_argument("--min-rows", type=int, default=50, help="fit a type only when it has at least this many records")
215
+ a = ap.parse_args(argv)
216
+ with open(a.records) as f:
217
+ recs = [json.loads(line) for line in f if line.strip()]
218
+ print(json.dumps(fit_by_type(recs, a.min_rows), indent=1))
219
+
220
+
221
+ if __name__ == "__main__":
222
+ main(sys.argv[1:])
decider/engine.py CHANGED
@@ -7,6 +7,7 @@ option-letter logits for all positions [B, T, K]; slots are gathered outside.
7
  import time, torch, torch._dynamo, torch.nn.functional as F
8
  from decider.model import DecisionModel, collate
9
  from decider.prompt import build, MAX_OPTIONS
 
10
 
11
  T_BUCKETS = [64, 128, 192, 256, 320, 384, 512, 640, 768, 1024, 1280, 1536, 2048]
12
  B_BUCKETS = [1, 2, 4, 8, 16, 32, 64]
@@ -46,11 +47,12 @@ def patch_conv():
46
 
47
  def read_slots(out, rows, slots, nopts, temperature, n_per_item):
48
  """One gather + one softmax + one device-to-host copy for the whole batch (was: three small kernels and a sync per item).
49
- out [B, T, K] logits; rows/slots/nopts: flat python lists, one entry per question; n_per_item: questions per item."""
 
50
  dev = out.device; idx = torch.tensor([rows, slots, nopts], dtype=torch.long).to(dev, non_blocking=True)
51
  lg = out[idx[0], idx[1]] # [N, K]
52
  lg = lg.masked_fill(torch.arange(lg.shape[1], device=dev)[None, :] >= idx[2][:, None], float("-inf"))
53
- p = torch.softmax(lg / temperature, -1).cpu()
54
  return list(torch.split(p, n_per_item))
55
 
56
 
@@ -140,14 +142,15 @@ class Engine:
140
 
141
  @torch.no_grad()
142
  def score_items(self, items, temperature=1.0):
143
- """items: list of dicts from prompt.build. Returns list of [n_q, MAX_OPTIONS] prob tensors (cpu)."""
 
144
  Tmax = max(len(it["ids"]) for it in items)
145
  T = _bucket(Tmax, T_BUCKETS) or -(-Tmax // LONG_STEP) * LONG_STEP
146
  B = (_bucket(len(items), B_BUCKETS) or len(items)) if T <= GRAPH_MAX_T else len(items)
147
  ids = fill_ids([it["ids"] for it in items], B, T, self.tok.pad_token_id)
148
  out = self.logits_all(ids.to(self.dev, non_blocking=True))
149
  return read_slots(out, [b for b, it in enumerate(items) for _ in it["slots"]], [s for it in items for s in it["slots"]],
150
- [n for it in items for n in it["nopts"]], temperature, [len(it["slots"]) for it in items])
151
 
152
  @torch.no_grad()
153
  def score_shared(self, items, temperature=1.0, min_prefix=192):
 
7
  import time, torch, torch._dynamo, torch.nn.functional as F
8
  from decider.model import DecisionModel, collate
9
  from decider.prompt import build, MAX_OPTIONS
10
+ from decider.temperature import scaled_softmax, slot_temperatures
11
 
12
  T_BUCKETS = [64, 128, 192, 256, 320, 384, 512, 640, 768, 1024, 1280, 1536, 2048]
13
  B_BUCKETS = [1, 2, 4, 8, 16, 32, 64]
 
47
 
48
  def read_slots(out, rows, slots, nopts, temperature, n_per_item):
49
  """One gather + one softmax + one device-to-host copy for the whole batch (was: three small kernels and a sync per item).
50
+ out [B, T, K] logits; rows/slots/nopts: flat python lists, one entry per question; n_per_item: questions per item.
51
+ temperature: a number for every question, or a flat list with one temperature per question (decider.temperature)."""
52
  dev = out.device; idx = torch.tensor([rows, slots, nopts], dtype=torch.long).to(dev, non_blocking=True)
53
  lg = out[idx[0], idx[1]] # [N, K]
54
  lg = lg.masked_fill(torch.arange(lg.shape[1], device=dev)[None, :] >= idx[2][:, None], float("-inf"))
55
+ p = scaled_softmax(lg, temperature).cpu()
56
  return list(torch.split(p, n_per_item))
57
 
58
 
 
142
 
143
  @torch.no_grad()
144
  def score_items(self, items, temperature=1.0):
145
+ """items: list of dicts from prompt.build. Returns list of [n_q, MAX_OPTIONS] prob tensors (cpu).
146
+ temperature: a number, or one entry per item (a number or one number per slot; decider.temperature.for_items)."""
147
  Tmax = max(len(it["ids"]) for it in items)
148
  T = _bucket(Tmax, T_BUCKETS) or -(-Tmax // LONG_STEP) * LONG_STEP
149
  B = (_bucket(len(items), B_BUCKETS) or len(items)) if T <= GRAPH_MAX_T else len(items)
150
  ids = fill_ids([it["ids"] for it in items], B, T, self.tok.pad_token_id)
151
  out = self.logits_all(ids.to(self.dev, non_blocking=True))
152
  return read_slots(out, [b for b, it in enumerate(items) for _ in it["slots"]], [s for it in items for s in it["slots"]],
153
+ [n for it in items for n in it["nopts"]], slot_temperatures(temperature, items), [len(it["slots"]) for it in items])
154
 
155
  @torch.no_grad()
156
  def score_shared(self, items, temperature=1.0, min_prefix=192):
decider/engine_v2.py CHANGED
@@ -19,6 +19,7 @@ import time, torch, torch.nn.functional as F
19
  from decider import shared_prefix
20
  from decider.engine import read_slots, fill_ids, patch_conv, set_attention_backend_policy
21
  from decider.model import DecisionModel
 
22
 
23
  T_BUCKETS = [64, 128, 192, 256, 320, 384, 512, 640, 768, 1024, 1280, 1536, 2048, 3072, 4096, 6144, 8192]
24
  B_BUCKETS = [1, 2, 4, 8, 16, 32]
@@ -150,9 +151,11 @@ class EngineV2:
150
  # ---- scoring -----------------------------------------------------------
151
  @torch.no_grad()
152
  def score_items(self, items, temperature=1.0):
153
- """items: dicts from prompt.build / build_rows. -> one [n_q, MAX_OPTIONS] cpu probability tensor per item."""
 
154
  if not items:
155
  return []
 
156
  Tmax = max(len(it["ids"]) for it in items)
157
  T = self.t_bucket(Tmax)
158
  if T is None:
@@ -168,7 +171,7 @@ class EngineV2:
168
  lg = self.logits_all(ids.to(self.dev, non_blocking=True))
169
  out += read_slots(lg, [b for b, it in enumerate(chunk) for _ in it["slots"]],
170
  [s for it in chunk for s in it["slots"]], [n for it in chunk for n in it["nopts"]],
171
- temperature, [len(it["slots"]) for it in chunk])
172
  return out
173
 
174
  @torch.no_grad()
 
19
  from decider import shared_prefix
20
  from decider.engine import read_slots, fill_ids, patch_conv, set_attention_backend_policy
21
  from decider.model import DecisionModel
22
+ from decider.temperature import item_slice, slot_temperatures
23
 
24
  T_BUCKETS = [64, 128, 192, 256, 320, 384, 512, 640, 768, 1024, 1280, 1536, 2048, 3072, 4096, 6144, 8192]
25
  B_BUCKETS = [1, 2, 4, 8, 16, 32]
 
151
  # ---- scoring -----------------------------------------------------------
152
  @torch.no_grad()
153
  def score_items(self, items, temperature=1.0):
154
+ """items: dicts from prompt.build / build_rows. -> one [n_q, MAX_OPTIONS] cpu probability tensor per item.
155
+ temperature: a number, or one entry per item (a number or one number per slot; decider.temperature.for_items)."""
156
  if not items:
157
  return []
158
+ slot_temperatures(temperature, items) # a length mismatch fails before any forward
159
  Tmax = max(len(it["ids"]) for it in items)
160
  T = self.t_bucket(Tmax)
161
  if T is None:
 
171
  lg = self.logits_all(ids.to(self.dev, non_blocking=True))
172
  out += read_slots(lg, [b for b, it in enumerate(chunk) for _ in it["slots"]],
173
  [s for it in chunk for s in it["slots"]], [n for it in chunk for n in it["nopts"]],
174
+ slot_temperatures(item_slice(temperature, i - len(chunk), i), chunk), [len(it["slots"]) for it in chunk])
175
  return out
176
 
177
  @torch.no_grad()
decider/infer.py CHANGED
@@ -12,6 +12,7 @@ import logging
12
  import torch
13
  from decider.model import DecisionModel, collate
14
  from decider.prompt import build, MAX_OPTIONS, resolve_layout, chat_template
 
15
  from dataclasses import dataclass
16
 
17
 
@@ -49,11 +50,15 @@ class _NoShuffle: # keep option order as given
49
 
50
 
51
  class CompiledSchema:
52
- def __init__(self, d, rqs, h, index): self.d, self.rqs, self.h, self.index = d, rqs, h, index
 
 
 
53
 
54
  def batch(self, states, max_state_tokens=32768):
55
  from decider.systemone import render_state, assemble
56
- probs = self.d._se.score(self.h, [render_state(s) for s in states], temperature=self.d.T_schema, max_ctx_tokens=max_state_tokens)
 
57
  return [{"model": self.d.name, "answers": assemble(self.rqs, self.index, [p.tolist() for p in pr])} for pr in probs]
58
 
59
  def __call__(self, state, max_state_tokens=32768):
@@ -67,10 +72,15 @@ class Decider:
67
  the optional MPS patch; CPU defaults to bfloat16. Set ``use_graphs=False``
68
  for eager execution or debugging.
69
  """
70
- def __init__(self, path, device=None, dtype=None, temperature=None, abstain_below=0.0, use_graphs=None):
71
  """The prompt layout comes from decider_config.json: "layout": "chat" (chat-trained checkpoints) wraps every prompt in the
72
  tokenizer's chat template (decider.prompt.build_chat); no "layout" key is the plain layout of every earlier model.
73
- An unknown layout raises ValueError before the weights are loaded."""
 
 
 
 
 
74
  if device is None:
75
  device = "cuda" if torch.cuda.is_available() else ("mps" if torch.backends.mps.is_available() else "cpu")
76
  if dtype is None:
@@ -85,8 +95,7 @@ class Decider:
85
  except Exception:
86
  pass
87
  self.layout = resolve_layout(cfg)
88
- if temperature is None:
89
- temperature = float(cfg.get("temperature", 1.0))
90
  self.neutralize_none = bool(cfg.get("neutralize_none", True)) # v4 and earlier learned the literal string as an abstain signal
91
  if use_graphs is None:
92
  use_graphs = str(device).startswith("cuda")
@@ -102,7 +111,7 @@ class Decider:
102
  self.dev = device; self.T = temperature; self.abstain_below = abstain_below
103
  self.name = "decider-" + str(cfg.get("version", "dev"))
104
  self.schema_first = bool(cfg.get("schema_first", False)) and self.eng is not None # default layout. Questions-first (the cacheable one) costs accuracy
105
- self.T_schema = float(cfg.get("temperature_schema_first", temperature)) # (about 1.5 points on fixed label sets, more elsewhere): opt in with schema()
106
  self.isolated_levels = bool(cfg.get("isolated_levels", False)) # Score levels judged one per row (v8+)
107
  self._se = None; self._schemas = {}
108
 
@@ -116,19 +125,22 @@ class Decider:
116
  assert 2 <= len(q["options"]) <= MAX_OPTIONS, f"2..{MAX_OPTIONS} options required"
117
  exs.append(Example(context, [Q(q["question"], list(q["options"]), 0) for q in qs], "infer"))
118
  items = [build(e, self.m.tok, _NoShuffle(), max_options=MAX_OPTIONS, max_ctx_tokens=max_ctx_tokens, chat=self.chat) for e in exs]
 
 
119
  return requests, items
120
 
121
  @torch.no_grad()
122
  def decide_batch(self, requests, max_ctx_tokens=1536):
123
  """requests: list of (context:str, questions:list[dict(question, options)]). One forward pass for everything."""
124
  requests, items = self._decide_items(requests, max_ctx_tokens)
 
125
  if self.eng is not None:
126
- probs = torch.cat(self.eng.score_items(items, temperature=self.T))
127
  else:
128
  b = collate(items, self.m.tok.pad_token_id)
129
  logits = self.m.slot_logits(b["input_ids"].to(self.dev), b["attention_mask"].to(self.dev), b["slot_idx"].to(self.dev),
130
  b["slot_batch"].to(self.dev), b["nopts"].to(self.dev))
131
- probs = torch.softmax(logits / self.T, -1).cpu()
132
  out, k = [], 0
133
  for context, qs in requests:
134
  res = []
@@ -176,27 +188,33 @@ class Decider:
176
  (later questions can then see earlier question texts)."""
177
  from decider.systemone import unique_tokens, assemble
178
  rqs, index, items = self._system_one_items(state, questions, independent, max_state_tokens, layout, isolated)
179
- with torch.no_grad():
180
- if self.eng is not None and len(items) > 1 and layout == "state_first":
181
- probs = self.eng.score_shared(items, temperature=self.T)
182
- else:
183
- probs = []; per = max(1, max_fwd_tokens // max(len(it["ids"]) for it in items))
184
- for i in range(0, len(items), per):
185
- if self.eng is not None:
186
- probs += self.eng.score_items(items[i:i + per], temperature=self.T)
187
- else:
188
- bt = collate(items[i:i + per], self.m.tok.pad_token_id)
189
- lg = self.m.slot_logits(*[bt[k].to(self.dev) for k in ("input_ids", "attention_mask", "slot_idx", "slot_batch", "nopts")])
190
- pr = torch.softmax(lg / self.T, -1).cpu(); c = 0
191
- for it in items[i:i + per]:
192
- probs.append(pr[c:c + len(it["slots"])]); c += len(it["slots"])
193
- flatp = [p.tolist() for ps in probs for p in ps]
194
  return {"model": self.name, "answers": assemble(rqs, index, flatp),
195
  "usage": {"input_tokens": unique_tokens(items), "output_tokens": 0}}
196
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
197
  def _system_one_items(self, state, questions, independent=True, max_state_tokens=32768, layout=None, isolated=None):
198
  """The prompt rows of system_one's uncached path: -> (rendered questions, answer index, items)."""
199
- from decider.systemone import render_state, render_question, plan_rows
200
  layout = layout or ("schema_first" if self.schema_first else "state_first")
201
  isolated = (self.isolated_levels if isolated is None else isolated) and independent
202
  ctx = render_state(state); rqs = {k: render_question(v) for k, v in questions.items()}
@@ -205,6 +223,9 @@ class Decider:
205
  rows = [[r] for r in flat] if independent else [flat]
206
  items = [build(Example(ctx, [Q(r["question"], opts(r), 0) for r in row]), self.m.tok, _NoShuffle(), max_options=MAX_OPTIONS,
207
  max_ctx_tokens=max_state_tokens, layout=layout, chat=self.chat) for row in rows]
 
 
 
208
  return rqs, index, items
209
 
210
  # ---- typed schema interface: {question: {"type": "bool"} | {"type": "choice", "options": [...]}
@@ -257,14 +278,14 @@ class Decider:
257
  for qtext, spec in schema.items():
258
  t = spec.get("type", "choice")
259
  if t == "bool":
260
- qs.append(dict(question=qtext, options=["no", "yes"]))
261
  elif t == "choice":
262
- qs.append(dict(question=qtext, options=list(spec["options"])))
263
  elif t == "scale":
264
  leg = spec["legend"]
265
  keys = sorted(leg, key=lambda k: float(k)) if isinstance(leg, dict) else list(range(len(leg)))
266
  labels = [f"{k}: {leg[k]}" if isinstance(leg, dict) else f"{i}: {leg[i]}" for i, k in enumerate(keys)]
267
- qs.append(dict(question=qtext, options=labels, _keys=keys, _legend=leg))
268
  else:
269
  raise ValueError(f"unknown field type {t}")
270
  return qs
 
12
  import torch
13
  from decider.model import DecisionModel, collate
14
  from decider.prompt import build, MAX_OPTIONS, resolve_layout, chat_template
15
+ from decider import temperature as TT
16
  from dataclasses import dataclass
17
 
18
 
 
50
 
51
 
52
  class CompiledSchema:
53
+ def __init__(self, d, rqs, h, index):
54
+ from decider.systemone import row_types
55
+ self.d, self.rqs, self.h, self.index = d, rqs, h, index
56
+ self.types = row_types(rqs, index) # answer type of every schema row (decider.temperature)
57
 
58
  def batch(self, states, max_state_tokens=32768):
59
  from decider.systemone import render_state, assemble
60
+ T = TT.for_types(self.d.T_schema, self.d.T_schema_by_type, self.types)
61
+ probs = self.d._se.score(self.h, [render_state(s) for s in states], temperature=T, max_ctx_tokens=max_state_tokens)
62
  return [{"model": self.d.name, "answers": assemble(self.rqs, self.index, [p.tolist() for p in pr])} for pr in probs]
63
 
64
  def __call__(self, state, max_state_tokens=32768):
 
72
  the optional MPS patch; CPU defaults to bfloat16. Set ``use_graphs=False``
73
  for eager execution or debugging.
74
  """
75
+ def __init__(self, path, device=None, dtype=None, temperature=None, abstain_below=0.0, use_graphs=None, temperature_by_type=None):
76
  """The prompt layout comes from decider_config.json: "layout": "chat" (chat-trained checkpoints) wraps every prompt in the
77
  tokenizer's chat template (decider.prompt.build_chat); no "layout" key is the plain layout of every earlier model.
78
+ An unknown layout raises ValueError before the weights are loaded.
79
+
80
+ Temperatures (decider.temperature): decider_config.json "temperature" and the optional "temperature_by_type"
81
+ {"choice": T, "noul": T, "score": T}. temperature= overrides "temperature" and switches the config's by-type map off;
82
+ temperature_by_type= sets the map explicitly. An invalid temperature or map key raises ValueError before the weights
83
+ are loaded."""
84
  if device is None:
85
  device = "cuda" if torch.cuda.is_available() else ("mps" if torch.backends.mps.is_available() else "cpu")
86
  if dtype is None:
 
95
  except Exception:
96
  pass
97
  self.layout = resolve_layout(cfg)
98
+ (temperature, self.T_by_type), (T_schema, self.T_schema_by_type) = TT.from_config(cfg, temperature, temperature_by_type)
 
99
  self.neutralize_none = bool(cfg.get("neutralize_none", True)) # v4 and earlier learned the literal string as an abstain signal
100
  if use_graphs is None:
101
  use_graphs = str(device).startswith("cuda")
 
111
  self.dev = device; self.T = temperature; self.abstain_below = abstain_below
112
  self.name = "decider-" + str(cfg.get("version", "dev"))
113
  self.schema_first = bool(cfg.get("schema_first", False)) and self.eng is not None # default layout. Questions-first (the cacheable one) costs accuracy
114
+ self.T_schema = T_schema # (about 1.5 points on fixed label sets, more elsewhere): opt in with schema()
115
  self.isolated_levels = bool(cfg.get("isolated_levels", False)) # Score levels judged one per row (v8+)
116
  self._se = None; self._schemas = {}
117
 
 
125
  assert 2 <= len(q["options"]) <= MAX_OPTIONS, f"2..{MAX_OPTIONS} options required"
126
  exs.append(Example(context, [Q(q["question"], list(q["options"]), 0) for q in qs], "infer"))
127
  items = [build(e, self.m.tok, _NoShuffle(), max_options=MAX_OPTIONS, max_ctx_tokens=max_ctx_tokens, chat=self.chat) for e in exs]
128
+ for it, (_, qs) in zip(items, requests): # answer type per slot: a plain question is "choice"
129
+ it["types"] = [q.get("_type", "choice") for q in qs]
130
  return requests, items
131
 
132
  @torch.no_grad()
133
  def decide_batch(self, requests, max_ctx_tokens=1536):
134
  """requests: list of (context:str, questions:list[dict(question, options)]). One forward pass for everything."""
135
  requests, items = self._decide_items(requests, max_ctx_tokens)
136
+ T = TT.for_items(self.T, self.T_by_type, items) # the scalar self.T when there is no by-type map
137
  if self.eng is not None:
138
+ probs = torch.cat(self.eng.score_items(items, temperature=T))
139
  else:
140
  b = collate(items, self.m.tok.pad_token_id)
141
  logits = self.m.slot_logits(b["input_ids"].to(self.dev), b["attention_mask"].to(self.dev), b["slot_idx"].to(self.dev),
142
  b["slot_batch"].to(self.dev), b["nopts"].to(self.dev))
143
+ probs = TT.scaled_softmax(logits, TT.slot_temperatures(T, items)).cpu()
144
  out, k = [], 0
145
  for context, qs in requests:
146
  res = []
 
188
  (later questions can then see earlier question texts)."""
189
  from decider.systemone import unique_tokens, assemble
190
  rqs, index, items = self._system_one_items(state, questions, independent, max_state_tokens, layout, isolated)
191
+ flatp = self._system_one_probs(items, layout, max_fwd_tokens, TT.for_items(self.T, self.T_by_type, items))
 
 
 
 
 
 
 
 
 
 
 
 
 
 
192
  return {"model": self.name, "answers": assemble(rqs, index, flatp),
193
  "usage": {"input_tokens": unique_tokens(items), "output_tokens": 0}}
194
 
195
+ @torch.no_grad()
196
+ def _system_one_probs(self, items, layout, max_fwd_tokens=65536, temperature=None):
197
+ """Score system_one's uncached rows. temperature: as Engine.score_items takes it (a number, or one entry per item).
198
+ -> one probability list per question row, in row order."""
199
+ if self.eng is not None and len(items) > 1 and layout == "state_first":
200
+ probs = self.eng.score_shared(items, temperature=temperature)
201
+ else:
202
+ probs = []; per = max(1, max_fwd_tokens // max(len(it["ids"]) for it in items))
203
+ for i in range(0, len(items), per):
204
+ T = TT.item_slice(temperature, i, i + per)
205
+ if self.eng is not None:
206
+ probs += self.eng.score_items(items[i:i + per], temperature=T)
207
+ else:
208
+ bt = collate(items[i:i + per], self.m.tok.pad_token_id)
209
+ lg = self.m.slot_logits(*[bt[k].to(self.dev) for k in ("input_ids", "attention_mask", "slot_idx", "slot_batch", "nopts")])
210
+ pr = TT.scaled_softmax(lg, TT.slot_temperatures(T, items[i:i + per])).cpu(); c = 0
211
+ for it in items[i:i + per]:
212
+ probs.append(pr[c:c + len(it["slots"])]); c += len(it["slots"])
213
+ return [p.tolist() for ps in probs for p in ps]
214
+
215
  def _system_one_items(self, state, questions, independent=True, max_state_tokens=32768, layout=None, isolated=None):
216
  """The prompt rows of system_one's uncached path: -> (rendered questions, answer index, items)."""
217
+ from decider.systemone import render_state, render_question, plan_rows, row_types
218
  layout = layout or ("schema_first" if self.schema_first else "state_first")
219
  isolated = (self.isolated_levels if isolated is None else isolated) and independent
220
  ctx = render_state(state); rqs = {k: render_question(v) for k, v in questions.items()}
 
223
  rows = [[r] for r in flat] if independent else [flat]
224
  items = [build(Example(ctx, [Q(r["question"], opts(r), 0) for r in row]), self.m.tok, _NoShuffle(), max_options=MAX_OPTIONS,
225
  max_ctx_tokens=max_state_tokens, layout=layout, chat=self.chat) for row in rows]
226
+ types = row_types(rqs, index) # answer type per row (decider.temperature)
227
+ for it, ts in zip(items, [[t] for t in types] if independent else [types]):
228
+ it["types"] = ts
229
  return rqs, index, items
230
 
231
  # ---- typed schema interface: {question: {"type": "bool"} | {"type": "choice", "options": [...]}
 
278
  for qtext, spec in schema.items():
279
  t = spec.get("type", "choice")
280
  if t == "bool":
281
+ qs.append(dict(question=qtext, options=["no", "yes"], _type="noul"))
282
  elif t == "choice":
283
+ qs.append(dict(question=qtext, options=list(spec["options"]), _type="choice"))
284
  elif t == "scale":
285
  leg = spec["legend"]
286
  keys = sorted(leg, key=lambda k: float(k)) if isinstance(leg, dict) else list(range(len(leg)))
287
  labels = [f"{k}: {leg[k]}" if isinstance(leg, dict) else f"{i}: {leg[i]}" for i, k in enumerate(keys)]
288
+ qs.append(dict(question=qtext, options=labels, _keys=keys, _legend=leg, _type="score"))
289
  else:
290
  raise ValueError(f"unknown field type {t}")
291
  return qs
decider/schema_engine.py CHANGED
@@ -124,12 +124,15 @@ class SchemaEngine:
124
  return next((t for t in TS_BUCKETS if t >= n_tokens), -(-n_tokens // 256) * 256)
125
 
126
  def score(self, h, contexts, temperature=1.0, max_ctx_tokens=1536):
127
- """-> one [n_questions, MAX_OPTIONS] probability tensor per context."""
128
  return self.score_rows(h, [self.tokenize(h, c, max_ctx_tokens) for c in contexts], temperature)
129
 
130
  @torch.no_grad()
131
  def score_rows(self, h, rows, temperature=1.0):
132
- """rows: [(suffix ids, slots)] from tokenize()."""
 
 
 
133
  Tmax = max(len(r[0]) for r in rows); Ts = next((t for t in TS_BUCKETS if t >= Tmax), None); n = len(rows)
134
  R = next((b for b in B_BUCKETS if b >= n), n) if Ts else n; Ts = Ts or -(-Tmax // 256) * 256
135
  ids = fill_ids([x for x, _ in rows for _ in range(h.P)], R * h.P, Ts, self.tok.pad_token_id).to(self.dev, non_blocking=True)
@@ -143,4 +146,5 @@ class SchemaEngine:
143
  rws = [r for r in range(n) for _ in range(h.nq)]; sls = [x for _, sl in rows for x in sl]
144
  else: # independent: one slot in each of the request's P rows
145
  rws = [r * h.P + p for r in range(n) for p in range(h.P)]; sls = [sl[0] for _, sl in rows for _ in range(h.P)]
146
- return read_slots(out, rws, sls, h.nopts * n, temperature, [h.nq] * n)
 
 
124
  return next((t for t in TS_BUCKETS if t >= n_tokens), -(-n_tokens // 256) * 256)
125
 
126
  def score(self, h, contexts, temperature=1.0, max_ctx_tokens=1536):
127
+ """-> one [n_questions, MAX_OPTIONS] probability tensor per context. temperature: as in score_rows."""
128
  return self.score_rows(h, [self.tokenize(h, c, max_ctx_tokens) for c in contexts], temperature)
129
 
130
  @torch.no_grad()
131
  def score_rows(self, h, rows, temperature=1.0):
132
+ """rows: [(suffix ids, slots)] from tokenize(). temperature: a number, or a list with one temperature per schema row
133
+ (h.nq values, in the order of prepare's questions), applied to every request."""
134
+ if isinstance(temperature, (list, tuple)) and len(temperature) != h.nq:
135
+ raise ValueError(f"temperature: {len(temperature)} values for a schema with {h.nq} rows")
136
  Tmax = max(len(r[0]) for r in rows); Ts = next((t for t in TS_BUCKETS if t >= Tmax), None); n = len(rows)
137
  R = next((b for b in B_BUCKETS if b >= n), n) if Ts else n; Ts = Ts or -(-Tmax // 256) * 256
138
  ids = fill_ids([x for x, _ in rows for _ in range(h.P)], R * h.P, Ts, self.tok.pad_token_id).to(self.dev, non_blocking=True)
 
146
  rws = [r for r in range(n) for _ in range(h.nq)]; sls = [x for _, sl in rows for x in sl]
147
  else: # independent: one slot in each of the request's P rows
148
  rws = [r * h.P + p for r in range(n) for p in range(h.P)]; sls = [sl[0] for _, sl in rows for _ in range(h.P)]
149
+ temps = list(temperature) * n if isinstance(temperature, (list, tuple)) else temperature
150
+ return read_slots(out, rws, sls, h.nopts * n, temps, [h.nq] * n)
decider/serve.py CHANGED
@@ -40,6 +40,12 @@ Variables (default): DECIDER_DEVICE (auto: cuda, else mps, else cpu) DECIDER_C
40
  DECIDER_T_BUCKETS DECIDER_B_BUCKETS DECIDER_GRAPH_TOKEN_BUDGET (32768) DECIDER_WARMUP (1) DECIDER_TOKENIZE_THREADS (8)
41
  DECIDER_MAX_ROWS (1024) DECIDER_MAX_ROW_TOKENS (DECIDER_MAX_STATE_TOKENS + 4096) DECIDER_MAX_REQUEST_TOKENS (1048576)
42
  DECIDER_MAX_QUEUE_ROWS (4096) DECIDER_TEMPERATURE DECIDER_SCHEMA_CACHE (0) DECIDER_SCHEMA_MIN_SEEN (2) DECIDER_SCHEMAS
 
 
 
 
 
 
43
  """
44
  import asyncio, json, os, time
45
  from concurrent.futures import ThreadPoolExecutor
@@ -47,6 +53,7 @@ from contextlib import asynccontextmanager
47
  from fastapi import FastAPI, HTTPException
48
  from pydantic import BaseModel
49
  from decider import systemone as S1
 
50
  from decider.batching import DEFAULT_MERGE_OVERHEAD_TOKENS, plan_batches
51
  from decider.prompt import build, MAX_OPTIONS, resolve_layout, chat_template
52
  from decider.prompt_fast import build_rows, unique_tokens
@@ -78,7 +85,7 @@ MAX_REQUEST_TOKENS = _env_int("DECIDER_MAX_REQUEST_TOKENS", 1 << 20) # sum
78
  MAX_QUEUE_ROWS = _env_int("DECIDER_MAX_QUEUE_ROWS", 4096) # rows admitted and not yet scored, over all requests
79
  DECIDE_MAX_CTX_TOKENS = 1536 # /decide context cap, unchanged from 1.0.x
80
 
81
- MODEL_NAME = "decider"; TEMP = 1.0; TEMP_SCHEMA = 1.0; RELEASE_DATE = "2026-09-17"; ISOLATED = False; NEUTRALIZE_NONE = True
82
  SCHEMA_FIRST = False; LAYOUT = "plain"; CHAT = None; se = None; squeue = None; schemas = {}; seen = {}
83
  eng = None; queue = None; gpu = None; cpu = None; batcher_task = None; schema_task = None
84
  outstanding = 0; REQ_SEQ = 0
@@ -118,6 +125,9 @@ def prepare(tok, state, questions, independent, isolated=False, max_state_tokens
118
  pairs = [(r["question"], list(r["options"])) for r in flat]
119
  rows = [[p] for p in pairs] if independent else [pairs]
120
  items, ctx_len = build_rows(tok, ctx, rows, max_ctx_tokens=max_state_tokens, chat=chat)
 
 
 
121
  return rqs, index, items, ctx_len
122
 
123
 
@@ -133,6 +143,7 @@ def _prepare_decide(context, schema):
133
  q["options"], q["_back"] = neutralize_options(q["options"])
134
  ex = Example(context, [Q(q["question"], list(q["options"]), 0) for q in qs])
135
  it = build(ex, eng.tok, _NoShuffle(), max_options=MAX_OPTIONS, max_ctx_tokens=DECIDE_MAX_CTX_TOKENS, chat=CHAT)
 
136
  return qs, it
137
 
138
 
@@ -188,11 +199,11 @@ def _release(n):
188
 
189
  # ---- GPU work ----------------------------------------------------------------
190
  def _score_items(items):
191
- return eng.score_items(items, temperature=TEMP)
192
 
193
 
194
  def _score_shared(items):
195
- return eng.score_shared(items, temperature=TEMP)
196
 
197
 
198
  async def _collect(q, wait_ms=None, adaptive_ms=0.0, idle_reset_ms=None):
@@ -294,6 +305,7 @@ def _schema_handle(questions, independent, compile=False):
294
  old = next(iter(schemas)); hid = schemas.pop(old)[1].id
295
  for k in [k for k in se.graphs if k[0] == hid]: del se.graphs[k]
296
  h = se.prepare(rows, independent=independent, compile=compile)
 
297
  schemas[key] = (rqs, h, index)
298
  return schemas[key]
299
 
@@ -308,7 +320,7 @@ def _worth_caching(questions, independent):
308
 
309
 
310
  def _score_schema(h, rows):
311
- return se.score_rows(h, rows, temperature=TEMP_SCHEMA)
312
 
313
 
314
  async def schema_batcher():
@@ -333,12 +345,13 @@ async def schema_batcher():
333
 
334
  # ---- start-up / shutdown -------------------------------------------------------
335
  def apply_config(cfg):
336
- global MODEL_NAME, TEMP, TEMP_SCHEMA, RELEASE_DATE, ISOLATED, NEUTRALIZE_NONE, SCHEMA_FIRST, LAYOUT
337
  LAYOUT = resolve_layout(cfg) # ValueError for an unknown layout, before the engine is built
338
  NEUTRALIZE_NONE = bool(cfg.get("neutralize_none", True))
339
  MODEL_NAME = "decider-" + str(cfg.get("version", "dev"))
340
- TEMP = float(os.environ.get("DECIDER_TEMPERATURE", cfg.get("temperature", 1.0)))
341
- TEMP_SCHEMA = float(cfg.get("temperature_schema_first", TEMP))
 
342
  RELEASE_DATE = str(cfg.get("release_date", RELEASE_DATE))
343
  ISOLATED = bool(cfg.get("isolated_levels", False))
344
  trained = bool(cfg.get("schema_first", False) or cfg.get("schema_first_trained", False))
@@ -403,7 +416,9 @@ async def _start():
403
  t = await loop.run_in_executor(gpu, lambda: eng.warmup(log=lambda s: print(s, flush=True)))
404
  print(f"[serve] captured {len(eng.graphs)} graphs in {t:.0f}s", flush=True)
405
  eng.seal()
406
- print("[serve] ready", json.dumps(dict(model=MODEL_NAME, layout=LAYOUT, temperature=TEMP, isolated_levels=ISOLATED, schema_first=SCHEMA_FIRST,
 
 
407
  shared=SHARED, graphs=len(eng.graphs), limits=dict(
408
  max_rows=MAX_ROWS, max_row_tokens=MAX_ROW_TOKENS, max_request_tokens=MAX_REQUEST_TOKENS,
409
  max_queue_rows=MAX_QUEUE_ROWS))), flush=True)
@@ -535,7 +550,9 @@ async def models():
535
  @app.get("/health")
536
  async def health():
537
  return {"ok": eng is not None and bool(getattr(eng, "sealed", True)) and _alive(), "model": MODEL,
538
- "device": str(getattr(eng, "dev", "")) if eng is not None else None, "layout": LAYOUT}
 
 
539
 
540
 
541
  @app.get("/stats")
 
40
  DECIDER_T_BUCKETS DECIDER_B_BUCKETS DECIDER_GRAPH_TOKEN_BUDGET (32768) DECIDER_WARMUP (1) DECIDER_TOKENIZE_THREADS (8)
41
  DECIDER_MAX_ROWS (1024) DECIDER_MAX_ROW_TOKENS (DECIDER_MAX_STATE_TOKENS + 4096) DECIDER_MAX_REQUEST_TOKENS (1048576)
42
  DECIDER_MAX_QUEUE_ROWS (4096) DECIDER_TEMPERATURE DECIDER_SCHEMA_CACHE (0) DECIDER_SCHEMA_MIN_SEEN (2) DECIDER_SCHEMAS
43
+
44
+ Temperatures (1.4.0, decider.temperature): decider_config.json "temperature", and optionally "temperature_by_type"
45
+ {"choice": T, "noul": T, "score": T} (a /decide "bool" field is "noul", a "scale" field is "score"; a missing type uses
46
+ "temperature"). Every row carries the answer type of each of its slots, so a batch that mixes requests and types still
47
+ applies each answer's own temperature. DECIDER_TEMPERATURE replaces "temperature" and switches the by-type map off.
48
+ /health and the ready line report the temperature every answer type gets.
49
  """
50
  import asyncio, json, os, time
51
  from concurrent.futures import ThreadPoolExecutor
 
53
  from fastapi import FastAPI, HTTPException
54
  from pydantic import BaseModel
55
  from decider import systemone as S1
56
+ from decider import temperature as TT
57
  from decider.batching import DEFAULT_MERGE_OVERHEAD_TOKENS, plan_batches
58
  from decider.prompt import build, MAX_OPTIONS, resolve_layout, chat_template
59
  from decider.prompt_fast import build_rows, unique_tokens
 
85
  MAX_QUEUE_ROWS = _env_int("DECIDER_MAX_QUEUE_ROWS", 4096) # rows admitted and not yet scored, over all requests
86
  DECIDE_MAX_CTX_TOKENS = 1536 # /decide context cap, unchanged from 1.0.x
87
 
88
+ MODEL_NAME = "decider"; TEMP = 1.0; TEMP_SCHEMA = 1.0; TEMP_BY_TYPE = {}; TEMP_SCHEMA_BY_TYPE = {}; RELEASE_DATE = "2026-09-17"; ISOLATED = False; NEUTRALIZE_NONE = True
89
  SCHEMA_FIRST = False; LAYOUT = "plain"; CHAT = None; se = None; squeue = None; schemas = {}; seen = {}
90
  eng = None; queue = None; gpu = None; cpu = None; batcher_task = None; schema_task = None
91
  outstanding = 0; REQ_SEQ = 0
 
125
  pairs = [(r["question"], list(r["options"])) for r in flat]
126
  rows = [[p] for p in pairs] if independent else [pairs]
127
  items, ctx_len = build_rows(tok, ctx, rows, max_ctx_tokens=max_state_tokens, chat=chat)
128
+ types = S1.row_types(rqs, index) # answer type per slot (decider.temperature)
129
+ for it, ts in zip(items, [[t] for t in types] if independent else [types]):
130
+ it["types"] = ts
131
  return rqs, index, items, ctx_len
132
 
133
 
 
143
  q["options"], q["_back"] = neutralize_options(q["options"])
144
  ex = Example(context, [Q(q["question"], list(q["options"]), 0) for q in qs])
145
  it = build(ex, eng.tok, _NoShuffle(), max_options=MAX_OPTIONS, max_ctx_tokens=DECIDE_MAX_CTX_TOKENS, chat=CHAT)
146
+ it["types"] = [q["_type"] for q in qs] # bool -> noul, scale -> score (decider.temperature)
147
  return qs, it
148
 
149
 
 
199
 
200
  # ---- GPU work ----------------------------------------------------------------
201
  def _score_items(items):
202
+ return eng.score_items(items, temperature=TT.for_items(TEMP, TEMP_BY_TYPE, items)) # the scalar TEMP without a by-type map
203
 
204
 
205
  def _score_shared(items):
206
+ return eng.score_shared(items, temperature=TT.for_items(TEMP, TEMP_BY_TYPE, items))
207
 
208
 
209
  async def _collect(q, wait_ms=None, adaptive_ms=0.0, idle_reset_ms=None):
 
305
  old = next(iter(schemas)); hid = schemas.pop(old)[1].id
306
  for k in [k for k in se.graphs if k[0] == hid]: del se.graphs[k]
307
  h = se.prepare(rows, independent=independent, compile=compile)
308
+ h.types = S1.row_types(rqs, index) # answer type per schema row (decider.temperature)
309
  schemas[key] = (rqs, h, index)
310
  return schemas[key]
311
 
 
320
 
321
 
322
  def _score_schema(h, rows):
323
+ return se.score_rows(h, rows, temperature=TT.for_types(TEMP_SCHEMA, TEMP_SCHEMA_BY_TYPE, h.types))
324
 
325
 
326
  async def schema_batcher():
 
345
 
346
  # ---- start-up / shutdown -------------------------------------------------------
347
  def apply_config(cfg):
348
+ global MODEL_NAME, TEMP, TEMP_SCHEMA, TEMP_BY_TYPE, TEMP_SCHEMA_BY_TYPE, RELEASE_DATE, ISOLATED, NEUTRALIZE_NONE, SCHEMA_FIRST, LAYOUT
349
  LAYOUT = resolve_layout(cfg) # ValueError for an unknown layout, before the engine is built
350
  NEUTRALIZE_NONE = bool(cfg.get("neutralize_none", True))
351
  MODEL_NAME = "decider-" + str(cfg.get("version", "dev"))
352
+ # ValueError for a temperature that is not finite and > 0 or an unknown by-type key, before the engine is built.
353
+ # DECIDER_TEMPERATURE replaces "temperature" and switches "temperature_by_type" off (decider.temperature).
354
+ (TEMP, TEMP_BY_TYPE), (TEMP_SCHEMA, TEMP_SCHEMA_BY_TYPE) = TT.from_config(cfg, os.environ.get("DECIDER_TEMPERATURE"))
355
  RELEASE_DATE = str(cfg.get("release_date", RELEASE_DATE))
356
  ISOLATED = bool(cfg.get("isolated_levels", False))
357
  trained = bool(cfg.get("schema_first", False) or cfg.get("schema_first_trained", False))
 
416
  t = await loop.run_in_executor(gpu, lambda: eng.warmup(log=lambda s: print(s, flush=True)))
417
  print(f"[serve] captured {len(eng.graphs)} graphs in {t:.0f}s", flush=True)
418
  eng.seal()
419
+ print("[serve] ready", json.dumps(dict(model=MODEL_NAME, layout=LAYOUT, temperature=TEMP, temperature_by_type=TT.effective(TEMP, TEMP_BY_TYPE),
420
+ **({"temperature_schema_first_by_type": TT.effective(TEMP_SCHEMA, TEMP_SCHEMA_BY_TYPE)} if SCHEMA_FIRST else {}),
421
+ isolated_levels=ISOLATED, schema_first=SCHEMA_FIRST,
422
  shared=SHARED, graphs=len(eng.graphs), limits=dict(
423
  max_rows=MAX_ROWS, max_row_tokens=MAX_ROW_TOKENS, max_request_tokens=MAX_REQUEST_TOKENS,
424
  max_queue_rows=MAX_QUEUE_ROWS))), flush=True)
 
550
  @app.get("/health")
551
  async def health():
552
  return {"ok": eng is not None and bool(getattr(eng, "sealed", True)) and _alive(), "model": MODEL,
553
+ "device": str(getattr(eng, "dev", "")) if eng is not None else None, "layout": LAYOUT,
554
+ "temperature": TEMP, "temperature_by_type": TT.effective(TEMP, TEMP_BY_TYPE),
555
+ **({"temperature_schema_first_by_type": TT.effective(TEMP_SCHEMA, TEMP_SCHEMA_BY_TYPE)} if SCHEMA_FIRST else {})}
556
 
557
 
558
  @app.get("/stats")
decider/shared_prefix.py CHANGED
@@ -33,6 +33,7 @@ import torch
33
  import torch.nn.functional as F
34
 
35
  from decider.engine import fill_ids, read_slots
 
36
 
37
  DEFAULT_FORK_GB = 8.0
38
  STATE_NAMES = ("keys", "values", "indexer_keys", "conv_states", "recurrent_states")
@@ -159,6 +160,7 @@ def score_shared(engine, items, temperature=1.0, min_prefix=192, budget_bytes=No
159
  request does not qualify (fewer than two rows, or a common prefix below `min_prefix`) and the caller should use
160
  `score_items`.
161
 
 
162
  `rows_per_fork` forces the chunk size; it exists for the tests that compare chunked against unchunked answers."""
163
  ids = [it["ids"] for it in items]
164
  n = len(ids)
@@ -167,6 +169,7 @@ def score_shared(engine, items, temperature=1.0, min_prefix=192, budget_bytes=No
167
  lcp = common_prefix_len(ids)
168
  if lcp < min_prefix:
169
  return None
 
170
  core, W, dev, pad = engine.core, engine.W, engine.dev, engine.tok.pad_token_id
171
  pre = torch.tensor(ids[0][:lcp], device=dev)[None]
172
  cache = core(input_ids=pre, use_cache=True).past_key_values
@@ -191,6 +194,7 @@ def score_shared(engine, items, temperature=1.0, min_prefix=192, budget_bytes=No
191
  sl = [s - lcp for it in part for s in it["slots"]]
192
  idx = torch.tensor([rows, sl], device=dev)
193
  out += read_slots(F.linear(h[idx[0], idx[1]], W).float()[:, None, :], list(range(len(rows))), [0] * len(rows),
194
- [k for it in part for k in it["nopts"]], temperature, [len(it["slots"]) for it in part])
 
195
  del fork, h, suf # drop this chunk's fork before the next one is built
196
  return out
 
33
  import torch.nn.functional as F
34
 
35
  from decider.engine import fill_ids, read_slots
36
+ from decider.temperature import item_slice, slot_temperatures
37
 
38
  DEFAULT_FORK_GB = 8.0
39
  STATE_NAMES = ("keys", "values", "indexer_keys", "conv_states", "recurrent_states")
 
160
  request does not qualify (fewer than two rows, or a common prefix below `min_prefix`) and the caller should use
161
  `score_items`.
162
 
163
+ `temperature`: a number, or one entry per item (decider.temperature.slot_temperatures).
164
  `rows_per_fork` forces the chunk size; it exists for the tests that compare chunked against unchunked answers."""
165
  ids = [it["ids"] for it in items]
166
  n = len(ids)
 
169
  lcp = common_prefix_len(ids)
170
  if lcp < min_prefix:
171
  return None
172
+ slot_temperatures(temperature, items) # a length mismatch fails before any forward
173
  core, W, dev, pad = engine.core, engine.W, engine.dev, engine.tok.pad_token_id
174
  pre = torch.tensor(ids[0][:lcp], device=dev)[None]
175
  cache = core(input_ids=pre, use_cache=True).past_key_values
 
194
  sl = [s - lcp for it in part for s in it["slots"]]
195
  idx = torch.tensor([rows, sl], device=dev)
196
  out += read_slots(F.linear(h[idx[0], idx[1]], W).float()[:, None, :], list(range(len(rows))), [0] * len(rows),
197
+ [k for it in part for k in it["nopts"]],
198
+ slot_temperatures(item_slice(temperature, i, i + b), part), [len(it["slots"]) for it in part])
199
  del fork, h, suf # drop this chunk's fork before the next one is built
200
  return out
decider/systemone.py CHANGED
@@ -3,7 +3,8 @@
3
  state str | dict | list JSON state is serialised compactly; questions may name a part by path (`ticket.messages[0].text`)
4
  questions {id: {"type": "choice", "instructions": ..., "criteria": {name: description | {...} | [...] | None}} up to 255 options
5
  {"type": "score", "instructions": ..., "criteria": [level 0 description, level 1 description, ...]} 2..10 levels
6
- {"type": "noul", "instructions": ..., "criteria": {"true": ..., "false": ...} (optional)}}
 
7
  ids are never shown to the model. `instructions` and every description may be a string or any JSON value.
8
  """
9
  import json, math
@@ -38,7 +39,15 @@ def render_state(state, index_arrays=True):
38
 
39
  def render_question(spec):
40
  """-> dict(question=str, options=[str], type=..., names=[...]) (names: what the answer reports for each option)"""
41
- t = spec.get("type", "choice"); ins = _txt(spec.get("instructions", spec.get("question", ""))); crit = spec.get("criteria", spec.get("options"))
 
 
 
 
 
 
 
 
42
  if not ins:
43
  raise ValueError("question without instructions")
44
  if t == "choice":
@@ -65,6 +74,12 @@ def render_question(spec):
65
  isolated=bool(spec.get("isolated", True)))
66
 
67
 
 
 
 
 
 
 
68
  # ---- isolated levels: every Score level is judged in its own row, without its number or its neighbours
69
  ISOLATED = "{q}\nProposed answer: {level}\nDoes the proposed answer fit?"
70
  _NUM = None
@@ -100,6 +115,15 @@ def plan_rows(rqs, isolated=True):
100
  return rows, index
101
 
102
 
 
 
 
 
 
 
 
 
 
103
  def assemble(rqs, index, probs):
104
  """probs: one probability list per row (plan_rows order) -> {id: answer}."""
105
  out = {}
@@ -118,16 +142,50 @@ def certainty(p):
118
  return max(0.0, 1.0 - h / math.log(len(p))) if len(p) > 1 else 1.0
119
 
120
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
121
  def format_answer(rq, p, nd=4):
122
- """rq: render_question output; p: probabilities in option order."""
 
 
123
  p = [float(x) for x in p[:len(rq["options"])]]; s = sum(p) or 1.0; p = [x / s for x in p]
124
  j = max(range(len(p)), key=p.__getitem__)
125
  if rq["type"] == "noul":
126
  return {"type": "noul", "noul": round(p[1], nd)}
127
  if rq["type"] == "choice":
128
- return {"type": "choice", "choice": rq["names"][j], "confidence": round(p[j], nd), "certainty": round(certainty(p), nd),
129
- "probabilities": {n: round(x, nd) for n, x in zip(rq["names"], p)}}
130
- return {"type": "score", "score": round(sum(i * x for i, x in enumerate(p)), 2), "confidence": round(p[j], nd), "certainty": round(certainty(p), nd),
 
131
  "legend": {str(i): d for i, d in enumerate(rq["legend"])}, "probabilities": {str(i): round(x, nd) for i, x in enumerate(p)}}
132
 
133
 
 
3
  state str | dict | list JSON state is serialised compactly; questions may name a part by path (`ticket.messages[0].text`)
4
  questions {id: {"type": "choice", "instructions": ..., "criteria": {name: description | {...} | [...] | None}} up to 255 options
5
  {"type": "score", "instructions": ..., "criteria": [level 0 description, level 1 description, ...]} 2..10 levels
6
+ {"type": "noul", "instructions": ... (optional), "criteria": {"true": ..., "false": ...} (optional)}}
7
+ a noul question needs instructions or at least one true/false description.
8
  ids are never shown to the model. `instructions` and every description may be a string or any JSON value.
9
  """
10
  import json, math
 
39
 
40
  def render_question(spec):
41
  """-> dict(question=str, options=[str], type=..., names=[...]) (names: what the answer reports for each option)"""
42
+ t = spec.get("type", "choice"); crit = spec.get("criteria", spec.get("options"))
43
+ raw = spec.get("instructions", spec.get("question", ""))
44
+ if t in ("noul", "bool") and raw in (None, ""): # the criteria carry the question (NOUL_WITHOUT_INSTRUCTIONS)
45
+ ins = NOUL_WITHOUT_INSTRUCTIONS
46
+ described = isinstance(crit, dict) and any(d not in (None, "") for d in (crit.get("true", crit.get(True)), crit.get("false", crit.get(False))))
47
+ if not described and (crit is None or isinstance(crit, dict)): # a non-map is rejected below with the criteria message
48
+ raise ValueError("noul question without instructions: criteria must describe true or false")
49
+ else:
50
+ ins = _txt(raw)
51
  if not ins:
52
  raise ValueError("question without instructions")
53
  if t == "choice":
 
74
  isolated=bool(spec.get("isolated", True)))
75
 
76
 
77
+ # A noul question may omit `instructions` (TypeSafe's OpenAPI file marks it optional). The question id is never shown to the
78
+ # model, so the question text is this fixed sentence and the true/false descriptions, rendered as the options "no: ..." and
79
+ # "yes: ...", say what is being asked. A request that gives instructions is rendered exactly as before.
80
+ NOUL_WITHOUT_INSTRUCTIONS = "Which answer fits the context?"
81
+
82
+
83
  # ---- isolated levels: every Score level is judged in its own row, without its number or its neighbours
84
  ISOLATED = "{q}\nProposed answer: {level}\nDoes the proposed answer fit?"
85
  _NUM = None
 
115
  return rows, index
116
 
117
 
118
+ def row_types(rqs, index):
119
+ """The answer type ("choice", "noul" or "score") of every plan_rows row, in row order. An isolated Score question's
120
+ yes/no level rows carry "score": together they are one Score answer (decider.temperature)."""
121
+ types = [None] * sum(n for _, _, _, n in index)
122
+ for k, _, s, n in index:
123
+ types[s:s + n] = [rqs[k]["type"]] * n
124
+ return types
125
+
126
+
127
  def assemble(rqs, index, probs):
128
  """probs: one probability list per row (plan_rows order) -> {id: answer}."""
129
  out = {}
 
142
  return max(0.0, 1.0 - h / math.log(len(p))) if len(p) > 1 else 1.0
143
 
144
 
145
+ def _clip01(x):
146
+ return min(1.0, max(0.0, x))
147
+
148
+
149
+ def _normalised(p):
150
+ """As the adapter's _normalize: a distribution with zero total counts as uniform."""
151
+ tot = sum(p)
152
+ return [1.0 / len(p)] * len(p) if tot == 0 else [x / tot for x in p]
153
+
154
+
155
+ def choice_confidence(p):
156
+ """TypeSafe's Choice confidence: the largest probability rescaled so that a uniform distribution gives 0 and all mass on one
157
+ option gives 1, (n * p_max - 1) / (n - 1); 1 for a single option."""
158
+ n = len(p); p = _normalised(p)
159
+ return 1.0 if n <= 1 else _clip01((n * max(p) - 1) / (n - 1))
160
+
161
+
162
+ def score_confidence(p):
163
+ """TypeSafe's Score confidence (system-one-adapter-python, confidence_metrics.score_confidence): 1 minus the expected distance
164
+ from the most likely level, divided by D = mean over levels i of |i - (n - 1)/2| (the mean distance of the levels from the
165
+ middle of the scale), floored at 0; 1 for a single level.
166
+ With two levels it equals choice_confidence."""
167
+ n = len(p); p = _normalised(p)
168
+ if n <= 1:
169
+ return 1.0
170
+ k = max(range(n), key=p.__getitem__)
171
+ spread = sum(x * abs(i - k) for i, x in enumerate(p))
172
+ uniform = sum(abs(i - (n - 1) / 2) for i in range(n)) / n
173
+ return _clip01(1.0 - spread / uniform)
174
+
175
+
176
  def format_answer(rq, p, nd=4):
177
+ """rq: render_question output; p: probabilities in option order.
178
+ `confidence` is TypeSafe's (choice_confidence / score_confidence); `x_p_max` is the largest probability, which was
179
+ `confidence` before 1.3.0."""
180
  p = [float(x) for x in p[:len(rq["options"])]]; s = sum(p) or 1.0; p = [x / s for x in p]
181
  j = max(range(len(p)), key=p.__getitem__)
182
  if rq["type"] == "noul":
183
  return {"type": "noul", "noul": round(p[1], nd)}
184
  if rq["type"] == "choice":
185
+ return {"type": "choice", "choice": rq["names"][j], "confidence": round(choice_confidence(p), nd), "x_p_max": round(p[j], nd),
186
+ "certainty": round(certainty(p), nd), "probabilities": {n: round(x, nd) for n, x in zip(rq["names"], p)}}
187
+ return {"type": "score", "score": round(sum(i * x for i, x in enumerate(p)), 2), "confidence": round(score_confidence(p), nd), "x_p_max": round(p[j], nd),
188
+ "certainty": round(certainty(p), nd),
189
  "legend": {str(i): d for i, d in enumerate(rq["legend"])}, "probabilities": {str(i): round(x, nd) for i, x in enumerate(p)}}
190
 
191
 
decider/temperature.py ADDED
@@ -0,0 +1,144 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Answer temperatures from decider_config.json (1.4.0).
2
+
3
+ Every answer is softmax(logits / T) over its option letters. decider_config.json sets T:
4
+
5
+ "temperature": 1.3 one value for every answer (the only form before 1.4.0)
6
+ "temperature_by_type": {"choice": 1.48, "noul": 2.22, "score": 1.38}
7
+ optional; one value per answer type, a missing type uses "temperature"
8
+ "temperature_schema_first": 1.18 optional; the schema cache (questions-first layout), as before
9
+ "temperature_schema_first_by_type": {...} optional; per answer type on the schema cache
10
+
11
+ The answer types are the /v1/systemone question types. A /decide field maps onto them: "choice" -> "choice", "bool" -> "noul",
12
+ "scale" -> "score". Plain `Decider.decide` questions (a question and its options, no type) are "choice". A Score question
13
+ read with isolated levels (one yes/no row per level) uses the "score" temperature on every one of its level rows: the rows form
14
+ one Score answer, and a temperature fitted on Score answers is fitted through that same readout (decider.calibrate).
15
+
16
+ On the state-first layout the temperature of an answer of type t is temperature_by_type[t], else temperature. On the schema
17
+ cache it is temperature_schema_first_by_type[t], else temperature_schema_first, else (no schema-first value at all) the
18
+ state-first temperature of t. An explicit override (Decider(temperature=...), DECIDER_TEMPERATURE) replaces the state-first
19
+ "temperature" and switches temperature_by_type off, so the override is the one temperature of every state-first answer, as it
20
+ was before 1.4.0.
21
+
22
+ A config without a by-type map gives every path the single number it gave in 1.3.0, and that number reaches the engines as the
23
+ same Python float, so the probabilities are bit-identical to 1.3.0.
24
+ """
25
+ import math
26
+
27
+ TYPES = ("choice", "noul", "score")
28
+ FIELD_TYPES = {"choice": "choice", "bool": "noul", "scale": "score"} # /decide field type -> answer type
29
+ _KEYS_TEXT = ('the keys are "choice", "noul" and "score" (a /decide "bool" field is "noul", a "scale" field is "score"; '
30
+ 'a /v1/systemone "bool" question is "noul")')
31
+
32
+
33
+ def positive(value, where):
34
+ """A temperature: a number (or a numeric string, as float() reads it, which 1.3.0 accepted for "temperature") that is finite
35
+ and > 0. Raises ValueError naming `where`."""
36
+ if isinstance(value, bool):
37
+ raise ValueError(f"{where} must be a finite number > 0, got {value!r}")
38
+ try:
39
+ v = float(value)
40
+ except (TypeError, ValueError):
41
+ raise ValueError(f"{where} must be a finite number > 0, got {value!r}") from None
42
+ if not math.isfinite(v) or v <= 0:
43
+ raise ValueError(f"{where} must be a finite number > 0, got {value!r}")
44
+ return v
45
+
46
+
47
+ def by_type(m, where):
48
+ """Validate a {answer type: temperature} map. None -> {}. Unknown keys, non-numbers and values that are not finite and > 0
49
+ raise ValueError."""
50
+ if m is None:
51
+ return {}
52
+ if not isinstance(m, dict):
53
+ raise ValueError(f"{where} must be a map {{answer type: temperature}}, got {type(m).__name__}; " + _KEYS_TEXT)
54
+ out = {}
55
+ for k, v in m.items():
56
+ if k not in TYPES:
57
+ raise ValueError(f"{where} has the unknown key {k!r}; " + _KEYS_TEXT)
58
+ if isinstance(v, bool) or not isinstance(v, (int, float)):
59
+ raise ValueError(f"{where}[{k!r}] must be a finite number > 0, got {v!r}")
60
+ out[k] = positive(v, f"{where}[{k!r}]")
61
+ return out
62
+
63
+
64
+ def from_config(cfg, temperature=None, temperature_by_type=None):
65
+ """-> ((T, by_type) for the state-first layout, (T, by_type) for the schema cache).
66
+
67
+ temperature / temperature_by_type: explicit overrides (Decider arguments, DECIDER_TEMPERATURE). An explicit temperature
68
+ without an explicit map switches the config's map off (see the module docstring)."""
69
+ cfg = cfg or {}
70
+ where = "decider_config.json"
71
+ if temperature is not None:
72
+ T = positive(temperature, "temperature")
73
+ m = by_type(temperature_by_type, "temperature_by_type") if temperature_by_type is not None else {}
74
+ else:
75
+ T = positive(cfg.get("temperature", 1.0), f'{where} "temperature"')
76
+ m = (by_type(temperature_by_type, "temperature_by_type") if temperature_by_type is not None
77
+ else by_type(cfg.get("temperature_by_type"), f'{where} "temperature_by_type"'))
78
+ ms = by_type(cfg.get("temperature_schema_first_by_type"), f'{where} "temperature_schema_first_by_type"')
79
+ if "temperature_schema_first" in cfg:
80
+ schema = (positive(cfg["temperature_schema_first"], f'{where} "temperature_schema_first"'), ms)
81
+ else:
82
+ schema = (T, {**m, **ms})
83
+ return (T, m), schema
84
+
85
+
86
+ def effective(T, m):
87
+ """{answer type: the temperature it gets} for reporting (/health, the ready line)."""
88
+ return {t: m.get(t, T) for t in TYPES}
89
+
90
+
91
+ def for_types(T, m, types):
92
+ """One temperature per slot, or the scalar T itself when there is no map (the 1.3.0 call)."""
93
+ if not m:
94
+ return T
95
+ return [m.get(t, T) for t in types]
96
+
97
+
98
+ def item_types(it):
99
+ """The answer type of every slot of a prompt item ("types", set where the item is built); None for an item without it."""
100
+ ts = it.get("types")
101
+ return list(ts) if ts is not None else [None] * len(it["slots"])
102
+
103
+
104
+ def for_items(T, m, items):
105
+ """The `temperature` argument of Engine.score_items / score_shared: the scalar T when there is no map (the 1.3.0 call),
106
+ else one list of per-slot temperatures per item."""
107
+ if not m:
108
+ return T
109
+ return [[m.get(t, T) for t in item_types(it)] for it in items]
110
+
111
+
112
+ def slot_temperatures(temperature, items):
113
+ """Engine side. temperature: a scalar (returned as it is), or one entry per item, each a scalar or one value per slot of
114
+ that item. -> the scalar, or a flat list with one temperature per slot in item order."""
115
+ if not isinstance(temperature, (list, tuple)):
116
+ return temperature
117
+ if len(temperature) != len(items):
118
+ raise ValueError(f"temperature: {len(temperature)} entries for {len(items)} items")
119
+ flat = []
120
+ for t, it in zip(temperature, items):
121
+ n = len(it["slots"])
122
+ if isinstance(t, (list, tuple)):
123
+ if len(t) != n:
124
+ raise ValueError(f"temperature: {len(t)} values for an item with {n} slots")
125
+ flat += list(t)
126
+ else:
127
+ flat += [t] * n
128
+ return flat
129
+
130
+
131
+ def item_slice(temperature, lo, hi):
132
+ """The per-item temperature entries of items[lo:hi] (a scalar is shared by every item)."""
133
+ return temperature[lo:hi] if isinstance(temperature, (list, tuple)) else temperature
134
+
135
+
136
+ def scaled_softmax(lg, temperature):
137
+ """softmax(lg / T) over the last axis. A scalar T is the 1.3.0 expression unchanged; a list gives one T per row of lg."""
138
+ import torch
139
+ if isinstance(temperature, (list, tuple)):
140
+ if len(temperature) != lg.shape[0]:
141
+ raise ValueError(f"temperature: {len(temperature)} values for {lg.shape[0]} slots")
142
+ t = torch.tensor(temperature, dtype=lg.dtype).to(lg.device, non_blocking=True)[:, None]
143
+ return torch.softmax(lg / t, -1)
144
+ return torch.softmax(lg / temperature, -1)
decider_config.json CHANGED
@@ -1,7 +1,12 @@
1
  {
2
- "temperature": 1.935,
 
 
 
 
 
3
  "neutralize_none": false,
4
- "version": "4b-v2",
5
  "base": "Mapika/decider-4b v1 + LoRA (merged); v1 is Qwen/Qwen3.5-4B-Base + one supervised pass over mixture v2",
6
  "layout": "plain",
7
  "max_options": 255,
@@ -10,5 +15,6 @@
10
  "schema_first_trained": false,
11
  "isolated_levels": true,
12
  "release_date": "2026-09-24",
13
- "stage": "decider-4b v1 + LoRA rank 64 (alpha 128) on attention and MLP, LR 1e-4, 2 epochs (1,517 steps of 65,536 tokens) over 29,356 rows in the plain state-first layout (generated decision families with code-computed answers, Qwen3.6-27B-written document questions kept when two independent answers agreed, human-labelled public sets, replay of v1's mixture v2), merged into the bf16 weights; no RL stage; temperature fitted by NLL on 61 in-task regression tasks (the 67 in-task tasks without banking77, clinc_oos, mmlu, arc, winogrande, hellaswag)"
 
14
  }
 
1
  {
2
+ "temperature": 1.099,
3
+ "temperature_by_type": {
4
+ "choice": 1.11,
5
+ "noul": 1.56,
6
+ "score": 1.287
7
+ },
8
  "neutralize_none": false,
9
+ "version": "4b-v2.1",
10
  "base": "Mapika/decider-4b v1 + LoRA (merged); v1 is Qwen/Qwen3.5-4B-Base + one supervised pass over mixture v2",
11
  "layout": "plain",
12
  "max_options": 255,
 
15
  "schema_first_trained": false,
16
  "isolated_levels": true,
17
  "release_date": "2026-09-24",
18
+ "requires": "decider-ai>=1.4.0 for temperature_by_type; older versions serve every answer at temperature",
19
+ "stage": "decider-4b v1 + LoRA rank 64 (alpha 128) on attention and MLP, LR 1e-4, 2 epochs (1,518 steps of 65,536 tokens) over v2's 29,325-row mix in the plain state-first layout (generated decision families with code-computed answers, Qwen3.6-27B-written document questions kept when two independent answers agreed, human-labelled public sets, replay of v1's mixture v2), with the replay rows trained toward v1's own answer distribution (KL to v1) instead of their labels, merged into the bf16 weights; no RL stage; temperature fitted by NLL on 61 in-task regression tasks (the 67 in-task tasks without banking77, clinc_oos, mmlu, arc, winogrande, hellaswag); temperature_by_type fitted with decider.calibrate.fit_by_type on the same regression rows plus our own validation rows (choice, noul and score answers)"
20
  }
eval_results.json CHANGED
@@ -1,1796 +1,1549 @@
1
  {
2
- "model": "decider-4b-v2",
3
- "temperature": 1.935,
4
- "temperature_fit": {
5
- "method": "one scalar, pooled row NLL, golden-section search on log T in [0.3, 5]; stored rounded to 3 decimals",
6
- "tasks": 61,
7
- "rows": 102804,
8
- "excluded_in_task_tasks": [
9
- "banking77",
10
- "clinc_oos",
11
- "mmlu",
12
- "arc",
13
- "winogrande",
14
- "hellaswag"
15
- ],
16
- "fitted": 1.9349278889417365,
17
- "grid_fit_temperature": 1.9420684917099538,
18
- "stored": 1.935,
19
- "note": "fitted on the in-task half of the public regression set without banking77, clinc_oos, mmlu, arc, winogrande and hellaswag (61 tasks); no JevBench or Decision Index item was used for training, selection or temperature"
20
  },
21
- "protocol": "decider-4b v1 + merged LoRA (rank 64, attention and MLP, 2 epochs, 29,356 rows). Public regression set rebuilt on this machine: eval half of the mixture, 95 tasks (67 in-task / 28 held-out), state-first plain layout, max_ctx 1536, large label sets sub-sampled to 10 options; metrics at the stored temperature, ECE with 15 bins. Other sets through decider-ai 1.2.1 at T 1.935: fixtures through h2h/score_public.py (eager, 32k context), Bespoke suite through decider.bench.public_suite, games and browser through the h2h/serve_public.py shim, behaviour probes; JevBench public items read once at T 1.719 and rescaled to T 1.935 from the stored probabilities (18 isolated-level Score items kept at T 1.719); Decision Index sample read at T 1.719 and rescaled likewise; ten text games greedy (argmax, independent of T). Readout only.",
22
- "aggregate": {
23
- "in_task": {
24
- "acc": 0.8241469454813016,
25
- "nll": 0.441108525216357,
26
- "ece": 0.04134314671333522,
27
- "tasks": 67
28
- },
29
- "heldout": {
30
- "acc": 0.7787589834608305,
31
- "nll": 0.5663163406508309,
32
- "ece": 0.08033093216064069,
33
- "tasks": 28
34
- }
35
  },
36
- "results": {
37
- "abstain_probe": {
38
- "heldout": true,
39
- "acc": 0.5820987654320988,
40
- "nll": 1.1091256141662598,
41
- "ece": 0.13279259307884878
42
- },
43
- "ade": {
44
- "heldout": true,
45
- "acc": 0.7886666666666666,
46
- "nll": 0.5347910523414612,
47
- "ece": 0.10660141297181445
48
- },
49
- "ag_news": {
50
- "heldout": false,
51
- "acc": 0.9033333333333333,
52
- "nll": 0.294931024312973,
53
- "ece": 0.04695633757114412
54
- },
55
- "agenttraj": {
56
- "heldout": false,
57
- "acc": 0.8973333333333333,
58
- "nll": 0.2499454766511917,
59
- "ece": 0.03513214365641278
60
- },
61
- "amazon_stars": {
62
- "heldout": false,
63
- "acc": 0.602,
64
- "nll": 0.9151002764701843,
65
- "ece": 0.035494098405043265
66
- },
67
- "arc": {
68
- "heldout": false,
69
- "acc": 0.939419795221843,
70
- "nll": 0.1868872344493866,
71
- "ece": 0.0222670579577062
72
- },
73
- "arena_pref": {
74
- "heldout": true,
75
- "acc": 0.5033333333333333,
76
- "nll": 1.2348569631576538,
77
- "ece": 0.21849741025765737
78
- },
79
- "banking77": {
80
- "heldout": false,
81
- "acc": 0.972,
82
- "nll": 0.10399508476257324,
83
- "ece": 0.0268547142942746
84
- },
85
- "bbc_news": {
86
- "heldout": true,
87
- "acc": 0.951,
88
- "nll": 0.18327508866786957,
89
- "ece": 0.06801495689153671
90
- },
91
- "bias_in_bios": {
92
- "heldout": false,
93
- "acc": 0.9486666666666667,
94
- "nll": 0.15249282121658325,
95
- "ece": 0.013438965007662774
96
- },
97
- "bitext_support": {
98
- "heldout": false,
99
- "acc": 0.998,
100
- "nll": 0.016967715695500374,
101
- "ece": 0.01042288645108536
102
- },
103
- "boolq": {
104
- "heldout": false,
105
- "acc": 0.9033333333333333,
106
- "nll": 0.24676379561424255,
107
- "ece": 0.029159295320510863
108
- },
109
- "cb": {
110
- "heldout": true,
111
- "acc": 0.8392857142857143,
112
- "nll": 0.4322512745857239,
113
- "ece": 0.11792815795966557
114
- },
115
- "civil_comments": {
116
- "heldout": false,
117
- "acc": 0.9328,
118
- "nll": 0.1662478744983673,
119
- "ece": 0.010289057048161809
120
- },
121
- "clinc_oos": {
122
- "heldout": false,
123
- "acc": 0.9826666666666667,
124
- "nll": 0.08267337828874588,
125
- "ece": 0.024905774275461832
126
- },
127
- "cola": {
128
- "heldout": false,
129
- "acc": 0.8331735378715245,
130
- "nll": 0.3790046274662018,
131
- "ece": 0.02507039284546103
132
- },
133
- "commonsense_qa": {
134
- "heldout": false,
135
- "acc": 0.823914823914824,
136
- "nll": 0.4847744107246399,
137
- "ece": 0.030845976944930447
138
- },
139
- "copa": {
140
- "heldout": false,
141
- "acc": 1.0,
142
- "nll": 0.030766163021326065,
143
- "ece": 0.028578150868415832
144
- },
145
- "counterfactual": {
146
- "heldout": false,
147
- "acc": 0.9573333333333334,
148
- "nll": 0.13998979330062866,
149
- "ece": 0.012234268069267251
150
- },
151
- "cr_reviews": {
152
- "heldout": true,
153
- "acc": 0.897742363877822,
154
- "nll": 0.24005697667598724,
155
- "ece": 0.045616757347289316
156
- },
157
- "dbpedia": {
158
- "heldout": false,
159
- "acc": 0.9933333333333333,
160
- "nll": 0.027618318796157837,
161
- "ece": 0.013420193096001943
162
- },
163
- "dbpedia_l2": {
164
- "heldout": true,
165
- "acc": 0.9506666666666667,
166
- "nll": 0.1401711255311966,
167
- "ece": 0.017278132518132504
168
- },
169
- "dbpedia_l3": {
170
- "heldout": true,
171
- "acc": 0.9913333333333333,
172
- "nll": 0.03723839297890663,
173
- "ece": 0.011898971815903936
174
- },
175
- "dolly_category": {
176
- "heldout": true,
177
- "acc": 0.374,
178
- "nll": 1.7009543180465698,
179
- "ece": 0.1419213018019994
180
- },
181
- "emotion": {
182
- "heldout": false,
183
- "acc": 0.8333333333333334,
184
- "nll": 0.4790882170200348,
185
- "ece": 0.025491972724596655
186
- },
187
- "enron_spam": {
188
- "heldout": false,
189
- "acc": 0.988,
190
- "nll": 0.03346330299973488,
191
- "ece": 0.006207752108573955
192
- },
193
- "fever": {
194
- "heldout": false,
195
- "acc": 0.8806666666666667,
196
- "nll": 0.32262682914733887,
197
- "ece": 0.01808753836154938
198
- },
199
- "fin_phrasebank": {
200
- "heldout": true,
201
- "acc": 0.6989690721649484,
202
- "nll": 0.6234146952629089,
203
- "ece": 0.07664692620026696
204
- },
205
- "fin_sentiment": {
206
- "heldout": true,
207
- "acc": 0.826,
208
- "nll": 0.42305630445480347,
209
- "ece": 0.051281696399052924
210
- },
211
- "glaive_tools": {
212
- "heldout": false,
213
- "acc": 0.9369436201780416,
214
- "nll": 0.2514188289642334,
215
- "ece": 0.10150587457047022
216
- },
217
- "go_emotions": {
218
- "heldout": false,
219
- "acc": 0.776,
220
- "nll": 0.6853362917900085,
221
- "ece": 0.03606782659888266
222
- },
223
- "hate_offensive": {
224
- "heldout": false,
225
- "acc": 0.9133333333333333,
226
- "nll": 0.26428458094596863,
227
- "ece": 0.014751521805922176
228
- },
229
- "hate_speech_scales": {
230
- "heldout": false,
231
- "acc": 0.5582666666666667,
232
- "nll": 0.9945353269577026,
233
- "ece": 0.03258252822558084
234
- },
235
- "hellaswag": {
236
- "heldout": false,
237
- "acc": 0.9346666666666666,
238
- "nll": 0.1996963769197464,
239
- "ece": 0.015657478451728855
240
- },
241
- "helpsteer2": {
242
- "heldout": false,
243
- "acc": 0.5980732177263969,
244
- "nll": 0.9590670466423035,
245
- "ece": 0.02256552288183126
246
- },
247
- "helpsteer3_pref": {
248
- "heldout": false,
249
- "acc": 0.38605442176870747,
250
- "nll": 1.5417133569717407,
251
- "ece": 0.061995073291314706
252
- },
253
- "hermes_tools": {
254
- "heldout": true,
255
- "acc": 0.7373333333333333,
256
- "nll": 0.5897114872932434,
257
- "ece": 0.1240579075018565
258
- },
259
- "hh_rlhf": {
260
- "heldout": false,
261
- "acc": 0.6451612903225806,
262
- "nll": 0.680734395980835,
263
- "ece": 0.11346792193349971
264
- },
265
- "hwu64": {
266
- "heldout": true,
267
- "acc": 0.9628252788104089,
268
- "nll": 0.12855161726474762,
269
- "ece": 0.03735405791093864
270
- },
271
- "imdb": {
272
- "heldout": false,
273
- "acc": 0.968,
274
- "nll": 0.09611714631319046,
275
- "ece": 0.01006534445285793
276
- },
277
- "insincere_questions": {
278
- "heldout": false,
279
- "acc": 0.9553333333333334,
280
- "nll": 0.12364762276411057,
281
- "ece": 0.01763951281706497
282
- },
283
- "liar2": {
284
- "heldout": false,
285
- "acc": 0.35,
286
- "nll": 1.513845682144165,
287
- "ece": 0.06293645719687144
288
- },
289
- "massive_intent": {
290
- "heldout": false,
291
- "acc": 0.9606666666666667,
292
- "nll": 0.13908076286315918,
293
- "ece": 0.02024710043271384
294
- },
295
- "massive_scenario": {
296
- "heldout": true,
297
- "acc": 0.7933333333333333,
298
- "nll": 0.5642920732498169,
299
- "ece": 0.05245566362142563
300
- },
301
- "medmcqa": {
302
- "heldout": false,
303
- "acc": 0.6233333333333333,
304
- "nll": 0.9161791801452637,
305
- "ece": 0.07326398479938506
306
- },
307
- "medqa": {
308
- "heldout": false,
309
- "acc": 0.6779261586802828,
310
- "nll": 0.8010168075561523,
311
- "ece": 0.0796375350744258
312
- },
313
- "mind2web": {
314
- "heldout": false,
315
- "acc": 0.822,
316
- "nll": 0.48613396286964417,
317
- "ece": 0.04646803086996078
318
- },
319
- "mmlu": {
320
- "heldout": false,
321
- "acc": 0.742,
322
- "nll": 0.6732416749000549,
323
- "ece": 0.04000420778989793
324
- },
325
- "mnli": {
326
- "heldout": false,
327
- "acc": 0.9026666666666666,
328
- "nll": 0.28677940368652344,
329
- "ece": 0.019106337686379753
330
- },
331
- "mrpc": {
332
- "heldout": false,
333
- "acc": 0.8431372549019608,
334
- "nll": 0.35424306988716125,
335
- "ece": 0.031691053334404445
336
- },
337
- "msmarco_rel": {
338
- "heldout": false,
339
- "acc": 0.6874135546334716,
340
- "nll": 0.586560845375061,
341
- "ece": 0.035539872212693564
342
- },
343
- "multirc": {
344
- "heldout": false,
345
- "acc": 0.9053333333333333,
346
- "nll": 0.2743209898471832,
347
- "ece": 0.025939866781234736
348
- },
349
- "newsgroups": {
350
- "heldout": false,
351
- "acc": 0.820109439124487,
352
- "nll": 0.5166534781455994,
353
- "ece": 0.0274887290539885
354
- },
355
- "offtopic_probe": {
356
- "heldout": true,
357
- "acc": 0.8234567901234567,
358
- "nll": 0.5037797093391418,
359
- "ece": 0.043301358119941055
360
- },
361
- "openbookqa": {
362
- "heldout": false,
363
- "acc": 0.904,
364
- "nll": 0.2716519832611084,
365
- "ece": 0.033541144907474535
366
- },
367
- "paws": {
368
- "heldout": true,
369
- "acc": 0.794,
370
- "nll": 0.5527914762496948,
371
- "ece": 0.12406627035140992
372
- },
373
- "piqa": {
374
- "heldout": false,
375
- "acc": 0.89,
376
- "nll": 0.2665766477584839,
377
- "ece": 0.026692868391672748
378
- },
379
- "prosocial_safety": {
380
- "heldout": false,
381
- "acc": 0.512,
382
- "nll": 1.1915110349655151,
383
- "ece": 0.030761619885762537
384
- },
385
- "pubmedqa": {
386
- "heldout": true,
387
- "acc": 0.758,
388
- "nll": 0.6072925329208374,
389
- "ece": 0.05524836438894271
390
- },
391
- "qasc": {
392
- "heldout": false,
393
- "acc": 0.8596112311015118,
394
- "nll": 0.47038528323173523,
395
- "ece": 0.06783940297094844
396
- },
397
- "qnli": {
398
- "heldout": false,
399
- "acc": 0.946,
400
- "nll": 0.14609812200069427,
401
- "ece": 0.01228535489241284
402
- },
403
- "qqp": {
404
- "heldout": false,
405
- "acc": 0.848,
406
- "nll": 0.3371011018753052,
407
- "ece": 0.03070641124248504
408
- },
409
- "quality": {
410
- "heldout": true,
411
- "acc": 0.5606666666666666,
412
- "nll": 1.1722065210342407,
413
- "ece": 0.17895789907375972
414
- },
415
- "quality_full": {
416
- "heldout": true,
417
- "acc": 0.505,
418
- "nll": 1.2171955108642578,
419
- "ece": 0.1881948218246301
420
- },
421
- "race": {
422
- "heldout": false,
423
- "acc": 0.8953333333333333,
424
- "nll": 0.33984363079071045,
425
- "ece": 0.02000871032476428
426
- },
427
- "reward_bench": {
428
- "heldout": true,
429
- "acc": 0.8526666666666667,
430
- "nll": 0.34169474244117737,
431
- "ece": 0.03198828514417014
432
- },
433
- "rte": {
434
- "heldout": false,
435
- "acc": 0.9205776173285198,
436
- "nll": 0.1940610110759735,
437
- "ece": 0.04438359793342838
438
- },
439
- "sciq": {
440
- "heldout": true,
441
- "acc": 0.991,
442
- "nll": 0.044908951967954636,
443
- "ece": 0.02056253519654279
444
- },
445
- "shp": {
446
- "heldout": false,
447
- "acc": 0.7353333333333333,
448
- "nll": 0.594969630241394,
449
- "ece": 0.1105338017543157
450
- },
451
- "sms_spam": {
452
- "heldout": false,
453
- "acc": 0.987,
454
- "nll": 0.055126290768384933,
455
- "ece": 0.006505105614662178
456
- },
457
- "snli": {
458
- "heldout": false,
459
- "acc": 0.9227642276422764,
460
- "nll": 0.24105508625507355,
461
- "ece": 0.03949038551105716
462
- },
463
- "social_iqa": {
464
- "heldout": true,
465
- "acc": 0.7593333333333333,
466
- "nll": 0.5866184830665588,
467
- "ece": 0.06180198764801027
468
- },
469
- "sst2": {
470
- "heldout": false,
471
- "acc": 0.948394495412844,
472
- "nll": 0.14050182700157166,
473
- "ece": 0.01752297238472409
474
- },
475
- "sst5": {
476
- "heldout": false,
477
- "acc": 0.5826666666666667,
478
- "nll": 0.9461587071418762,
479
- "ece": 0.0451480305393537
480
- },
481
- "strategyqa": {
482
- "heldout": true,
483
- "acc": 0.6317321688500728,
484
- "nll": 0.6970252990722656,
485
- "ece": 0.1567338213129335
486
- },
487
- "student_questions": {
488
- "heldout": true,
489
- "acc": 0.94,
490
- "nll": 0.20459797978401184,
491
- "ece": 0.05185317244132365
492
- },
493
- "subjectivity": {
494
- "heldout": false,
495
- "acc": 0.9566666666666667,
496
- "nll": 0.11102409660816193,
497
- "ece": 0.015209745287895248
498
- },
499
- "support_tickets": {
500
- "heldout": false,
501
- "acc": 0.5586666666666666,
502
- "nll": 1.0297108888626099,
503
- "ece": 0.024460139827595817
504
- },
505
- "synth": {
506
- "heldout": false,
507
- "acc": 0.8066666666666666,
508
- "nll": 0.4917465150356293,
509
- "ece": 0.04993985931078593
510
- },
511
- "toolace": {
512
- "heldout": false,
513
- "acc": 0.931,
514
- "nll": 0.2593076229095459,
515
- "ece": 0.06617412987351419
516
- },
517
- "toxic_chat": {
518
- "heldout": false,
519
- "acc": 0.977,
520
- "nll": 0.06245163083076477,
521
- "ece": 0.00981356755892436
522
- },
523
- "trec": {
524
- "heldout": true,
525
- "acc": 0.842,
526
- "nll": 0.5022116899490356,
527
- "ece": 0.02651941215991974
528
- },
529
- "truthfulqa": {
530
- "heldout": true,
531
- "acc": 0.653610771113831,
532
- "nll": 1.0122824907302856,
533
- "ece": 0.06321512999849775
534
- },
535
- "tweet_emotion": {
536
- "heldout": false,
537
- "acc": 0.8578465869106263,
538
- "nll": 0.41329097747802734,
539
- "ece": 0.039397706623063994
540
- },
541
- "tweet_hate": {
542
- "heldout": false,
543
- "acc": 0.5126666666666667,
544
- "nll": 1.136211633682251,
545
- "ece": 0.34860085622469583
546
- },
547
- "tweet_irony": {
548
- "heldout": true,
549
- "acc": 0.826530612244898,
550
- "nll": 0.39445212483406067,
551
- "ece": 0.027230415189144572
552
- },
553
- "tweet_offensive": {
554
- "heldout": false,
555
- "acc": 0.8523255813953489,
556
- "nll": 0.3407946228981018,
557
- "ece": 0.021613172181817006
558
- },
559
- "tweet_sentiment": {
560
- "heldout": false,
561
- "acc": 0.7206666666666667,
562
- "nll": 0.610216498374939,
563
- "ece": 0.028348857482274364
564
- },
565
- "ultrafeedback_pref": {
566
- "heldout": false,
567
- "acc": 0.7713333333333333,
568
- "nll": 0.4769253432750702,
569
- "ece": 0.06639519715309146
570
- },
571
- "wic": {
572
- "heldout": false,
573
- "acc": 0.7429467084639498,
574
- "nll": 0.5646101832389832,
575
- "ece": 0.09329287749846528
576
- },
577
- "wiki_qa": {
578
- "heldout": false,
579
- "acc": 0.9089874857792947,
580
- "nll": 0.2449624091386795,
581
- "ece": 0.020241137725365718
582
- },
583
- "winogrande": {
584
- "heldout": false,
585
- "acc": 0.8089976322020521,
586
- "nll": 0.48967406153678894,
587
- "ece": 0.08959193594736103
588
- },
589
- "xstory_cloze": {
590
- "heldout": true,
591
- "acc": 0.9706666666666667,
592
- "nll": 0.07805304229259491,
593
- "ece": 0.01724668137232463
594
- },
595
- "yahoo_topics": {
596
- "heldout": false,
597
- "acc": 0.7613333333333333,
598
- "nll": 0.7648204565048218,
599
- "ece": 0.0657656339406967
600
- },
601
- "yelp": {
602
- "heldout": false,
603
- "acc": 0.7033333333333334,
604
- "nll": 0.7055407166481018,
605
- "ece": 0.04224825153748195
606
  }
607
  },
608
- "other_sets": {
609
- "general_validation_847": {
610
- "acc": 0.8501,
611
- "nll": 0.4186,
612
- "ece": 0.0305
613
- },
614
- "typesafe_102": {
615
- "acc": 0.8627,
616
- "nll": 0.4101,
617
- "ece": 0.0723,
618
- "tv_to_reference": 0.1961
619
- },
620
- "openjev_5252": {
621
- "acc": 0.6672,
622
- "nll": 0.789,
623
- "ece": 0.1309,
624
- "macro_acc_3_sources": 0.8098
625
- },
626
- "mind2web_1770": {
627
- "acc": 0.874,
628
- "nll": 0.3935,
629
- "ece": 0.0512
630
- },
631
- "jevbench_public_items": {
632
- "easy": 1.0,
633
- "standard": 0.9861,
634
- "hard": 0.6757,
635
- "easy_top_label_ece": 0.0078,
636
- "standard_top_label_ece": 0.0854,
637
- "hard_top_label_ece": 0.0708,
638
- "note": "read once at T 1.719 after selection; ECE rescaled to T 1.935 from stored probabilities",
639
- "score_items_kept_at_T_1719": 18,
640
- "hard_mean_confidence": 0.7401,
641
- "hard_conf_ge_0_9_coverage": 0.3063,
642
- "hard_conf_ge_0_9_accuracy": 0.9412
643
- },
644
- "bespoke_public_suite": {
645
- "macro": 0.7725,
646
- "micro": 0.7809,
647
- "macro_untrained": 0.7514,
648
- "per_subset": {
649
- "vitaminc-dev": 0.778,
650
- "massive-en-US": 0.8686,
651
- "massive-de-DE": 0.8371,
652
- "boolq": 0.86,
653
- "squad2": 0.7926,
654
- "paws": 0.832,
655
- "multinli": 0.9264,
656
- "civil_comments": 0.8567,
657
- "aegis2": 0.812,
658
- "helpsteer2": 0.4659,
659
- "summeval-relevance": 0.4625,
660
- "summeval-consistency": 0.8264,
661
- "pubmedqa": 0.724
662
- }
663
- },
664
- "zero_shot_games_win_rate": {
665
- "sampled_all": 0.2244,
666
- "greedy_all": 0.2778,
667
- "sampled_bags": 0.3789,
668
- "greedy_bags": 0.4844
669
- },
670
- "text_games_zero_shot_5_episodes": {
671
- "pong": {
672
- "model": -21.0,
673
- "teacher": 8.0,
674
- "random": -20.4
675
- },
676
- "freeway": {
677
- "model": 1.0,
678
- "teacher": 5.0,
679
- "random": 0.0
680
- },
681
- "breakout": {
682
- "model": 12.0,
683
- "teacher": 22.0,
684
- "random": 0.8
685
  },
686
- "frozenlake": {
687
- "model": 0.0,
688
- "teacher": 1.0,
689
- "random": 0.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
690
  },
691
- "cliffwalking": {
692
- "model": -60.0,
693
- "teacher": -13.0,
694
- "random": -574.8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
695
  },
696
- "blackjack": {
697
- "model": -0.6,
698
- "teacher": -0.6,
699
- "random": -1.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
700
  },
701
- "minigrid_empty": {
702
- "model": 0.0,
703
- "teacher": 0.961328125,
704
- "random": 0.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
705
  },
706
- "minigrid_lavagap": {
707
- "model": 0.0,
708
- "teacher": 0.19173469387755102,
709
- "random": 0.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
710
  },
711
- "minigrid_doorkey": {
712
- "model": 0.0,
713
- "teacher": 0.0,
714
- "random": 0.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
715
  },
716
- "babyai_goto": {
717
- "model": 0.3521875,
718
- "teacher": 0.3409375,
719
- "random": 0.1184375
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
720
  }
721
  },
722
- "miniwob_live_22x8": {
723
- "sampled": 0.8807,
724
- "sampled_held_out": 0.75,
725
- "greedy": 0.9261,
726
- "greedy_held_out": 0.8125
727
- },
728
- "decision_index_4k_sample": {
729
- "note": "read at T 1.719 through decider-ai 1.2.1; calibration rescaled to T 1.935 from stored probabilities (argmax unchanged); sample index measured at T 1.719",
730
- "sample_index_T1719": 48.3,
731
- "ece_bw": 0.0738,
732
- "acc_bw": 0.6373,
733
- "conf_bw": 0.6901,
734
- "over95_bw": 0.0123,
735
- "brier_bw": 0.4788,
736
- "pooled_ece": 0.0305,
737
- "n": 27662,
738
- "benchmarks": 33
739
- },
740
- "behaviour_probes": {
741
- "note": "decider.probes.batteries / applications / isolated through decider-ai 1.2.1; the isolated-level probe reads raw logits (T=1)",
742
- "generic": 0.95,
743
- "specific": 1.0,
744
- "catchall": 0.95,
745
- "abstain_battery": "8/8",
746
- "router_tier": 0.935,
747
- "needs_live_data": 0.839,
748
- "command_risk": 0.911,
749
- "destructive_called_safe": 0,
750
- "touches_outside_project": 0.933,
751
- "browser_element": 0.938,
752
- "browser_action": 0.938,
753
- "isolated_minus_listwise_max_abs_points": 2.0,
754
- "per_level_fit_sum_range": [
755
- 0.92,
756
- 1.0
757
- ],
758
- "listwise_ece_five_rating_sets_T1": [
759
- 0.127,
760
- 0.265
761
- ]
762
- },
763
- "speed": {
764
- "note": "same weights, measured on 2026-09-24 in the candidate readout session (server config T 1.719; the temperature does not change speed); one unshared B300, bf16",
765
- "eager_game_state_median_ms": 34.6,
766
- "v1_same_session_ms": 32.4,
767
- "graphs_compile_single_request_ms": 5.2,
768
- "fp8_single_request_ms": 5.0,
769
- "batch32_decisions_per_s": 1178,
770
- "batch32_fp8_decisions_per_s": 1349,
771
- "http_decide_1_client_req_per_s": 72.8,
772
- "http_decide_1_client_p50_ms": 13.3,
773
- "http_decide_64_clients_req_per_s": 189.7
774
- }
775
- },
776
- "paired": {
777
- "regression_set_per_task_bootstrap": {
778
- "v2-v1 T_1.05|in_task": {
779
- "acc": {
780
- "tasks": 67,
781
- "delta": -0.00963326990956669,
782
- "ci95": [
783
- -0.012258285462742567,
784
- -0.007113672993330958
785
- ],
786
- "higher": 6,
787
- "lower": 54,
788
- "equal": 7
789
- },
790
- "nll": {
791
- "tasks": 67,
792
- "delta": 0.0367151916395428,
793
- "ci95": [
794
- 0.02853969128380196,
795
- 0.04618769902040932
796
- ],
797
- "higher": 64,
798
- "lower": 3,
799
- "equal": 0
800
- },
801
- "ece": {
802
- "tasks": 67,
803
- "delta": 0.014469173095775816,
804
- "ci95": [
805
- 0.008371143739489972,
806
- 0.021221732726272202
807
- ],
808
- "higher": 49,
809
- "lower": 18,
810
- "equal": 0
811
  }
812
  },
813
- "v2-v1 T_1.05|heldout": {
814
- "acc": {
815
- "tasks": 28,
816
- "delta": -0.008823480116377602,
817
- "ci95": [
818
- -0.019676785964620195,
819
- 0.001414172129592744
820
- ],
821
- "higher": 7,
822
- "lower": 21,
823
- "equal": 0
824
- },
825
- "nll": {
826
- "tasks": 28,
827
- "delta": 0.00806199632851141,
828
- "ci95": [
829
- -0.024438296351581807,
830
- 0.03533247855292367
831
- ],
832
- "higher": 20,
833
- "lower": 8,
834
- "equal": 0
835
- },
836
- "ece": {
837
- "tasks": 28,
838
- "delta": 0.009488148203351523,
839
- "ci95": [
840
- -0.002820037805815392,
841
- 0.02130860449786515
842
- ],
843
- "higher": 20,
844
- "lower": 8,
845
- "equal": 0
846
  }
847
  },
848
- "v2-v10|in_task": {
849
- "acc": {
850
- "tasks": 67,
851
- "delta": 0.01875302115583182,
852
- "ci95": [
853
- 0.01015619959044516,
854
- 0.027614901925500205
855
- ],
856
- "higher": 44,
857
- "lower": 21,
858
- "equal": 2
859
- },
860
- "nll": {
861
- "tasks": 67,
862
- "delta": -0.032433275216773375,
863
- "ci95": [
864
- -0.05438354930479024,
865
- -0.011119020796283635
866
- ],
867
- "higher": 26,
868
- "lower": 41,
869
- "equal": 0
870
- },
871
- "ece": {
872
- "tasks": 67,
873
- "delta": 0.004320193883140634,
874
- "ci95": [
875
- -0.0023582249916685393,
876
- 0.011625304997016925
877
- ],
878
- "higher": 34,
879
- "lower": 33,
880
- "equal": 0
881
  }
882
  },
883
- "v2-v10|heldout": {
884
- "acc": {
885
- "tasks": 28,
886
- "delta": 0.02370397041876381,
887
- "ci95": [
888
- 0.011139798470057867,
889
- 0.03700944122337325
890
- ],
891
- "higher": 22,
892
- "lower": 6,
893
- "equal": 0
894
- },
895
- "nll": {
896
- "tasks": 28,
897
- "delta": -0.05593727236347539,
898
- "ci95": [
899
- -0.093489749350452,
900
- -0.023956677370837752
901
- ],
902
- "higher": 6,
903
- "lower": 22,
904
- "equal": 0
905
- },
906
- "ece": {
907
- "tasks": 28,
908
- "delta": -0.003365505320264168,
909
- "ci95": [
910
- -0.01585441365540306,
911
- 0.008293014476259788
912
- ],
913
- "higher": 13,
914
- "lower": 15,
915
- "equal": 0
916
  }
917
  },
918
- "v2-35b|in_task": {
919
- "acc": {
920
- "tasks": 67,
921
- "delta": -0.031028858088253602,
922
- "ci95": [
923
- -0.039777794555657386,
924
- -0.02317882618464841
925
- ],
926
- "higher": 3,
927
- "lower": 61,
928
- "equal": 3
929
- },
930
- "nll": {
931
- "tasks": 67,
932
- "delta": 0.08398698643544938,
933
- "ci95": [
934
- 0.06702507901571886,
935
- 0.10335495943019389
936
- ],
937
- "higher": 67,
938
- "lower": 0,
939
- "equal": 0
940
- },
941
- "ece": {
942
- "tasks": 67,
943
- "delta": 0.015340690748236678,
944
- "ci95": [
945
- 0.009872991235955714,
946
- 0.021626024452136343
947
- ],
948
- "higher": 55,
949
- "lower": 12,
950
- "equal": 0
951
  }
952
  },
953
- "v2-35b|heldout": {
954
- "acc": {
955
- "tasks": 28,
956
- "delta": -0.03173775153445216,
957
- "ci95": [
958
- -0.045258765155078186,
959
- -0.018006466515002026
960
- ],
961
- "higher": 4,
962
- "lower": 24,
963
- "equal": 0
964
- },
965
- "nll": {
966
- "tasks": 28,
967
- "delta": 0.06927653994145137,
968
- "ci95": [
969
- 0.03497178333179493,
970
- 0.10177155091660098
971
- ],
972
- "higher": 24,
973
- "lower": 4,
974
- "equal": 0
975
- },
976
- "ece": {
977
- "tasks": 28,
978
- "delta": 0.01149923111649334,
979
- "ci95": [
980
- -0.010094514078614029,
981
- 0.03067631731368639
982
- ],
983
- "higher": 20,
984
- "lower": 8,
985
- "equal": 0
986
- }
987
- }
988
- },
989
- "fixtures_row_bootstrap_minus_v1": {
990
- "general_validation": {
991
- "accuracy": {
992
- "delta": -0.010625737898465172,
993
- "ci95": [
994
- -0.025974025974025976,
995
- 0.004722550177095631
996
- ]
997
- },
998
- "nll": {
999
- "delta": 0.0015418035959596215,
1000
- "ci95": [
1001
- -0.018830652980342588,
1002
- 0.02145239419471463
1003
- ]
1004
  }
1005
  },
1006
- "typesafe": {
1007
- "accuracy": {
1008
- "delta": 0.049019607843137254,
1009
- "ci95": [
1010
- -0.00980392156862745,
1011
- 0.10784313725490197
1012
- ]
1013
- },
1014
- "nll": {
1015
- "delta": -0.20073007691013034,
1016
- "ci95": [
1017
- -0.34779699906684913,
1018
- -0.07224570606300662
1019
- ]
 
 
 
 
1020
  }
1021
  },
1022
- "openjev_external": {
1023
- "accuracy": {
1024
- "delta": 0.028370144706778372,
1025
- "ci95": [
1026
- 0.017707539984767706,
1027
- 0.03865194211728865
1028
- ]
1029
- },
1030
- "nll": {
1031
- "delta": -0.10404728527555326,
1032
- "ci95": [
1033
- -0.12081249076065353,
1034
- -0.08734350246247681
1035
- ]
 
 
 
 
1036
  }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1037
  },
1038
- "mind2web": {
1039
- "accuracy": {
1040
- "delta": -0.0096045197740113,
1041
- "ci95": [
1042
- -0.01977401129943503,
1043
- 0.0005649717514124294
1044
- ]
1045
- },
1046
- "nll": {
1047
- "delta": 0.027492282187420017,
1048
- "ci95": [
1049
- 0.010634569199404512,
1050
- 0.04390014037543978
1051
- ]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1052
  }
1053
  }
1054
  },
1055
- "games_and_browser": {
1056
- "games": {
1057
- "v2-v1|sampled|all": {
1058
- "n": 234,
1059
- "a": 0.22435897435897437,
1060
- "b": 0.2777777777777778,
1061
- "delta": -0.053418803418803416,
1062
- "ci95": [
1063
- -0.07585470085470085,
1064
- -0.030982905982905984
1065
- ]
1066
- },
1067
- "v2-v1|sampled|bag_game": {
1068
- "n": 64,
1069
- "a": 0.37890625,
1070
- "b": 0.56640625,
1071
- "delta": -0.1875,
1072
- "ci95": [
1073
- -0.24609375,
1074
- -0.12890625
1075
- ]
1076
- },
1077
- "v2-v1|sampled|stochastic_grid": {
1078
- "n": 64,
1079
- "a": 0.16015625,
1080
- "b": 0.1640625,
1081
- "delta": -0.00390625,
1082
- "ci95": [
1083
- -0.02734375,
1084
- 0.01953125
1085
- ]
1086
- },
1087
- "v2-v1|sampled|tic_tac_toe": {
1088
- "n": 74,
1089
- "a": 0.24324324324324326,
1090
- "b": 0.24324324324324326,
1091
- "delta": 0.0,
1092
- "ci95": [
1093
- -0.02702702702702703,
1094
- 0.02702702702702703
1095
- ]
1096
- },
1097
- "v2-v1|sampled|exact_minesweeper": {
1098
- "n": 32,
1099
- "a": 0.0,
1100
- "b": 0.0078125,
1101
- "delta": -0.0078125,
1102
- "ci95": [
1103
- -0.0234375,
1104
- 0.0
1105
- ]
1106
- },
1107
- "v2-v1|greedy|all": {
1108
- "n": 234,
1109
- "a": 0.2777777777777778,
1110
- "b": 0.2905982905982906,
1111
- "delta": -0.01282051282051282,
1112
- "ci95": [
1113
- -0.05555555555555555,
1114
- 0.02564102564102564
1115
- ]
1116
- },
1117
- "v2-v1|greedy|bag_game": {
1118
- "n": 64,
1119
- "a": 0.484375,
1120
- "b": 0.625,
1121
- "delta": -0.140625,
1122
- "ci95": [
1123
- -0.234375,
1124
- -0.0625
1125
- ]
1126
- },
1127
- "v2-v1|greedy|stochastic_grid": {
1128
- "n": 64,
1129
- "a": 0.1875,
1130
- "b": 0.125,
1131
- "delta": 0.0625,
1132
- "ci95": [
1133
- 0.0,
1134
- 0.140625
1135
- ]
1136
- },
1137
- "v2-v1|greedy|tic_tac_toe": {
1138
- "n": 74,
1139
- "a": 0.2972972972972973,
1140
- "b": 0.2702702702702703,
1141
- "delta": 0.02702702702702703,
1142
- "ci95": [
1143
- -0.04054054054054054,
1144
- 0.0945945945945946
1145
- ]
1146
- },
1147
- "v2-v1|greedy|exact_minesweeper": {
1148
- "n": 32,
1149
- "a": 0.0,
1150
- "b": 0.0,
1151
- "delta": 0.0,
1152
- "ci95": [
1153
- 0.0,
1154
- 0.0
1155
- ]
1156
- },
1157
- "v2-v2_T1719|sampled|all": {
1158
- "n": 234,
1159
- "a": 0.22435897435897437,
1160
- "b": 0.23397435897435898,
1161
- "delta": -0.009615384615384616,
1162
- "ci95": [
1163
- -0.021367521367521368,
1164
- 0.0010683760683760685
1165
- ]
1166
- },
1167
- "v2-v2_T1719|sampled|bag_game": {
1168
- "n": 64,
1169
- "a": 0.37890625,
1170
- "b": 0.390625,
1171
- "delta": -0.01171875,
1172
- "ci95": [
1173
- -0.04296875,
1174
- 0.01953125
1175
- ]
1176
- },
1177
- "v2-v2_T1719|sampled|stochastic_grid": {
1178
- "n": 64,
1179
- "a": 0.16015625,
1180
- "b": 0.1796875,
1181
- "delta": -0.01953125,
1182
- "ci95": [
1183
- -0.04296875,
1184
- 0.0
1185
- ]
1186
- },
1187
- "v2-v2_T1719|sampled|tic_tac_toe": {
1188
- "n": 74,
1189
- "a": 0.24324324324324326,
1190
- "b": 0.24662162162162163,
1191
- "delta": -0.0033783783783783786,
1192
- "ci95": [
1193
- -0.016891891891891893,
1194
- 0.006756756756756757
1195
- ]
1196
- },
1197
- "v2-v2_T1719|sampled|exact_minesweeper": {
1198
- "n": 32,
1199
- "a": 0.0,
1200
- "b": 0.0,
1201
- "delta": 0.0,
1202
- "ci95": [
1203
- 0.0,
1204
- 0.0
1205
- ]
1206
- },
1207
- "v2-v2_T1719|greedy|all": {
1208
- "n": 234,
1209
- "a": 0.2777777777777778,
1210
- "b": 0.2777777777777778,
1211
- "delta": 0.0,
1212
- "ci95": [
1213
- 0.0,
1214
- 0.0
1215
- ]
1216
- },
1217
- "v2-v2_T1719|greedy|bag_game": {
1218
- "n": 64,
1219
- "a": 0.484375,
1220
- "b": 0.484375,
1221
- "delta": 0.0,
1222
- "ci95": [
1223
- 0.0,
1224
- 0.0
1225
- ]
1226
- },
1227
- "v2-v2_T1719|greedy|stochastic_grid": {
1228
- "n": 64,
1229
- "a": 0.1875,
1230
- "b": 0.1875,
1231
- "delta": 0.0,
1232
- "ci95": [
1233
- 0.0,
1234
- 0.0
1235
- ]
1236
- },
1237
- "v2-v2_T1719|greedy|tic_tac_toe": {
1238
- "n": 74,
1239
- "a": 0.2972972972972973,
1240
- "b": 0.2972972972972973,
1241
- "delta": 0.0,
1242
- "ci95": [
1243
- 0.0,
1244
- 0.0
1245
- ]
1246
- },
1247
- "v2-v2_T1719|greedy|exact_minesweeper": {
1248
- "n": 32,
1249
- "a": 0.0,
1250
- "b": 0.0,
1251
- "delta": 0.0,
1252
- "ci95": [
1253
- 0.0,
1254
- 0.0
1255
- ]
1256
- },
1257
- "v2-v10|sampled|all": {
1258
- "n": 234,
1259
- "a": 0.22435897435897437,
1260
- "b": 0.23717948717948717,
1261
- "delta": -0.01282051282051282,
1262
- "ci95": [
1263
- -0.041666666666666664,
1264
- 0.014957264957264958
1265
- ]
1266
- },
1267
- "v2-v10|sampled|bag_game": {
1268
- "n": 64,
1269
- "a": 0.37890625,
1270
- "b": 0.4140625,
1271
- "delta": -0.03515625,
1272
- "ci95": [
1273
- -0.08984375,
1274
- 0.01953125
1275
- ]
1276
- },
1277
- "v2-v10|sampled|stochastic_grid": {
1278
- "n": 64,
1279
- "a": 0.16015625,
1280
- "b": 0.1875,
1281
- "delta": -0.02734375,
1282
- "ci95": [
1283
- -0.10546875,
1284
- 0.04296875
1285
- ]
1286
- },
1287
- "v2-v10|sampled|tic_tac_toe": {
1288
- "n": 74,
1289
- "a": 0.24324324324324326,
1290
- "b": 0.22972972972972974,
1291
- "delta": 0.013513513513513514,
1292
- "ci95": [
1293
- -0.02364864864864865,
1294
- 0.05405405405405406
1295
- ]
1296
- },
1297
- "v2-v10|sampled|exact_minesweeper": {
1298
- "n": 32,
1299
- "a": 0.0,
1300
- "b": 0.0,
1301
- "delta": 0.0,
1302
- "ci95": [
1303
- 0.0,
1304
- 0.0
1305
- ]
1306
- },
1307
- "v2-v10|greedy|all": {
1308
- "n": 234,
1309
- "a": 0.2777777777777778,
1310
- "b": 0.26495726495726496,
1311
- "delta": 0.01282051282051282,
1312
- "ci95": [
1313
- -0.042735042735042736,
1314
- 0.06837606837606838
1315
- ]
1316
- },
1317
- "v2-v10|greedy|bag_game": {
1318
- "n": 64,
1319
- "a": 0.484375,
1320
- "b": 0.578125,
1321
- "delta": -0.09375,
1322
- "ci95": [
1323
- -0.203125,
1324
- 0.015625
1325
- ]
1326
- },
1327
- "v2-v10|greedy|stochastic_grid": {
1328
- "n": 64,
1329
- "a": 0.1875,
1330
- "b": 0.15625,
1331
- "delta": 0.03125,
1332
- "ci95": [
1333
- -0.078125,
1334
- 0.140625
1335
- ]
1336
- },
1337
- "v2-v10|greedy|tic_tac_toe": {
1338
- "n": 74,
1339
- "a": 0.2972972972972973,
1340
- "b": 0.20270270270270271,
1341
- "delta": 0.0945945945945946,
1342
- "ci95": [
1343
- -0.013513513513513514,
1344
- 0.20270270270270271
1345
- ]
1346
- },
1347
- "v2-v10|greedy|exact_minesweeper": {
1348
- "n": 32,
1349
- "a": 0.0,
1350
- "b": 0.0,
1351
- "delta": 0.0,
1352
- "ci95": [
1353
- 0.0,
1354
- 0.0
1355
- ]
1356
- },
1357
- "v2-35b|sampled|all": {
1358
- "n": 234,
1359
- "a": 0.22435897435897437,
1360
- "b": 0.24145299145299146,
1361
- "delta": -0.017094017094017096,
1362
- "ci95": [
1363
- -0.03739316239316239,
1364
- 0.002136752136752137
1365
- ]
1366
- },
1367
- "v2-35b|sampled|bag_game": {
1368
- "n": 64,
1369
- "a": 0.37890625,
1370
- "b": 0.41796875,
1371
- "delta": -0.0390625,
1372
- "ci95": [
1373
- -0.0859375,
1374
- 0.0078125
1375
- ]
1376
- },
1377
- "v2-35b|sampled|stochastic_grid": {
1378
- "n": 64,
1379
- "a": 0.16015625,
1380
- "b": 0.16796875,
1381
- "delta": -0.0078125,
1382
- "ci95": [
1383
- -0.04296875,
1384
- 0.02734375
1385
- ]
1386
- },
1387
- "v2-35b|sampled|tic_tac_toe": {
1388
- "n": 74,
1389
- "a": 0.24324324324324326,
1390
- "b": 0.25675675675675674,
1391
- "delta": -0.013513513513513514,
1392
- "ci95": [
1393
- -0.05405405405405406,
1394
- 0.02027027027027027
1395
- ]
1396
- },
1397
- "v2-35b|sampled|exact_minesweeper": {
1398
- "n": 32,
1399
- "a": 0.0,
1400
- "b": 0.0,
1401
- "delta": 0.0,
1402
- "ci95": [
1403
- 0.0,
1404
- 0.0
1405
- ]
1406
- },
1407
- "v2-35b|greedy|all": {
1408
- "n": 234,
1409
- "a": 0.2777777777777778,
1410
- "b": 0.3717948717948718,
1411
- "delta": -0.09401709401709402,
1412
- "ci95": [
1413
- -0.14957264957264957,
1414
- -0.038461538461538464
1415
- ]
1416
- },
1417
- "v2-35b|greedy|bag_game": {
1418
- "n": 64,
1419
- "a": 0.484375,
1420
- "b": 0.625,
1421
- "delta": -0.140625,
1422
- "ci95": [
1423
- -0.234375,
1424
- -0.0625
1425
- ]
1426
- },
1427
- "v2-35b|greedy|stochastic_grid": {
1428
- "n": 64,
1429
- "a": 0.1875,
1430
- "b": 0.328125,
1431
- "delta": -0.140625,
1432
- "ci95": [
1433
- -0.25,
1434
- -0.03125
1435
- ]
1436
- },
1437
- "v2-35b|greedy|tic_tac_toe": {
1438
- "n": 74,
1439
- "a": 0.2972972972972973,
1440
- "b": 0.33783783783783783,
1441
- "delta": -0.04054054054054054,
1442
- "ci95": [
1443
- -0.16216216216216217,
1444
- 0.08108108108108109
1445
- ]
1446
- },
1447
- "v2-35b|greedy|exact_minesweeper": {
1448
- "n": 32,
1449
- "a": 0.0,
1450
- "b": 0.03125,
1451
- "delta": -0.03125,
1452
- "ci95": [
1453
- -0.09375,
1454
- 0.0
1455
- ]
1456
- }
1457
  },
1458
- "browser": {
1459
- "v2-v1|greedy|all": {
1460
- "n": 176,
1461
- "a": 0.9261363636363636,
1462
- "b": 0.9147727272727273,
1463
- "delta": 0.011363636363636364,
1464
- "ci95": [
1465
- -0.017045454545454544,
1466
- 0.03977272727272727
1467
- ]
1468
- },
1469
- "v2-v1|greedy|rewarded": {
1470
- "n": 128,
1471
- "a": 0.96875,
1472
- "b": 0.9765625,
1473
- "delta": -0.0078125,
1474
- "ci95": [
1475
- -0.0234375,
1476
- 0.0
1477
- ]
1478
- },
1479
- "v2-v1|greedy|held_out": {
1480
- "n": 48,
1481
- "a": 0.8125,
1482
- "b": 0.75,
1483
- "delta": 0.0625,
1484
- "ci95": [
1485
- -0.020833333333333332,
1486
- 0.16666666666666666
1487
- ]
1488
- },
1489
- "v2-v1|sampled|all": {
1490
- "n": 176,
1491
- "a": 0.8806818181818182,
1492
- "b": 0.9090909090909091,
1493
- "delta": -0.028409090909090908,
1494
- "ci95": [
1495
- -0.07386363636363637,
1496
- 0.017045454545454544
1497
- ]
1498
- },
1499
- "v2-v1|sampled|rewarded": {
1500
- "n": 128,
1501
- "a": 0.9296875,
1502
- "b": 0.9609375,
1503
- "delta": -0.03125,
1504
- "ci95": [
1505
- -0.078125,
1506
- 0.0078125
1507
- ]
1508
- },
1509
- "v2-v1|sampled|held_out": {
1510
- "n": 48,
1511
- "a": 0.75,
1512
- "b": 0.7708333333333334,
1513
- "delta": -0.020833333333333332,
1514
- "ci95": [
1515
- -0.14583333333333334,
1516
- 0.10416666666666667
1517
- ]
1518
- },
1519
- "v2-v2_T1719|greedy|all": {
1520
- "n": 176,
1521
- "a": 0.9261363636363636,
1522
- "b": 0.9261363636363636,
1523
- "delta": 0.0,
1524
- "ci95": [
1525
- 0.0,
1526
- 0.0
1527
- ]
1528
- },
1529
- "v2-v2_T1719|greedy|rewarded": {
1530
- "n": 128,
1531
- "a": 0.96875,
1532
- "b": 0.96875,
1533
- "delta": 0.0,
1534
- "ci95": [
1535
- 0.0,
1536
- 0.0
1537
- ]
1538
- },
1539
- "v2-v2_T1719|greedy|held_out": {
1540
- "n": 48,
1541
- "a": 0.8125,
1542
- "b": 0.8125,
1543
- "delta": 0.0,
1544
- "ci95": [
1545
- 0.0,
1546
- 0.0
1547
- ]
1548
- },
1549
- "v2-v2_T1719|sampled|all": {
1550
- "n": 176,
1551
- "a": 0.8806818181818182,
1552
- "b": 0.8636363636363636,
1553
- "delta": 0.017045454545454544,
1554
- "ci95": [
1555
- -0.022727272727272728,
1556
- 0.0625
1557
- ]
1558
- },
1559
- "v2-v2_T1719|sampled|rewarded": {
1560
- "n": 128,
1561
- "a": 0.9296875,
1562
- "b": 0.8984375,
1563
- "delta": 0.03125,
1564
- "ci95": [
1565
- -0.015625,
1566
- 0.078125
1567
- ]
1568
- },
1569
- "v2-v2_T1719|sampled|held_out": {
1570
- "n": 48,
1571
- "a": 0.75,
1572
- "b": 0.7708333333333334,
1573
- "delta": -0.020833333333333332,
1574
- "ci95": [
1575
- -0.10416666666666667,
1576
- 0.0625
1577
- ]
1578
- },
1579
- "v2-v10|greedy|all": {
1580
- "n": 176,
1581
- "a": 0.9261363636363636,
1582
- "b": 0.9090909090909091,
1583
- "delta": 0.017045454545454544,
1584
- "ci95": [
1585
- -0.03977272727272727,
1586
- 0.07386363636363637
1587
- ]
1588
- },
1589
- "v2-v10|greedy|rewarded": {
1590
- "n": 128,
1591
- "a": 0.96875,
1592
- "b": 0.90625,
1593
- "delta": 0.0625,
1594
- "ci95": [
1595
- 0.0078125,
1596
- 0.125
1597
- ]
1598
- },
1599
- "v2-v10|greedy|held_out": {
1600
- "n": 48,
1601
- "a": 0.8125,
1602
- "b": 0.9166666666666666,
1603
- "delta": -0.10416666666666667,
1604
- "ci95": [
1605
- -0.22916666666666666,
1606
- 0.020833333333333332
1607
- ]
1608
- },
1609
- "v2-v10|sampled|all": {
1610
- "n": 176,
1611
- "a": 0.8806818181818182,
1612
- "b": 0.9318181818181818,
1613
- "delta": -0.05113636363636364,
1614
- "ci95": [
1615
- -0.10795454545454546,
1616
- 0.005681818181818182
1617
- ]
1618
- },
1619
- "v2-v10|sampled|rewarded": {
1620
- "n": 128,
1621
- "a": 0.9296875,
1622
- "b": 0.9375,
1623
- "delta": -0.0078125,
1624
- "ci95": [
1625
- -0.0703125,
1626
- 0.0546875
1627
- ]
1628
- },
1629
- "v2-v10|sampled|held_out": {
1630
- "n": 48,
1631
- "a": 0.75,
1632
- "b": 0.9166666666666666,
1633
- "delta": -0.16666666666666666,
1634
- "ci95": [
1635
- -0.3125,
1636
- -0.041666666666666664
1637
- ]
1638
- },
1639
- "v2-35b|greedy|all": {
1640
- "n": 176,
1641
- "a": 0.9261363636363636,
1642
- "b": 0.9715909090909091,
1643
- "delta": -0.045454545454545456,
1644
- "ci95": [
1645
- -0.07954545454545454,
1646
- -0.017045454545454544
1647
- ]
1648
- },
1649
- "v2-35b|greedy|rewarded": {
1650
- "n": 128,
1651
- "a": 0.96875,
1652
- "b": 0.96875,
1653
- "delta": 0.0,
1654
- "ci95": [
1655
- 0.0,
1656
- 0.0
1657
- ]
1658
- },
1659
- "v2-35b|greedy|held_out": {
1660
- "n": 48,
1661
- "a": 0.8125,
1662
- "b": 0.9791666666666666,
1663
- "delta": -0.16666666666666666,
1664
- "ci95": [
1665
- -0.2708333333333333,
1666
- -0.0625
1667
- ]
1668
- },
1669
- "v2-35b|sampled|all": {
1670
- "n": 176,
1671
- "a": 0.8806818181818182,
1672
- "b": 0.8636363636363636,
1673
- "delta": 0.017045454545454544,
1674
- "ci95": [
1675
- -0.03409090909090909,
1676
- 0.06832386363636415
1677
- ]
1678
- },
1679
- "v2-35b|sampled|rewarded": {
1680
- "n": 128,
1681
- "a": 0.9296875,
1682
- "b": 0.890625,
1683
- "delta": 0.0390625,
1684
- "ci95": [
1685
- -0.015625,
1686
- 0.09375
1687
- ]
1688
- },
1689
- "v2-35b|sampled|held_out": {
1690
- "n": 48,
1691
- "a": 0.75,
1692
- "b": 0.7916666666666666,
1693
- "delta": -0.041666666666666664,
1694
- "ci95": [
1695
- -0.16666666666666666,
1696
- 0.08333333333333333
1697
- ]
1698
- },
1699
- "per_task|greedy": {
1700
- "click-button": 1.0,
1701
- "click-button-sequence": 1.0,
1702
- "click-checkboxes": 1.0,
1703
- "click-checkboxes-large": 0.625,
1704
- "click-checkboxes-soft": 1.0,
1705
- "click-checkboxes-transfer": 1.0,
1706
- "click-collapsible": 1.0,
1707
- "click-collapsible-2": 0.5,
1708
- "click-color": 1.0,
1709
- "click-dialog": 1.0,
1710
- "click-dialog-2": 1.0,
1711
- "click-link": 1.0,
1712
- "click-option": 1.0,
1713
- "click-tab": 1.0,
1714
- "click-tab-2": 0.875,
1715
- "click-tab-2-hard": 1.0,
1716
- "click-test": 1.0,
1717
- "click-test-2": 1.0,
1718
- "click-widget": 1.0,
1719
- "focus-text": 1.0,
1720
- "focus-text-2": 0.375,
1721
- "navigate-tree": 1.0
1722
- },
1723
- "per_task|sampled": {
1724
- "click-button": 1.0,
1725
- "click-button-sequence": 1.0,
1726
- "click-checkboxes": 1.0,
1727
- "click-checkboxes-large": 0.75,
1728
- "click-checkboxes-soft": 0.875,
1729
- "click-checkboxes-transfer": 0.875,
1730
- "click-collapsible": 1.0,
1731
- "click-collapsible-2": 0.75,
1732
- "click-color": 1.0,
1733
- "click-dialog": 1.0,
1734
- "click-dialog-2": 0.5,
1735
- "click-link": 1.0,
1736
- "click-option": 0.875,
1737
- "click-tab": 0.625,
1738
- "click-tab-2": 0.75,
1739
- "click-tab-2-hard": 0.625,
1740
- "click-test": 1.0,
1741
- "click-test-2": 1.0,
1742
- "click-widget": 1.0,
1743
- "focus-text": 1.0,
1744
- "focus-text-2": 0.75,
1745
- "navigate-tree": 1.0
1746
  }
1747
  }
1748
  },
1749
- "jevbench_item_bootstrap_minus_v1": {
1750
- "easy": {
1751
- "n": 48,
1752
- "a": 1.0,
1753
- "b": 1.0,
1754
- "delta": 0.0,
1755
- "ci95": [
1756
- 0.0,
1757
- 0.0
1758
- ],
1759
- "a_only": 0,
1760
- "b_only": 0,
1761
- "argmax_differs": 0
1762
  },
1763
- "standard": {
1764
- "n": 72,
1765
- "a": 0.9861111111111112,
1766
- "b": 0.9583333333333334,
1767
- "delta": 0.027777777777777776,
1768
- "ci95": [
1769
- 0.0,
1770
- 0.06944444444444445
1771
- ],
1772
- "a_only": 2,
1773
- "b_only": 0,
1774
- "argmax_differs": 2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1775
  },
1776
- "hard": {
1777
- "n": 111,
1778
- "a": 0.6756756756756757,
1779
- "b": 0.5495495495495496,
1780
- "delta": 0.12612612612612611,
1781
- "ci95": [
1782
- 0.04504504504504504,
1783
- 0.20743243243243326
1784
- ],
1785
- "a_only": 20,
1786
- "b_only": 6,
1787
- "argmax_differs": 32
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1788
  }
1789
  }
1790
  },
1791
  "references": {
1792
- "decider_4b_v1": "Mapika/decider-4b revision v1, T 1.05",
1793
- "decider_2b_v10": "Mapika/decider-2b, T 1.30",
1794
- "decider_35b_a3b": "Mapika/decider-35b-a3b, T 1.075"
1795
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1796
  }
 
1
  {
2
+ "model": "decider-4b-v2.1",
3
+ "temperature": 1.099,
4
+ "temperature_by_type": {
5
+ "choice": 1.11,
6
+ "noul": 1.56,
7
+ "score": 1.287
 
 
 
 
 
 
 
 
 
 
 
 
8
  },
9
+ "temperature_fit": {
10
+ "global": "one scalar, pooled row NLL on the in-task half of the public regression set without banking77, clinc_oos, mmlu, arc, winogrande and hellaswag (61 tasks, 102,804 rows)",
11
+ "by_type": "decider.calibrate.fit_by_type (decider-ai 1.4.0), min_rows 50, on the same regression rows plus our own validation rows (val_jb, val_teacher, val_teacher2, cal_human without mmlu, arc and banking77, val_replay without six excluded families)",
12
+ "rows": {
13
+ "choice": 108927,
14
+ "noul": 1677,
15
+ "score": 535
16
+ },
17
+ "note": "no JevBench or Decision Index item was used for training, selection or temperature"
 
 
 
 
 
18
  },
19
+ "protocol": "Regression set: eval half of the mixture, 95 tasks (67 in-task / 28 held-out), plain state-first layout, large label sets sub-sampled to 10 options, ECE 15 bins, per-task mean. Own sets (heldout_jb, test_teacher2, cal_human, guard): stored T=1 logits read at the served temperatures, ECE 10 bins top label. Card lanes (fixtures, games, browser, probes, Bespoke, issue #9) through decider-ai 1.4.0 with the map; parents through 1.3.0 at their released temperatures, same rows and seeds; intervals are 95% paired bootstrap. JevBench public items read in process through 1.4.0 with the map; parents' files read earlier through 1.2.1.",
20
+ "regression": {
21
+ "map": {
22
+ "reg_in": 0.8308,
23
+ "reg_out": 0.7838,
24
+ "reg_in_nll": 0.4145,
25
+ "reg_out_nll": 0.5689,
26
+ "reg_in_ece": 0.0309,
27
+ "reg_out_ece": 0.0775
28
+ },
29
+ "global_T": {
30
+ "reg_in": 0.8308,
31
+ "reg_out": 0.7838,
32
+ "reg_in_nll": 0.4145,
33
+ "reg_out_nll": 0.5703,
34
+ "reg_in_ece": 0.0308,
35
+ "reg_out_ece": 0.0781
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
  }
37
  },
38
+ "own_sets": {
39
+ "map": {
40
+ "heldout_jb": {
41
+ "all": {
42
+ "n": 5000,
43
+ "accuracy": 0.556,
44
+ "nll": 1.2571,
45
+ "ece": 0.1466,
46
+ "mean_conf": 0.7026
47
+ },
48
+ "by_type": {
49
+ "choice": {
50
+ "n": 3293,
51
+ "accuracy": 0.5056,
52
+ "nll": 1.4995,
53
+ "ece": 0.1701,
54
+ "mean_conf": 0.6757
55
+ },
56
+ "noul": {
57
+ "n": 1289,
58
+ "accuracy": 0.7114,
59
+ "nll": 0.6074,
60
+ "ece": 0.1013,
61
+ "mean_conf": 0.8127
62
+ },
63
+ "score": {
64
+ "n": 418,
65
+ "accuracy": 0.4737,
66
+ "nll": 1.3504,
67
+ "ece": 0.1081,
68
+ "mean_conf": 0.5747
69
+ }
70
+ }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
71
  },
72
+ "test_teacher2": {
73
+ "all": {
74
+ "n": 449,
75
+ "accuracy": 0.8196,
76
+ "nll": 0.4304,
77
+ "ece": 0.043,
78
+ "mean_conf": 0.8559
79
+ },
80
+ "by_type": {
81
+ "choice": {
82
+ "n": 257,
83
+ "accuracy": 0.8016,
84
+ "nll": 0.5041,
85
+ "ece": 0.0569,
86
+ "mean_conf": 0.8585
87
+ },
88
+ "noul": {
89
+ "n": 130,
90
+ "accuracy": 0.8692,
91
+ "nll": 0.2489,
92
+ "ece": 0.0479,
93
+ "mean_conf": 0.905
94
+ },
95
+ "score": {
96
+ "n": 62,
97
+ "accuracy": 0.7903,
98
+ "nll": 0.5051,
99
+ "ece": 0.0863,
100
+ "mean_conf": 0.7423
101
+ }
102
+ }
103
  },
104
+ "guard": {
105
+ "all": {
106
+ "n": 2994,
107
+ "accuracy": 0.8186,
108
+ "nll": 0.4866,
109
+ "ece": 0.0168,
110
+ "mean_conf": 0.8344
111
+ },
112
+ "by_type": {
113
+ "choice": {
114
+ "n": 2994,
115
+ "accuracy": 0.8186,
116
+ "nll": 0.4866,
117
+ "ece": 0.0168,
118
+ "mean_conf": 0.8344
119
+ },
120
+ "noul": null,
121
+ "score": null
122
+ }
123
  },
124
+ "val_jb": {
125
+ "all": {
126
+ "n": 5000,
127
+ "accuracy": 0.796,
128
+ "nll": 0.5264,
129
+ "ece": 0.0157,
130
+ "mean_conf": 0.8117
131
+ },
132
+ "by_type": {
133
+ "choice": {
134
+ "n": 3374,
135
+ "accuracy": 0.7647,
136
+ "nll": 0.6258,
137
+ "ece": 0.0289,
138
+ "mean_conf": 0.7935
139
+ },
140
+ "noul": {
141
+ "n": 1273,
142
+ "accuracy": 0.901,
143
+ "nll": 0.2233,
144
+ "ece": 0.0161,
145
+ "mean_conf": 0.8849
146
+ },
147
+ "score": {
148
+ "n": 353,
149
+ "accuracy": 0.7167,
150
+ "nll": 0.6681,
151
+ "ece": 0.0317,
152
+ "mean_conf": 0.722
153
+ }
154
+ }
155
  },
156
+ "val_teacher": {
157
+ "all": {
158
+ "n": 944,
159
+ "accuracy": 0.8379,
160
+ "nll": 0.4169,
161
+ "ece": 0.0369,
162
+ "mean_conf": 0.8619
163
+ },
164
+ "by_type": {
165
+ "choice": {
166
+ "n": 540,
167
+ "accuracy": 0.8315,
168
+ "nll": 0.4643,
169
+ "ece": 0.0464,
170
+ "mean_conf": 0.8547
171
+ },
172
+ "noul": {
173
+ "n": 284,
174
+ "accuracy": 0.8873,
175
+ "nll": 0.2931,
176
+ "ece": 0.0453,
177
+ "mean_conf": 0.9088
178
+ },
179
+ "score": {
180
+ "n": 120,
181
+ "accuracy": 0.75,
182
+ "nll": 0.4969,
183
+ "ece": 0.0576,
184
+ "mean_conf": 0.7831
185
+ }
186
+ }
187
  },
188
+ "val_teacher2": {
189
+ "all": {
190
+ "n": 447,
191
+ "accuracy": 0.8322,
192
+ "nll": 0.4435,
193
+ "ece": 0.0627,
194
+ "mean_conf": 0.8747
195
+ },
196
+ "by_type": {
197
+ "choice": {
198
+ "n": 265,
199
+ "accuracy": 0.8377,
200
+ "nll": 0.4219,
201
+ "ece": 0.0503,
202
+ "mean_conf": 0.8858
203
+ },
204
+ "noul": {
205
+ "n": 120,
206
+ "accuracy": 0.85,
207
+ "nll": 0.3702,
208
+ "ece": 0.0827,
209
+ "mean_conf": 0.9034
210
+ },
211
+ "score": {
212
+ "n": 62,
213
+ "accuracy": 0.7742,
214
+ "nll": 0.6773,
215
+ "ece": 0.1441,
216
+ "mean_conf": 0.772
217
+ }
218
+ }
219
  },
220
+ "cal_human": {
221
+ "all": {
222
+ "n": 1595,
223
+ "accuracy": 0.8564,
224
+ "nll": 0.3892,
225
+ "ece": 0.0205,
226
+ "mean_conf": 0.8717
227
+ },
228
+ "by_type": {
229
+ "choice": {
230
+ "n": 1595,
231
+ "accuracy": 0.8564,
232
+ "nll": 0.3892,
233
+ "ece": 0.0205,
234
+ "mean_conf": 0.8717
235
+ },
236
+ "noul": null,
237
+ "score": null
238
+ }
239
  },
240
+ "val_replay4": {
241
+ "all": {
242
+ "n": 1000,
243
+ "accuracy": 0.889,
244
+ "nll": 0.2848,
245
+ "ece": 0.0297,
246
+ "mean_conf": 0.8702
247
+ },
248
+ "by_type": {
249
+ "choice": {
250
+ "n": 1000,
251
+ "accuracy": 0.889,
252
+ "nll": 0.2848,
253
+ "ece": 0.0297,
254
+ "mean_conf": 0.8702
255
+ },
256
+ "noul": null,
257
+ "score": null
258
+ }
259
  }
260
  },
261
+ "global_T": {
262
+ "heldout_jb": {
263
+ "all": {
264
+ "n": 5000,
265
+ "accuracy": 0.5562,
266
+ "nll": 1.2937,
267
+ "ece": 0.163,
268
+ "mean_conf": 0.7192
269
+ },
270
+ "by_type": {
271
+ "choice": {
272
+ "n": 3293,
273
+ "accuracy": 0.5056,
274
+ "nll": 1.5059,
275
+ "ece": 0.1724,
276
+ "mean_conf": 0.678
277
+ },
278
+ "noul": {
279
+ "n": 1289,
280
+ "accuracy": 0.7114,
281
+ "nll": 0.7112,
282
+ "ece": 0.1463,
283
+ "mean_conf": 0.8577
284
+ },
285
+ "score": {
286
+ "n": 418,
287
+ "accuracy": 0.4761,
288
+ "nll": 1.4183,
289
+ "ece": 0.1402,
290
+ "mean_conf": 0.6163
291
+ }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
292
  }
293
  },
294
+ "test_teacher2": {
295
+ "all": {
296
+ "n": 449,
297
+ "accuracy": 0.8196,
298
+ "nll": 0.4381,
299
+ "ece": 0.0497,
300
+ "mean_conf": 0.8664
301
+ },
302
+ "by_type": {
303
+ "choice": {
304
+ "n": 257,
305
+ "accuracy": 0.8016,
306
+ "nll": 0.5058,
307
+ "ece": 0.058,
308
+ "mean_conf": 0.8596
309
+ },
310
+ "noul": {
311
+ "n": 130,
312
+ "accuracy": 0.8692,
313
+ "nll": 0.2755,
314
+ "ece": 0.0602,
315
+ "mean_conf": 0.9294
316
+ },
317
+ "score": {
318
+ "n": 62,
319
+ "accuracy": 0.7903,
320
+ "nll": 0.4984,
321
+ "ece": 0.0798,
322
+ "mean_conf": 0.7623
323
+ }
 
 
 
324
  }
325
  },
326
+ "guard": {
327
+ "all": {
328
+ "n": 2994,
329
+ "accuracy": 0.8186,
330
+ "nll": 0.4871,
331
+ "ece": 0.0194,
332
+ "mean_conf": 0.836
333
+ },
334
+ "by_type": {
335
+ "choice": {
336
+ "n": 2994,
337
+ "accuracy": 0.8186,
338
+ "nll": 0.4871,
339
+ "ece": 0.0194,
340
+ "mean_conf": 0.836
341
+ },
342
+ "noul": null,
343
+ "score": null
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
344
  }
345
  },
346
+ "val_jb": {
347
+ "all": {
348
+ "n": 5000,
349
+ "accuracy": 0.796,
350
+ "nll": 0.5269,
351
+ "ece": 0.0264,
352
+ "mean_conf": 0.8224
353
+ },
354
+ "by_type": {
355
+ "choice": {
356
+ "n": 3374,
357
+ "accuracy": 0.7647,
358
+ "nll": 0.6267,
359
+ "ece": 0.0303,
360
+ "mean_conf": 0.795
361
+ },
362
+ "noul": {
363
+ "n": 1273,
364
+ "accuracy": 0.901,
365
+ "nll": 0.223,
366
+ "ece": 0.0172,
367
+ "mean_conf": 0.9146
368
+ },
369
+ "score": {
370
+ "n": 353,
371
+ "accuracy": 0.7167,
372
+ "nll": 0.6687,
373
+ "ece": 0.0363,
374
+ "mean_conf": 0.7514
375
+ }
 
 
 
376
  }
377
  },
378
+ "val_teacher": {
379
+ "all": {
380
+ "n": 944,
381
+ "accuracy": 0.8379,
382
+ "nll": 0.433,
383
+ "ece": 0.046,
384
+ "mean_conf": 0.8722
385
+ },
386
+ "by_type": {
387
+ "choice": {
388
+ "n": 540,
389
+ "accuracy": 0.8315,
390
+ "nll": 0.4654,
391
+ "ece": 0.0435,
392
+ "mean_conf": 0.8558
393
+ },
394
+ "noul": {
395
+ "n": 284,
396
+ "accuracy": 0.8873,
397
+ "nll": 0.3405,
398
+ "ece": 0.0634,
399
+ "mean_conf": 0.9331
400
+ },
401
+ "score": {
402
+ "n": 120,
403
+ "accuracy": 0.75,
404
+ "nll": 0.5064,
405
+ "ece": 0.0603,
406
+ "mean_conf": 0.8016
407
+ }
 
 
 
408
  }
409
  },
410
+ "val_teacher2": {
411
+ "all": {
412
+ "n": 447,
413
+ "accuracy": 0.8277,
414
+ "nll": 0.4695,
415
+ "ece": 0.0806,
416
+ "mean_conf": 0.8852
417
+ },
418
+ "by_type": {
419
+ "choice": {
420
+ "n": 265,
421
+ "accuracy": 0.8377,
422
+ "nll": 0.4233,
423
+ "ece": 0.0546,
424
+ "mean_conf": 0.8868
425
+ },
426
+ "noul": {
427
+ "n": 120,
428
+ "accuracy": 0.85,
429
+ "nll": 0.4441,
430
+ "ece": 0.1059,
431
+ "mean_conf": 0.9311
432
+ },
433
+ "score": {
434
+ "n": 62,
435
+ "accuracy": 0.7419,
436
+ "nll": 0.7165,
437
+ "ece": 0.1556,
438
+ "mean_conf": 0.7893
439
+ }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
440
  }
441
  },
442
+ "cal_human": {
443
+ "all": {
444
+ "n": 1595,
445
+ "accuracy": 0.8564,
446
+ "nll": 0.3893,
447
+ "ece": 0.0195,
448
+ "mean_conf": 0.8733
449
+ },
450
+ "by_type": {
451
+ "choice": {
452
+ "n": 1595,
453
+ "accuracy": 0.8564,
454
+ "nll": 0.3893,
455
+ "ece": 0.0195,
456
+ "mean_conf": 0.8733
457
+ },
458
+ "noul": null,
459
+ "score": null
460
  }
461
  },
462
+ "val_replay4": {
463
+ "all": {
464
+ "n": 1000,
465
+ "accuracy": 0.889,
466
+ "nll": 0.2843,
467
+ "ece": 0.0287,
468
+ "mean_conf": 0.8713
469
+ },
470
+ "by_type": {
471
+ "choice": {
472
+ "n": 1000,
473
+ "accuracy": 0.889,
474
+ "nll": 0.2843,
475
+ "ece": 0.0287,
476
+ "mean_conf": 0.8713
477
+ },
478
+ "noul": null,
479
+ "score": null
480
  }
481
+ }
482
+ }
483
+ },
484
+ "fixtures": {
485
+ "general_validation": {
486
+ "this": {
487
+ "rows": 847,
488
+ "accuracy": 0.8500590318772137,
489
+ "macro_accuracy": 0.8552272255409907,
490
+ "groups": 65,
491
+ "nll": 0.43237405419192004,
492
+ "ece": 0.03505942904541031,
493
+ "brier": 0.22353081852050935,
494
+ "mean_confidence": 0.8244136475475226
495
  },
496
+ "paired": {
497
+ "cand4b_B4_e2_tt-cand4b_B4_e2": {
498
+ "accuracy": {
499
+ "delta": 0.0,
500
+ "ci95": [
501
+ 0.0,
502
+ 0.0
503
+ ]
504
+ },
505
+ "nll": {
506
+ "delta": 0.00012970495542040443,
507
+ "ci95": [
508
+ -0.00039863639304211823,
509
+ 0.0005942988193083918
510
+ ]
511
+ }
512
+ },
513
+ "cand4b_B4_e2_tt-v1_4b_p130": {
514
+ "accuracy": {
515
+ "delta": -0.010625737898465172,
516
+ "ci95": [
517
+ -0.021251475796930343,
518
+ 0.0
519
+ ]
520
+ },
521
+ "nll": {
522
+ "delta": 0.015307099810435664,
523
+ "ci95": [
524
+ 0.0030912070436396634,
525
+ 0.028028207832146323
526
+ ]
527
+ }
528
+ },
529
+ "cand4b_B4_e2_tt-v2_4b_p130": {
530
+ "accuracy": {
531
+ "delta": 0.0,
532
+ "ci95": [
533
+ -0.015348288075560802,
534
+ 0.015348288075560802
535
+ ]
536
+ },
537
+ "nll": {
538
+ "delta": 0.013765296214476044,
539
+ "ci95": [
540
+ -0.006225581100979938,
541
+ 0.0345539752628619
542
+ ]
543
+ }
544
  }
545
  }
546
  },
547
+ "typesafe": {
548
+ "this": {
549
+ "rows": 102,
550
+ "accuracy": 0.8431372549019608,
551
+ "macro_accuracy": 0.8464696223316912,
552
+ "groups": 4,
553
+ "nll": 0.44128428424617644,
554
+ "ece": 0.06267209380280735,
555
+ "brier": 0.24132394423929263,
556
+ "mean_confidence": 0.8996653276331285,
557
+ "mean_total_variation_to_reference": 0.19246147475275616
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
558
  },
559
+ "paired": {
560
+ "cand4b_B4_e2_tt-cand4b_B4_e2": {
561
+ "accuracy": {
562
+ "delta": 0.0,
563
+ "ci95": [
564
+ 0.0,
565
+ 0.0
566
+ ]
567
+ },
568
+ "nll": {
569
+ "delta": -0.0017070072924830098,
570
+ "ci95": [
571
+ -0.003544773928864442,
572
+ -8.629601238631066e-05
573
+ ]
574
+ }
575
+ },
576
+ "cand4b_B4_e2_tt-v1_4b_p130": {
577
+ "accuracy": {
578
+ "delta": 0.029411764705882353,
579
+ "ci95": [
580
+ -0.029411764705882353,
581
+ 0.08823529411764706
582
+ ]
583
+ },
584
+ "nll": {
585
+ "delta": -0.16956755666314244,
586
+ "ci95": [
587
+ -0.3236270765220259,
588
+ -0.035752501984007506
589
+ ]
590
+ }
591
+ },
592
+ "cand4b_B4_e2_tt-v2_4b_p130": {
593
+ "accuracy": {
594
+ "delta": -0.0196078431372549,
595
+ "ci95": [
596
+ -0.0784313725490196,
597
+ 0.0392156862745098
598
+ ]
599
+ },
600
+ "nll": {
601
+ "delta": 0.03116252024698791,
602
+ "ci95": [
603
+ -0.06909971743191749,
604
+ 0.13434316002254298
605
+ ]
606
+ }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
607
  }
608
  }
609
  },
610
+ "openjev_external": {
611
+ "this": {
612
+ "rows": 5252,
613
+ "accuracy": 0.6597486671744097,
614
+ "macro_accuracy": 0.8063444444444444,
615
+ "groups": 3,
616
+ "nll": 0.8457289344223444,
617
+ "ece": 0.1539857582574297,
618
+ "brier": 0.5022800176356607,
619
+ "mean_confidence": 0.8111537299704116
 
 
 
620
  },
621
+ "paired": {
622
+ "cand4b_B4_e2_tt-cand4b_B4_e2": {
623
+ "accuracy": {
624
+ "delta": 0.0,
625
+ "ci95": [
626
+ 0.0,
627
+ 0.0
628
+ ]
629
+ },
630
+ "nll": {
631
+ "delta": -0.003617502816960803,
632
+ "ci95": [
633
+ -0.003919285763621306,
634
+ -0.00332506503930084
635
+ ]
636
+ }
637
+ },
638
+ "cand4b_B4_e2_tt-v1_4b_p130": {
639
+ "accuracy": {
640
+ "delta": 0.020944402132520943,
641
+ "ci95": [
642
+ 0.012376237623762377,
643
+ 0.029322162985529324
644
+ ]
645
+ },
646
+ "nll": {
647
+ "delta": -0.047333755296254304,
648
+ "ci95": [
649
+ -0.059559262901540976,
650
+ -0.035193633735333484
651
+ ]
652
+ }
653
+ },
654
+ "cand4b_B4_e2_tt-v2_4b_p130": {
655
+ "accuracy": {
656
+ "delta": -0.007425742574257425,
657
+ "ci95": [
658
+ -0.015993907083015995,
659
+ 0.0017136329017517135
660
+ ]
661
+ },
662
+ "nll": {
663
+ "delta": 0.05671352997929895,
664
+ "ci95": [
665
+ 0.04359541016972144,
666
+ 0.07033520922345497
667
+ ]
668
+ }
669
+ }
670
+ }
671
+ },
672
+ "mind2web": {
673
+ "this": {
674
+ "rows": 1770,
675
+ "accuracy": 0.8751412429378531,
676
+ "macro_accuracy": 0.8773196126201471,
677
+ "groups": 3,
678
+ "nll": 0.37523612083453944,
679
+ "ece": 0.016833950116135952,
680
+ "brier": 0.18402882486455008,
681
+ "mean_confidence": 0.8624463401776923
682
  },
683
+ "paired": {
684
+ "cand4b_B4_e2_tt-cand4b_B4_e2": {
685
+ "accuracy": {
686
+ "delta": 0.0,
687
+ "ci95": [
688
+ 0.0,
689
+ 0.0
690
+ ]
691
+ },
692
+ "nll": {
693
+ "delta": 0.00026909733615400604,
694
+ "ci95": [
695
+ -4.997823592689937e-05,
696
+ 0.0006061414472582659
697
+ ]
698
+ }
699
+ },
700
+ "cand4b_B4_e2_tt-v1_4b_p130": {
701
+ "accuracy": {
702
+ "delta": -0.00847457627118644,
703
+ "ci95": [
704
+ -0.01694915254237288,
705
+ 0.0
706
+ ]
707
+ },
708
+ "nll": {
709
+ "delta": 0.00924747244030173,
710
+ "ci95": [
711
+ -0.004092843962589085,
712
+ 0.022512164121517977
713
+ ]
714
+ }
715
+ },
716
+ "cand4b_B4_e2_tt-v2_4b_p130": {
717
+ "accuracy": {
718
+ "delta": 0.0011299435028248588,
719
+ "ci95": [
720
+ -0.00847457627118644,
721
+ 0.011299435028248588
722
+ ]
723
+ },
724
+ "nll": {
725
+ "delta": -0.018244809747118285,
726
+ "ci95": [
727
+ -0.032040489157299956,
728
+ -0.005027723688309381
729
+ ]
730
+ }
731
+ }
732
+ }
733
+ }
734
+ },
735
+ "games": {
736
+ "cand4b_B4_e2_tt-cand4b_B4_e2|sampled|all": {
737
+ "n": 234,
738
+ "a": 0.2692307692307692,
739
+ "b": 0.2692307692307692,
740
+ "delta": 0.0,
741
+ "ci95": [
742
+ -0.004273504273504274,
743
+ 0.004273504273504274
744
+ ]
745
+ },
746
+ "cand4b_B4_e2_tt-cand4b_B4_e2|sampled|bag_game": {
747
+ "n": 64,
748
+ "a": 0.51953125,
749
+ "b": 0.51953125,
750
+ "delta": 0.0,
751
+ "ci95": [
752
+ -0.01171875,
753
+ 0.01171875
754
+ ]
755
+ },
756
+ "cand4b_B4_e2_tt-cand4b_B4_e2|sampled|stochastic_grid": {
757
+ "n": 64,
758
+ "a": 0.16796875,
759
+ "b": 0.16796875,
760
+ "delta": 0.0,
761
+ "ci95": [
762
+ -0.01171875,
763
+ 0.01171875
764
+ ]
765
+ },
766
+ "cand4b_B4_e2_tt-cand4b_B4_e2|sampled|tic_tac_toe": {
767
+ "n": 74,
768
+ "a": 0.2533783783783784,
769
+ "b": 0.2533783783783784,
770
+ "delta": 0.0,
771
+ "ci95": [
772
+ 0.0,
773
+ 0.0
774
+ ]
775
+ },
776
+ "cand4b_B4_e2_tt-cand4b_B4_e2|sampled|exact_minesweeper": {
777
+ "n": 32,
778
+ "a": 0.0078125,
779
+ "b": 0.0078125,
780
+ "delta": 0.0,
781
+ "ci95": [
782
+ 0.0,
783
+ 0.0
784
+ ]
785
+ },
786
+ "cand4b_B4_e2_tt-cand4b_B4_e2|greedy|all": {
787
+ "n": 234,
788
+ "a": 0.25213675213675213,
789
+ "b": 0.25213675213675213,
790
+ "delta": 0.0,
791
+ "ci95": [
792
+ 0.0,
793
+ 0.0
794
+ ]
795
+ },
796
+ "cand4b_B4_e2_tt-cand4b_B4_e2|greedy|bag_game": {
797
+ "n": 64,
798
+ "a": 0.53125,
799
+ "b": 0.53125,
800
+ "delta": 0.0,
801
+ "ci95": [
802
+ 0.0,
803
+ 0.0
804
+ ]
805
+ },
806
+ "cand4b_B4_e2_tt-cand4b_B4_e2|greedy|stochastic_grid": {
807
+ "n": 64,
808
+ "a": 0.109375,
809
+ "b": 0.109375,
810
+ "delta": 0.0,
811
+ "ci95": [
812
+ 0.0,
813
+ 0.0
814
+ ]
815
+ },
816
+ "cand4b_B4_e2_tt-cand4b_B4_e2|greedy|tic_tac_toe": {
817
+ "n": 74,
818
+ "a": 0.24324324324324326,
819
+ "b": 0.24324324324324326,
820
+ "delta": 0.0,
821
+ "ci95": [
822
+ 0.0,
823
+ 0.0
824
+ ]
825
+ },
826
+ "cand4b_B4_e2_tt-cand4b_B4_e2|greedy|exact_minesweeper": {
827
+ "n": 32,
828
+ "a": 0.0,
829
+ "b": 0.0,
830
+ "delta": 0.0,
831
+ "ci95": [
832
+ 0.0,
833
+ 0.0
834
+ ]
835
+ },
836
+ "cand4b_B4_e2_tt-v1_4b_p130|sampled|all": {
837
+ "n": 234,
838
+ "a": 0.2692307692307692,
839
+ "b": 0.2777777777777778,
840
+ "delta": -0.008547008547008548,
841
+ "ci95": [
842
+ -0.02564102564102564,
843
+ 0.009615384615384616
844
+ ]
845
+ },
846
+ "cand4b_B4_e2_tt-v1_4b_p130|sampled|bag_game": {
847
+ "n": 64,
848
+ "a": 0.51953125,
849
+ "b": 0.56640625,
850
+ "delta": -0.046875,
851
+ "ci95": [
852
+ -0.08984375,
853
+ -0.00390625
854
+ ]
855
+ },
856
+ "cand4b_B4_e2_tt-v1_4b_p130|sampled|stochastic_grid": {
857
+ "n": 64,
858
+ "a": 0.16796875,
859
+ "b": 0.1640625,
860
+ "delta": 0.00390625,
861
+ "ci95": [
862
+ -0.02734375,
863
+ 0.03515625
864
+ ]
865
+ },
866
+ "cand4b_B4_e2_tt-v1_4b_p130|sampled|tic_tac_toe": {
867
+ "n": 74,
868
+ "a": 0.2533783783783784,
869
+ "b": 0.24324324324324326,
870
+ "delta": 0.010135135135135136,
871
+ "ci95": [
872
+ -0.013513513513513514,
873
+ 0.037162162162162164
874
+ ]
875
+ },
876
+ "cand4b_B4_e2_tt-v1_4b_p130|sampled|exact_minesweeper": {
877
+ "n": 32,
878
+ "a": 0.0078125,
879
+ "b": 0.0078125,
880
+ "delta": 0.0,
881
+ "ci95": [
882
+ -0.0234375,
883
+ 0.0234375
884
+ ]
885
+ },
886
+ "cand4b_B4_e2_tt-v1_4b_p130|greedy|all": {
887
+ "n": 234,
888
+ "a": 0.25213675213675213,
889
+ "b": 0.2905982905982906,
890
+ "delta": -0.038461538461538464,
891
+ "ci95": [
892
+ -0.07264957264957266,
893
+ -0.004273504273504274
894
+ ]
895
+ },
896
+ "cand4b_B4_e2_tt-v1_4b_p130|greedy|bag_game": {
897
+ "n": 64,
898
+ "a": 0.53125,
899
+ "b": 0.625,
900
+ "delta": -0.09375,
901
+ "ci95": [
902
+ -0.171875,
903
+ -0.03125
904
+ ]
905
+ },
906
+ "cand4b_B4_e2_tt-v1_4b_p130|greedy|stochastic_grid": {
907
+ "n": 64,
908
+ "a": 0.109375,
909
+ "b": 0.125,
910
+ "delta": -0.015625,
911
+ "ci95": [
912
+ -0.07851562499999987,
913
+ 0.046875
914
+ ]
915
+ },
916
+ "cand4b_B4_e2_tt-v1_4b_p130|greedy|tic_tac_toe": {
917
+ "n": 74,
918
+ "a": 0.24324324324324326,
919
+ "b": 0.2702702702702703,
920
+ "delta": -0.02702702702702703,
921
+ "ci95": [
922
+ -0.0945945945945946,
923
+ 0.04054054054054054
924
+ ]
925
+ },
926
+ "cand4b_B4_e2_tt-v1_4b_p130|greedy|exact_minesweeper": {
927
+ "n": 32,
928
+ "a": 0.0,
929
+ "b": 0.0,
930
+ "delta": 0.0,
931
+ "ci95": [
932
+ 0.0,
933
+ 0.0
934
+ ]
935
+ },
936
+ "cand4b_B4_e2_tt-v2_4b_p130|sampled|all": {
937
+ "n": 234,
938
+ "a": 0.2692307692307692,
939
+ "b": 0.22435897435897437,
940
+ "delta": 0.04487179487179487,
941
+ "ci95": [
942
+ 0.024572649572649572,
943
+ 0.06623931623931624
944
+ ]
945
+ },
946
+ "cand4b_B4_e2_tt-v2_4b_p130|sampled|bag_game": {
947
+ "n": 64,
948
+ "a": 0.51953125,
949
+ "b": 0.37890625,
950
+ "delta": 0.140625,
951
+ "ci95": [
952
+ 0.08203125,
953
+ 0.19921875
954
+ ]
955
+ },
956
+ "cand4b_B4_e2_tt-v2_4b_p130|sampled|stochastic_grid": {
957
+ "n": 64,
958
+ "a": 0.16796875,
959
+ "b": 0.16015625,
960
+ "delta": 0.0078125,
961
+ "ci95": [
962
+ -0.0234375,
963
+ 0.0390625
964
+ ]
965
+ },
966
+ "cand4b_B4_e2_tt-v2_4b_p130|sampled|tic_tac_toe": {
967
+ "n": 74,
968
+ "a": 0.2533783783783784,
969
+ "b": 0.24324324324324326,
970
+ "delta": 0.010135135135135136,
971
+ "ci95": [
972
+ -0.010135135135135136,
973
+ 0.033783783783783786
974
+ ]
975
+ },
976
+ "cand4b_B4_e2_tt-v2_4b_p130|sampled|exact_minesweeper": {
977
+ "n": 32,
978
+ "a": 0.0078125,
979
+ "b": 0.0,
980
+ "delta": 0.0078125,
981
+ "ci95": [
982
+ 0.0,
983
+ 0.0234375
984
+ ]
985
+ },
986
+ "cand4b_B4_e2_tt-v2_4b_p130|greedy|all": {
987
+ "n": 234,
988
+ "a": 0.25213675213675213,
989
+ "b": 0.2777777777777778,
990
+ "delta": -0.02564102564102564,
991
+ "ci95": [
992
+ -0.0641025641025641,
993
+ 0.01282051282051282
994
+ ]
995
+ },
996
+ "cand4b_B4_e2_tt-v2_4b_p130|greedy|bag_game": {
997
+ "n": 64,
998
+ "a": 0.53125,
999
+ "b": 0.484375,
1000
+ "delta": 0.046875,
1001
+ "ci95": [
1002
+ -0.03125,
1003
+ 0.125
1004
+ ]
1005
+ },
1006
+ "cand4b_B4_e2_tt-v2_4b_p130|greedy|stochastic_grid": {
1007
+ "n": 64,
1008
+ "a": 0.109375,
1009
+ "b": 0.1875,
1010
+ "delta": -0.078125,
1011
+ "ci95": [
1012
+ -0.15625,
1013
+ 0.0
1014
+ ]
1015
+ },
1016
+ "cand4b_B4_e2_tt-v2_4b_p130|greedy|tic_tac_toe": {
1017
+ "n": 74,
1018
+ "a": 0.24324324324324326,
1019
+ "b": 0.2972972972972973,
1020
+ "delta": -0.05405405405405406,
1021
+ "ci95": [
1022
+ -0.12162162162162163,
1023
+ 0.02702702702702703
1024
+ ]
1025
+ },
1026
+ "cand4b_B4_e2_tt-v2_4b_p130|greedy|exact_minesweeper": {
1027
+ "n": 32,
1028
+ "a": 0.0,
1029
+ "b": 0.0,
1030
+ "delta": 0.0,
1031
+ "ci95": [
1032
+ 0.0,
1033
+ 0.0
1034
+ ]
1035
+ }
1036
+ },
1037
+ "browser": {
1038
+ "cand4b_B4_e2_tt-cand4b_B4_e2|sampled|all": {
1039
+ "n": 176,
1040
+ "a": 0.9318181818181818,
1041
+ "b": 0.9318181818181818,
1042
+ "delta": 0.0,
1043
+ "ci95": [
1044
+ 0.0,
1045
+ 0.0
1046
+ ]
1047
+ },
1048
+ "cand4b_B4_e2_tt-cand4b_B4_e2|sampled|rewarded": {
1049
+ "n": 128,
1050
+ "a": 0.953125,
1051
+ "b": 0.953125,
1052
+ "delta": 0.0,
1053
+ "ci95": [
1054
+ 0.0,
1055
+ 0.0
1056
+ ]
1057
+ },
1058
+ "cand4b_B4_e2_tt-cand4b_B4_e2|sampled|held_out": {
1059
+ "n": 48,
1060
+ "a": 0.875,
1061
+ "b": 0.875,
1062
+ "delta": 0.0,
1063
+ "ci95": [
1064
+ 0.0,
1065
+ 0.0
1066
+ ]
1067
+ },
1068
+ "cand4b_B4_e2_tt-cand4b_B4_e2|greedy|all": {
1069
+ "n": 176,
1070
+ "a": 0.9375,
1071
+ "b": 0.9375,
1072
+ "delta": 0.0,
1073
+ "ci95": [
1074
+ 0.0,
1075
+ 0.0
1076
+ ]
1077
+ },
1078
+ "cand4b_B4_e2_tt-cand4b_B4_e2|greedy|rewarded": {
1079
+ "n": 128,
1080
+ "a": 0.9609375,
1081
+ "b": 0.9609375,
1082
+ "delta": 0.0,
1083
+ "ci95": [
1084
+ 0.0,
1085
+ 0.0
1086
+ ]
1087
+ },
1088
+ "cand4b_B4_e2_tt-cand4b_B4_e2|greedy|held_out": {
1089
+ "n": 48,
1090
+ "a": 0.875,
1091
+ "b": 0.875,
1092
+ "delta": 0.0,
1093
+ "ci95": [
1094
+ 0.0,
1095
+ 0.0
1096
+ ]
1097
+ },
1098
+ "cand4b_B4_e2_tt-v1_4b_p130|sampled|all": {
1099
+ "n": 176,
1100
+ "a": 0.9318181818181818,
1101
+ "b": 0.9090909090909091,
1102
+ "delta": 0.022727272727272728,
1103
+ "ci95": [
1104
+ -0.017045454545454544,
1105
+ 0.0625
1106
+ ]
1107
+ },
1108
+ "cand4b_B4_e2_tt-v1_4b_p130|sampled|rewarded": {
1109
+ "n": 128,
1110
+ "a": 0.953125,
1111
+ "b": 0.9609375,
1112
+ "delta": -0.0078125,
1113
+ "ci95": [
1114
+ -0.0390625,
1115
+ 0.0234375
1116
+ ]
1117
+ },
1118
+ "cand4b_B4_e2_tt-v1_4b_p130|sampled|held_out": {
1119
+ "n": 48,
1120
+ "a": 0.875,
1121
+ "b": 0.7708333333333334,
1122
+ "delta": 0.10416666666666667,
1123
+ "ci95": [
1124
+ 0.0,
1125
+ 0.20833333333333334
1126
+ ]
1127
+ },
1128
+ "cand4b_B4_e2_tt-v1_4b_p130|greedy|all": {
1129
+ "n": 176,
1130
+ "a": 0.9375,
1131
+ "b": 0.9147727272727273,
1132
+ "delta": 0.022727272727272728,
1133
+ "ci95": [
1134
+ -0.005681818181818182,
1135
+ 0.056818181818181816
1136
+ ]
1137
+ },
1138
+ "cand4b_B4_e2_tt-v1_4b_p130|greedy|rewarded": {
1139
+ "n": 128,
1140
+ "a": 0.9609375,
1141
+ "b": 0.9765625,
1142
+ "delta": -0.015625,
1143
+ "ci95": [
1144
+ -0.0390625,
1145
+ 0.0
1146
+ ]
1147
+ },
1148
+ "cand4b_B4_e2_tt-v1_4b_p130|greedy|held_out": {
1149
+ "n": 48,
1150
+ "a": 0.875,
1151
+ "b": 0.75,
1152
+ "delta": 0.125,
1153
+ "ci95": [
1154
+ 0.041666666666666664,
1155
+ 0.22916666666666666
1156
+ ]
1157
+ },
1158
+ "cand4b_B4_e2_tt-v2_4b_p130|sampled|all": {
1159
+ "n": 176,
1160
+ "a": 0.9318181818181818,
1161
+ "b": 0.8806818181818182,
1162
+ "delta": 0.05113636363636364,
1163
+ "ci95": [
1164
+ 0.005681818181818182,
1165
+ 0.09659090909090909
1166
+ ]
1167
+ },
1168
+ "cand4b_B4_e2_tt-v2_4b_p130|sampled|rewarded": {
1169
+ "n": 128,
1170
+ "a": 0.953125,
1171
+ "b": 0.9296875,
1172
+ "delta": 0.0234375,
1173
+ "ci95": [
1174
+ -0.0234375,
1175
+ 0.0703125
1176
+ ]
1177
+ },
1178
+ "cand4b_B4_e2_tt-v2_4b_p130|sampled|held_out": {
1179
+ "n": 48,
1180
+ "a": 0.875,
1181
+ "b": 0.75,
1182
+ "delta": 0.125,
1183
+ "ci95": [
1184
+ 0.0,
1185
+ 0.25
1186
+ ]
1187
+ },
1188
+ "cand4b_B4_e2_tt-v2_4b_p130|greedy|all": {
1189
+ "n": 176,
1190
+ "a": 0.9375,
1191
+ "b": 0.9261363636363636,
1192
+ "delta": 0.011363636363636364,
1193
+ "ci95": [
1194
+ -0.017045454545454544,
1195
+ 0.045454545454545456
1196
+ ]
1197
+ },
1198
+ "cand4b_B4_e2_tt-v2_4b_p130|greedy|rewarded": {
1199
+ "n": 128,
1200
+ "a": 0.9609375,
1201
+ "b": 0.96875,
1202
+ "delta": -0.0078125,
1203
+ "ci95": [
1204
+ -0.0234375,
1205
+ 0.0
1206
+ ]
1207
+ },
1208
+ "cand4b_B4_e2_tt-v2_4b_p130|greedy|held_out": {
1209
+ "n": 48,
1210
+ "a": 0.875,
1211
+ "b": 0.8125,
1212
+ "delta": 0.0625,
1213
+ "ci95": [
1214
+ -0.041666666666666664,
1215
+ 0.16666666666666666
1216
+ ]
1217
+ }
1218
+ },
1219
+ "text_games_greedy": {
1220
+ "pong": -5.0,
1221
+ "freeway": 0.0,
1222
+ "breakout": 35.0,
1223
+ "frozenlake": 0.0,
1224
+ "cliffwalking": -13.0,
1225
+ "blackjack": -0.6,
1226
+ "minigrid_empty": 0.0,
1227
+ "minigrid_lavagap": 0.0,
1228
+ "minigrid_doorkey": 0.0,
1229
+ "babyai_goto": 0.1915625
1230
+ },
1231
+ "probes": {
1232
+ "router_n": 31,
1233
+ "router_tier": 0.968,
1234
+ "needs_live_data": 0.839,
1235
+ "commands_n": 45,
1236
+ "command_risk": 0.889,
1237
+ "destructive_called_safe": 0,
1238
+ "outside_project": 0.933,
1239
+ "browser_element": 0.938,
1240
+ "browser_action": 0.75,
1241
+ "batteries": {
1242
+ "generic": 1.0,
1243
+ "specific": 1.0,
1244
+ "catchall": 0.95,
1245
+ "noul_acc": 1.0,
1246
+ "noul_brier": 0.01,
1247
+ "n_choice": 60,
1248
+ "n_noul": 49,
1249
+ "abstain_battery": "8/8"
1250
+ }
1251
+ },
1252
+ "issue9": {
1253
+ "c_1": {
1254
+ "pred": 26,
1255
+ "gold": 6,
1256
+ "ok": false,
1257
+ "p_gold": 0.198,
1258
+ "p_pred": 0.764,
1259
+ "T": 1.099
1260
+ },
1261
+ "c_2": {
1262
+ "pred": 20,
1263
+ "gold": 20,
1264
+ "ok": true,
1265
+ "p_gold": 0.409,
1266
+ "p_pred": 0.409,
1267
+ "T": 1.099
1268
+ },
1269
+ "b_10": {
1270
+ "pred": 0,
1271
+ "gold": 0,
1272
+ "ok": true,
1273
+ "p_gold": 0.732,
1274
+ "p_pred": 0.732,
1275
+ "T": 1.099
1276
+ },
1277
+ "a2_4": {
1278
+ "pred": 4,
1279
+ "gold": 4,
1280
+ "ok": true,
1281
+ "p_gold": 0.577,
1282
+ "p_pred": 0.577,
1283
+ "T": 1.099
1284
+ }
1285
+ },
1286
+ "bespoke": {
1287
+ "vitaminc-dev": {
1288
+ "n": 599,
1289
+ "type": "choice",
1290
+ "acc": 0.7429,
1291
+ "ece": 0.1456,
1292
+ "brier": 0.3945,
1293
+ "nll": 0.7789,
1294
+ "trained": false,
1295
+ "sec": 11.8
1296
+ },
1297
+ "massive-en-US": {
1298
+ "n": 350,
1299
+ "type": "choice",
1300
+ "acc": 0.8629,
1301
+ "ece": 0.0411,
1302
+ "brier": 0.1936,
1303
+ "nll": 0.4173,
1304
+ "trained": true,
1305
+ "sec": 6.2
1306
+ },
1307
+ "massive-de-DE": {
1308
+ "n": 350,
1309
+ "type": "choice",
1310
+ "acc": 0.84,
1311
+ "ece": 0.0297,
1312
+ "brier": 0.2269,
1313
+ "nll": 0.4947,
1314
+ "trained": false,
1315
+ "sec": 4.5
1316
+ },
1317
+ "boolq": {
1318
+ "n": 300,
1319
+ "type": "noul",
1320
+ "acc": 0.8667,
1321
+ "ece": 0.0319,
1322
+ "brier": 0.1862,
1323
+ "nll": 0.3007,
1324
+ "trained": true,
1325
+ "sec": 6.2
1326
+ },
1327
+ "squad2": {
1328
+ "n": 299,
1329
+ "type": "noul",
1330
+ "acc": 0.7893,
1331
+ "ece": 0.0441,
1332
+ "brier": 0.2955,
1333
+ "nll": 0.452,
1334
+ "trained": false,
1335
+ "sec": 7.2
1336
+ },
1337
+ "paws": {
1338
+ "n": 250,
1339
+ "type": "noul",
1340
+ "acc": 0.764,
1341
+ "ece": 0.1238,
1342
+ "brier": 0.3398,
1343
+ "nll": 0.5538,
1344
+ "trained": true,
1345
+ "sec": 3.0
1346
+ },
1347
+ "multinli": {
1348
+ "n": 299,
1349
+ "type": "choice",
1350
+ "acc": 0.9331,
1351
+ "ece": 0.0103,
1352
+ "brier": 0.1089,
1353
+ "nll": 0.2099,
1354
+ "trained": true,
1355
+ "sec": 3.1
1356
+ },
1357
+ "civil_comments": {
1358
+ "n": 300,
1359
+ "type": "noul",
1360
+ "acc": 0.8733,
1361
+ "ece": 0.0566,
1362
+ "brier": 0.1917,
1363
+ "nll": 0.3199,
1364
+ "trained": true,
1365
+ "sec": 3.2
1366
+ },
1367
+ "aegis2": {
1368
+ "n": 250,
1369
+ "type": "noul",
1370
+ "acc": 0.804,
1371
+ "ece": 0.0467,
1372
+ "brier": 0.2797,
1373
+ "nll": 0.4527,
1374
+ "trained": false,
1375
+ "sec": 6.1
1376
+ },
1377
+ "helpsteer2": {
1378
+ "n": 249,
1379
+ "type": "score",
1380
+ "acc": 0.4498,
1381
+ "ece": 0.142,
1382
+ "brier": 0.6713,
1383
+ "nll": 1.2502,
1384
+ "trained": true,
1385
+ "sec": 25.8,
1386
+ "score_mae": 0.797
1387
+ },
1388
+ "summeval-relevance": {
1389
+ "n": 240,
1390
+ "type": "score",
1391
+ "acc": 0.3958,
1392
+ "ece": 0.0696,
1393
+ "brier": 0.6844,
1394
+ "nll": 1.2541,
1395
+ "trained": false,
1396
+ "sec": 20.1,
1397
+ "score_mae": 0.6172
1398
+ },
1399
+ "summeval-consistency": {
1400
+ "n": 144,
1401
+ "type": "score",
1402
+ "acc": 0.7986,
1403
+ "ece": 0.3095,
1404
+ "brier": 0.4741,
1405
+ "nll": 0.9007,
1406
+ "trained": false,
1407
+ "sec": 12.2,
1408
+ "score_mae": 0.7974
1409
+ },
1410
+ "pubmedqa": {
1411
+ "n": 250,
1412
+ "type": "choice",
1413
+ "acc": 0.708,
1414
+ "ece": 0.0863,
1415
+ "brier": 0.4012,
1416
+ "nll": 0.7205,
1417
+ "trained": true,
1418
+ "sec": 3.8
1419
+ },
1420
+ "_macro": 0.756,
1421
+ "_micro": 0.7652,
1422
+ "_macro_untrained": 0.7284
1423
+ },
1424
+ "jevbench_public": {
1425
+ "easy": {
1426
+ "n": 48,
1427
+ "accuracy": 1.0,
1428
+ "invalid": 0,
1429
+ "ece": 0.0027979166666670663,
1430
+ "by_family": {
1431
+ "extraction": {
1432
+ "n": 12,
1433
+ "accuracy": 1.0
1434
+ },
1435
+ "fact": {
1436
+ "n": 12,
1437
+ "accuracy": 1.0
1438
+ },
1439
+ "intent": {
1440
+ "n": 12,
1441
+ "accuracy": 1.0
1442
+ },
1443
+ "tool_selection": {
1444
+ "n": 12,
1445
+ "accuracy": 1.0
1446
+ }
1447
+ }
1448
+ },
1449
+ "standard": {
1450
+ "n": 72,
1451
+ "accuracy": 0.9861111111111112,
1452
+ "invalid": 0,
1453
+ "ece": 0.0753152777777778,
1454
+ "by_family": {
1455
+ "adequacy": {
1456
+ "n": 12,
1457
+ "accuracy": 0.9166666666666666
1458
+ },
1459
+ "extraction": {
1460
+ "n": 12,
1461
+ "accuracy": 1.0
1462
+ },
1463
+ "intent": {
1464
+ "n": 12,
1465
+ "accuracy": 1.0
1466
+ },
1467
+ "ordinal": {
1468
+ "n": 12,
1469
+ "accuracy": 1.0
1470
+ },
1471
+ "policy": {
1472
+ "n": 12,
1473
+ "accuracy": 1.0
1474
+ },
1475
+ "routing": {
1476
+ "n": 12,
1477
+ "accuracy": 1.0
1478
+ }
1479
+ }
1480
+ },
1481
+ "hard": {
1482
+ "n": 111,
1483
+ "accuracy": 0.6486486486486487,
1484
+ "invalid": 0,
1485
+ "ece": 0.18370720720720732,
1486
+ "by_family": {
1487
+ "adversarial": {
1488
+ "n": 6,
1489
+ "accuracy": 1.0
1490
+ },
1491
+ "ambiguous": {
1492
+ "n": 7,
1493
+ "accuracy": 0.7142857142857143
1494
+ },
1495
+ "judge_hard": {
1496
+ "n": 17,
1497
+ "accuracy": 0.5294117647058824
1498
+ },
1499
+ "long_policy": {
1500
+ "n": 19,
1501
+ "accuracy": 0.47368421052631576
1502
+ },
1503
+ "multi_hop": {
1504
+ "n": 18,
1505
+ "accuracy": 0.8333333333333334
1506
+ },
1507
+ "probability": {
1508
+ "n": 10,
1509
+ "accuracy": 0.8
1510
+ },
1511
+ "routing_hard": {
1512
+ "n": 5,
1513
+ "accuracy": 1.0
1514
+ },
1515
+ "temporal_numeric": {
1516
+ "n": 15,
1517
+ "accuracy": 0.3333333333333333
1518
+ },
1519
+ "tradeoff": {
1520
+ "n": 6,
1521
+ "accuracy": 0.3333333333333333
1522
+ },
1523
+ "trap": {
1524
+ "n": 8,
1525
+ "accuracy": 1.0
1526
+ }
1527
  }
1528
  }
1529
  },
1530
  "references": {
1531
+ "v1_4b": {
1532
+ "reg_in": 0.8338,
1533
+ "reg_out": 0.7876,
1534
+ "reg_in_nll": 0.4044,
1535
+ "reg_out_nll": 0.5583,
1536
+ "reg_in_ece": 0.0269,
1537
+ "reg_out_ece": 0.0708
1538
+ },
1539
+ "v2_4b": {
1540
+ "reg_in": 0.8241,
1541
+ "reg_out": 0.7788,
1542
+ "reg_in_nll": 0.4411,
1543
+ "reg_out_nll": 0.5663,
1544
+ "reg_in_ece": 0.0413,
1545
+ "reg_out_ece": 0.0803
1546
+ }
1547
+ },
1548
+ "speed": "not measured again: the architecture and size equal the parent's, and the temperature map does not change the computation"
1549
  }
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:69e6895461c425c6469cd304838a2e5673141613c2da04782026e37c1481d936
3
  size 8411558400
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ee8ce585b3cedd93206dd149b09b4bdd683174874211f85c77a36090b90c9fdd
3
  size 8411558400
tokenizer_config.json CHANGED
@@ -10,7 +10,7 @@
10
  "errors": "replace",
11
  "image_token": "<|image_pad|>",
12
  "is_local": true,
13
- "local_files_only": true,
14
  "model_max_length": 262144,
15
  "model_specific_special_tokens": {
16
  "audio_bos_token": "<|audio_start|>",
 
10
  "errors": "replace",
11
  "image_token": "<|image_pad|>",
12
  "is_local": true,
13
+ "local_files_only": false,
14
  "model_max_length": 262144,
15
  "model_specific_special_tokens": {
16
  "audio_bos_token": "<|audio_start|>",