decider-4b v2.1
Browse filesdecider-4b v1 + merged LoRA (rank 64, attention and MLP) with replay trained toward v1's distribution; temperature 1.099, temperature_by_type choice 1.110 / noul 1.560 / score 1.287 (decider-ai 1.4.0). v2 is under the tag v2, v1 under v1.
- README.md +265 -302
- decider/calibrate.py +222 -0
- decider/engine.py +7 -4
- decider/engine_v2.py +5 -2
- decider/infer.py +49 -28
- decider/schema_engine.py +7 -3
- decider/serve.py +26 -9
- decider/shared_prefix.py +5 -1
- decider/systemone.py +64 -6
- decider/temperature.py +144 -0
- decider_config.json +9 -3
- eval_results.json +1505 -1752
- model.safetensors +1 -1
- tokenizer_config.json +1 -1
README.md
CHANGED
|
@@ -17,93 +17,126 @@ Base model: [Qwen/Qwen3.5-4B-Base](https://huggingface.co/Qwen/Qwen3.5-4B-Base):
|
|
| 17 |
attention and 24 with gated delta-net linear attention, hidden size 2,560. Two supervised stages. Stage 1 (decider-4b v1): one
|
| 18 |
pass of cross-entropy on the slot readout over mixture v2, the public decision mixture of
|
| 19 |
[decider-2b](https://huggingface.co/Mapika/decider-2b) plus 26 further public decision datasets and ten programmatically
|
| 20 |
-
generated families with verifiable gold (742M tokens). Stage 2 (v2): a LoRA of rank 64 on the attention and MLP weights,
|
| 21 |
-
trained for 2 epochs on 29,
|
| 22 |
-
reinforcement-learning stage. **This repository holds v2**, the bf16
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 68 |
|
| 69 |
**Regressions, stated plainly.**
|
| 70 |
-
*
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 91 |
|
| 92 |
```python
|
| 93 |
from huggingface_hub import snapshot_download
|
| 94 |
from decider.infer import Decider
|
| 95 |
-
d = Decider(snapshot_download("Mapika/decider-4b", revision="
|
| 96 |
```
|
| 97 |
|
| 98 |
For the HTTP server, set `DECIDER_MODEL` to the same downloaded folder.
|
| 99 |
|
| 100 |
-
**How v2 was chosen.** The training run had a pre-registered rule for replacing v1: the candidate had to beat an earlier
|
| 101 |
-
candidate (the Qwen3.5-4B instruct model plus a LoRA on 15,768 of the same hard-decision rows) on our own held-out hard sets. v2 did not meet it:
|
| 102 |
-
it was 0.8 points short on the held-out teacher-written set and 0.5 points short on the held-out generated families, that is, it
|
| 103 |
-
tied the earlier candidate there. The earlier candidate had lost v1's everyday skills and v2 keeps them within about 1 point, so
|
| 104 |
-
v2 was then measured on every set of this card, and the release was decided on that full comparison with v1. The JevBench public
|
| 105 |
-
items were read once for v2, after selection; they were not used for training, selection or the temperature.
|
| 106 |
-
|
| 107 |
## The decider family
|
| 108 |
|
| 109 |
All six repositories share one interface (`decider.infer.Decider`, `POST /v1/systemone` in TypeSafe's format) and one
|
|
@@ -111,8 +144,8 @@ readout: the letter logits at an answer slot, softmaxed over the options. Pick b
|
|
| 111 |
|
| 112 |
| model | base | weights | use it for | numbers |
|
| 113 |
|---|---|---|---|---|
|
| 114 |
-
| [decider-2b](https://huggingface.co/Mapika/decider-2b)
|
| 115 |
-
| [decider-4b](https://huggingface.co/Mapika/decider-4b) v2 | Qwen3.5-4B-Base | 8.4 GB bf16 | the middle point: knowledge questions and hard judgments above the 2B in a dense 8.4 GB model; no RL stage; v1 under the
|
| 116 |
| [decider-35b-a3b](https://huggingface.co/Mapika/decider-35b-a3b) v1 | Qwen3.5-35B-A3B-Base (3B active) | 65 GB bf16 | when accuracy is worth 3 to 4 times the cost per decision: knowledge and multi-step questions, long policies | 0.855 / 0.810, above the 2B on 93 of 95 tasks; JevBench hard 0.676; Bespoke 0.774; no RL stage |
|
| 117 |
| [decider-35b-a3b-nvfp4](https://huggingface.co/Mapika/decider-35b-a3b-nvfp4) | the 35B in NVFP4 | 19.6 GB | the 35B on Blackwell through vLLM or TensorRT-LLM | 1.0 to 1.5 points under bf16 on the measured fixtures |
|
| 118 |
| [decider-0.8b](https://huggingface.co/Mapika/decider-0.8b) | Qwen3.5-0.8B-Base | 1.4 GB bf16 | the smallest: routing, yes/no and short-state lookups within 1 to 4 points of the 2B, 1.5x faster | 0.776 / 0.707 on the single-run protocol (2B: 0.809 / 0.739) |
|
|
@@ -133,19 +166,22 @@ d.decide("My card was charged twice for the same purchase.",
|
|
| 133 |
|
| 134 |
The API is the same as decider-2b's: `decide_batch` scores many states with many questions in one call, `abstain_below=t`
|
| 135 |
returns `None` under a confidence threshold, a question can have 2 to 255 options, and `system_one` / `decider.serve` accept
|
| 136 |
-
TypeSafe's `POST /v1/systemone` request shape
|
| 137 |
-
|
| 138 |
-
|
| 139 |
|
| 140 |
Requirements: `torch`, `transformers>=5`, and `flash-linear-attention` (Triton kernels for the Qwen3.5 linear-attention layers;
|
| 141 |
-
the model runs without it but several times slower). The weights take 8.4 GB in bf16. v2 uses the plain prompt layout, as v1
|
| 142 |
-
|
| 143 |
-
|
| 144 |
-
probabilities
|
| 145 |
-
the
|
| 146 |
-
|
| 147 |
-
|
| 148 |
-
|
|
|
|
|
|
|
|
|
|
| 149 |
|
| 150 |
Without the helper package, the same computation in plain `transformers`:
|
| 151 |
|
|
@@ -159,7 +195,7 @@ ids = tok(prompt, return_tensors="pt").to("cuda")
|
|
| 159 |
with torch.no_grad():
|
| 160 |
logits = m(**ids).logits[0, -1]
|
| 161 |
letters = [tok.encode(L, add_special_tokens=False)[0] for L in "ABC"]
|
| 162 |
-
probs = torch.softmax(logits[letters].float() / 1.
|
| 163 |
```
|
| 164 |
|
| 165 |
## How it works
|
|
@@ -167,8 +203,13 @@ probs = torch.softmax(logits[letters].float() / 1.935, -1) # 1.935 is the st
|
|
| 167 |
The prompt is `Context: ...` followed by, for each question, the question text, the lettered options `(A) ... (B) ...` and an
|
| 168 |
answer slot `Answer k: (`. The hidden state at each slot is projected with the option-letter rows of the LM head and softmaxed
|
| 169 |
over the valid letters, divided by the temperature in `decider_config.json`. Letters are never generated, so all slots are read
|
| 170 |
-
from one pass.
|
| 171 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 172 |
|
| 173 |
## Training
|
| 174 |
|
|
@@ -201,250 +242,171 @@ schedule, only the optimizer changed), AdamW on the bf16 parameters beat AdamW w
|
|
| 201 |
through and moves the weights further from the base model; without it, updates below the bf16 resolution round away and more of
|
| 202 |
the base model's knowledge is kept.
|
| 203 |
|
| 204 |
-
**Stage 2 (v2).** A LoRA of rank 64 (alpha 128) on the attention and MLP weights of v1, trained
|
| 205 |
-
|
| 206 |
-
weights:
|
| 207 |
|
| 208 |
-
| source | rows | content |
|
| 209 |
-
|---|---|---|
|
| 210 |
-
| generated decision families | 8,000 | ten families (temporal and numeric decisions, subtle answer judgment, long policies, multi-hop lookup, abstention, probability, constrained trade-offs, safety judgment, paraphrase sensitivity, adversarial traps); the answers are computed by the generating code |
|
| 211 |
-
| questions over business documents, written by Qwen3.6-27B with thinking on | 11,356 | in two rounds, one realistic business document (up to 34 domains and 22 document kinds) plus three or four typed questions per writer call; each question was answered twice more by the same model in fresh contexts, with shuffled options and without the writer's answer, and kept only when both answers agreed with the writer's (89% and 91% kept) |
|
| 212 |
-
| human-labelled public sets (training halves) | 3,
|
| 213 |
-
| replay of mixture v2 | 6,
|
|
|
|
|
|
|
|
|
|
| 214 |
|
| 215 |
| stage 2 | |
|
| 216 |
|---|---|
|
| 217 |
| trainable parameters | LoRA rank 64, alpha 128, on the attention and MLP projections; merged after training |
|
| 218 |
-
|
|
| 219 |
-
|
|
|
|
|
| 220 |
|
| 221 |
No JevBench item and no Decision Index item was used for training, for writing the generators or the document questions, for
|
| 222 |
-
selecting the checkpoint, or for the
|
| 223 |
written from the family names that JevBench publishes for its sealed set, not from its items. Every training row was checked
|
| 224 |
-
against every evaluation file used for selection
|
| 225 |
-
|
| 226 |
-
|
| 227 |
-
|
| 228 |
-
|
| 229 |
-
|
| 230 |
-
|
| 231 |
-
|
| 232 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 233 |
|
| 234 |
## Evaluation
|
| 235 |
|
| 236 |
-
All v2 numbers are
|
| 237 |
-
|
| 238 |
|
| 239 |
**Public regression set**, rebuilt on this machine (95 tasks: 67 in-task, 28 held-out; large label sets sub-sampled to 10
|
| 240 |
-
options;
|
| 241 |
-
error with 15 bins
|
| 242 |
|
| 243 |
| model | in-task acc / NLL / ECE (67 tasks) | held-out acc / NLL / ECE (28 tasks) |
|
| 244 |
|---|---|---|
|
| 245 |
-
| decider-2b v10, T=1.30 | 0.
|
|
|
|
| 246 |
| decider-4b v1, T=1.05 | 0.834 / 0.404 / 0.027 | 0.788 / 0.558 / 0.071 |
|
| 247 |
-
|
|
|
|
|
| 248 |
| decider-35b-a3b v1, T=1.08 | 0.855 / 0.357 / 0.026 | 0.810 / 0.497 / 0.069 |
|
| 249 |
|
| 250 |
-
|
| 251 |
-
|
| 252 |
-
emotion −3.6, ADE −2.9, abstention probe −2.3, SST-5 −2.1, HelpSteer3 preference −2.1, LIAR −2.1, the others under 2). The
|
| 253 |
-
largest gains are on knowledge and reasoning tasks: MedQA +13.4, MedMCQA +12.3, Winogrande +11.7, TruthfulQA +11.6, MMLU +11.3,
|
| 254 |
-
OpenBookQA +8.8, StrategyQA +8.0. Mean accuracy over nine knowledge tasks (MMLU, ARC, HellaSwag, MedQA, MedMCQA, Social IQa,
|
| 255 |
-
COPA, TruthfulQA, TREC): 0.797 (v1 0.801) against v10's 0.707 and the 35B's 0.862. Against the 35B: −3.1 in-task (−4.0 to −2.3)
|
| 256 |
-
and −3.2 held-out points (−4.5 to −1.8), lower on 85 of 95 tasks and higher on 7; the largest gaps are MedQA −17.7, HelpSteer3
|
| 257 |
-
preference −16.6, MedMCQA −11.9, StrategyQA −10.8, TruthfulQA −10.0.
|
| 258 |
|
| 259 |
-
|
| 260 |
-
|
|
|
|
| 261 |
|
| 262 |
-
|
| 263 |
-
|
| 264 |
-
|
| 265 |
-
|
| 266 |
-
| arena_pref | 0.483 / 0.189 | 0.503 / 0.218 | 0.521 / 0.121 |
|
| 267 |
-
| bbc_news | 0.927 / 0.013 | 0.951 / 0.068 | 0.944 / 0.027 |
|
| 268 |
-
| cb | 0.857 / 0.093 | 0.839 / 0.118 | 0.893 / 0.084 |
|
| 269 |
-
| cr_reviews | 0.903 / 0.031 | 0.898 / 0.046 | 0.914 / 0.033 |
|
| 270 |
-
| dbpedia_l2 | 0.950 / 0.018 | 0.951 / 0.017 | 0.961 / 0.010 |
|
| 271 |
-
| dbpedia_l3 | 0.987 / 0.005 | 0.991 / 0.012 | 0.992 / 0.004 |
|
| 272 |
-
| dolly_category | 0.299 / 0.203 | 0.374 / 0.142 | 0.354 / 0.098 |
|
| 273 |
-
| fin_phrasebank | 0.694 / 0.042 | 0.699 / 0.077 | 0.759 / 0.110 |
|
| 274 |
-
| fin_sentiment | 0.793 / 0.058 | 0.826 / 0.051 | 0.839 / 0.136 |
|
| 275 |
-
| hermes_tools | 0.723 / 0.208 | 0.737 / 0.124 | 0.799 / 0.085 |
|
| 276 |
-
| hwu64 | 0.961 / 0.030 | 0.963 / 0.037 | 0.975 / 0.022 |
|
| 277 |
-
| massive_scenario | 0.756 / 0.041 | 0.793 / 0.052 | 0.799 / 0.027 |
|
| 278 |
-
| offtopic_probe | 0.841 / 0.027 | 0.823 / 0.043 | 0.870 / 0.038 |
|
| 279 |
-
| paws | 0.724 / 0.145 | 0.794 / 0.124 | 0.729 / 0.169 |
|
| 280 |
-
| pubmedqa | 0.756 / 0.085 | 0.758 / 0.055 | 0.820 / 0.078 |
|
| 281 |
-
| quality | 0.494 / 0.233 | 0.561 / 0.179 | 0.632 / 0.096 |
|
| 282 |
-
| quality_full | 0.508 / 0.198 | 0.505 / 0.188 | 0.565 / 0.112 |
|
| 283 |
-
| reward_bench | 0.819 / 0.045 | 0.853 / 0.032 | 0.919 / 0.024 |
|
| 284 |
-
| sciq | 0.982 / 0.024 | 0.991 / 0.021 | 0.993 / 0.011 |
|
| 285 |
-
| social_iqa | 0.708 / 0.077 | 0.759 / 0.062 | 0.823 / 0.025 |
|
| 286 |
-
| strategyqa | 0.552 / 0.138 | 0.632 / 0.157 | 0.739 / 0.036 |
|
| 287 |
-
| student_questions | 0.925 / 0.045 | 0.940 / 0.052 | 0.954 / 0.090 |
|
| 288 |
-
| trec | 0.784 / 0.066 | 0.842 / 0.027 | 0.832 / 0.160 |
|
| 289 |
-
| truthfulqa | 0.537 / 0.090 | 0.654 / 0.063 | 0.754 / 0.068 |
|
| 290 |
-
| tweet_irony | 0.795 / 0.052 | 0.827 / 0.027 | 0.861 / 0.129 |
|
| 291 |
-
| xstory_cloze | 0.962 / 0.017 | 0.971 / 0.017 | 0.995 / 0.016 |
|
| 292 |
-
|
| 293 |
-
</details>
|
| 294 |
-
|
| 295 |
-
**On the same rows as decider-2b v10 and decider-35b-a3b.** Every row below is scored by all three models on identical inputs
|
| 296 |
-
and seeds. Intervals are 95% paired bootstrap intervals (rows for the fixtures, boards for the games, task-seed pairs for the
|
| 297 |
-
browser). The table under [Changes from v1](#changes-from-v1) has v1 on the same rows.
|
| 298 |
-
|
| 299 |
-
| | decider-2b v10 | **decider-4b v2** | decider-35b-a3b | 4B minus 2B | 4B minus 35B |
|
| 300 |
-
|---|---|---|---|---|---|
|
| 301 |
-
| 847 in-task validation rows, accuracy / NLL | 83.2% / 0.444 | 85.0% / 0.419 | 90.0% / 0.329 | +1.8 (−0.5 to +3.9) | −5.0 (−7.1 to −2.8) |
|
| 302 |
-
| OpenJev, 5,252 rows, accuracy / NLL | 63.3% / 0.916 | 66.7% / 0.789 | 68.3% / 0.752 | +3.5 (+2.2 to +4.7) | −1.6 (−2.8 to −0.3) |
|
| 303 |
-
| Mind2Web element and action choice, 1,770 rows | 82.7% / 0.543 | 87.4% / 0.393 | 89.6% / 0.316 | +4.8 (+3.1 to +6.4) | −2.2 (−3.7 to −0.6) |
|
| 304 |
-
| TypeSafe workflow decisions, 102 rows, accuracy / NLL | 80.4% / 0.585 | 86.3% / 0.410 | 86.3% / 0.342 | +5.9 (−2.9 to +14.7) | 0.0 (−5.9 to +4.9) |
|
| 305 |
-
| Bespoke's public suite, 13 subsets, macro / micro | 0.704 / 0.711 | 0.773 / 0.781 | 0.774 / 0.787 | | |
|
| 306 |
-
| JevBench public items, easy / standard / hard accuracy | 1.000 / 0.889 / 0.459 | 1.000 / 0.986 / 0.676 | 1.000 / 0.972 / 0.676 | | |
|
| 307 |
-
| live MiniWoB++ click tasks, 22 tasks x 8 seeds, greedy play | 90.9% | 92.6% | 97.2% | +1.7 (−4.0 to +7.4) | −4.5 (−8.0 to −1.7) |
|
| 308 |
-
| the same, 6 tasks v10 never used for reward, greedy | 91.7% | 81.2% | 97.9% | −10.4 (−22.9 to +2.1) | −16.7 (−27.1 to −6.2) |
|
| 309 |
-
| live MiniWoB++ click tasks, sampled play | 93.2% | 88.1% | 86.4% | −5.1 (−10.8 to +0.6) | +1.7 (−3.4 to +6.8) |
|
| 310 |
-
| the same, 6 held-out tasks, sampled | 91.7% | 75.0% | 79.2% | −16.7 (−31.2 to −4.2) | −4.2 (−16.7 to +8.3) |
|
| 311 |
-
| zero-shot games, win rate, sampled play (234 boards) | 23.7% | 22.4% | 24.1% | −1.3 (−4.2 to +1.5) | −1.7 (−3.7 to +0.2) |
|
| 312 |
-
| zero-shot games, greedy play | 26.5% | 27.8% | 37.2% | +1.3 (−4.3 to +6.8) | −9.4 (−15.0 to −3.8) |
|
| 313 |
-
| bag-draw games alone, sampled | 41.4% | 37.9% | 41.8% | −3.5 (−9.0 to +2.0) | −3.9 (−8.6 to +0.8) |
|
| 314 |
-
|
| 315 |
-
On the two fixtures whose rows are in neither model's training data (TypeSafe, OpenJev) v2 is above the 2B (OpenJev +3.5, interval
|
| 316 |
-
excludes zero; TypeSafe +5.9, interval includes zero) and on TypeSafe level with the 35B. Stage 2 contains teacher-written and
|
| 317 |
-
human-labelled judgment rows, not TypeSafe or OpenJev rows. The browser rows show the missing RL stage: greedy play is level
|
| 318 |
-
with v10 over all tasks and 10 points below it on the six tasks that v10's RL never rewarded (interval includes zero), where the
|
| 319 |
-
4B's weakest tasks are focus-text-2 (0.375) and click-collapsible-2 (0.5). On the bag-draw games v1 won about 15 points more
|
| 320 |
-
often than either other model in sampled play; v2 does not (37.9% against 41.4% and 41.8%). On tic-tac-toe, the slippery grid
|
| 321 |
-
and minesweeper every model is near the random floor.
|
| 322 |
-
|
| 323 |
-
**Ten text games, zero-shot** (`decider.games.play`, five episodes per game, greedy, eager; the 2B was trained on the first
|
| 324 |
-
four, the 4B and the 35B on none): Pong −21 (teacher 8, 2B 8, 35B −21), Breakout 12 (22, 22, 6), CliffWalking −60 (teacher −13,
|
| 325 |
-
2B −13, 35B −1,248), MiniGrid-Empty 0 (teacher 0.96, both others 0), Freeway 1 (teacher 5, 2B 0, 35B 1), FrozenLake 0 (teacher 1,
|
| 326 |
-
others 0), Blackjack −0.6 (teacher −0.6, 2B −1, 35B −0.6), MiniGrid-LavaGap 0, MiniGrid-DoorKey 0 (teacher 0), BabyAI-GoTo 0.35
|
| 327 |
-
(teacher 0.34, 2B 0, 35B 0.19). v1 reached the teacher on CliffWalking (−13) and read 0.54 on BabyAI-GoTo; v2 keeps neither.
|
| 328 |
-
Neither version learns Pong from the text state. Greedy play takes the most probable option, so these results do not depend on
|
| 329 |
-
the temperature.
|
| 330 |
|
| 331 |
**JevBench public items** (231 items of [Benchmark Heaven](https://benchmarkheaven.com/jev-models); argmax over the exact
|
| 332 |
-
label set with the request the harness's TypeSafe adapter builds).
|
| 333 |
-
|
| 334 |
-
|
| 335 |
-
0.
|
| 336 |
-
|
| 337 |
-
|
| 338 |
-
temporal-numeric 0.27 (0.13). Top-label ECE is 0.008 / 0.085 / 0.071 by tier (v1, read again on 2026-09-24: 0.001 / 0.047 / 0.288;
|
| 339 |
-
the 35B's hard tier 0.151). On the hard tier the mean confidence is 0.74 at accuracy 0.68; 31% of hard items are answered at confidence 0.9 or more,
|
| 340 |
-
with accuracy 0.94 on them (v1: 50% at 0.73). The 18 Score items were answered with isolated levels, whose per-level values are
|
| 341 |
-
not stored, so they keep their T 1.719 probabilities inside these ECE values. The public hard tier is 111 items (95% interval
|
| 342 |
-
about ±9 points), and the gain over v1 there (+12.6) is larger than the gain on our own held-out hard sets.
|
| 343 |
|
| 344 |
**Bespoke's public suite** (13 human-labelled subsets, 3,880 records in Jev's wire format, answered through `system_one` as
|
| 345 |
-
shipped).
|
|
|
|
| 346 |
|
| 347 |
-
| subset (type) | decider-
|
| 348 |
-
|---|---|---|---|
|
| 349 |
-
| vitaminc-dev (choice) | 0.
|
| 350 |
-
| massive-en-US (choice; trained) | 0.
|
| 351 |
-
| massive-de-DE (choice
|
| 352 |
-
| boolq (noul; trained) | 0.
|
| 353 |
-
| squad2 (noul) | 0.
|
| 354 |
-
| paws (noul; trained) | 0.
|
| 355 |
-
| multinli (choice; trained) | 0.
|
| 356 |
-
| civil_comments (noul; trained) | 0.
|
| 357 |
-
| aegis2 (noul) | 0.
|
| 358 |
-
| helpsteer2 (score; trained) | 0.
|
| 359 |
-
| summeval-relevance (score) | 0.
|
| 360 |
-
| summeval-consistency (score) | 0.
|
| 361 |
-
| pubmedqa (choice; trained) | 0.
|
| 362 |
-
| **macro / micro** | 0.
|
| 363 |
-
|
| 364 |
-
|
| 365 |
-
|
| 366 |
-
|
| 367 |
-
|
| 368 |
-
|
| 369 |
-
|
| 370 |
-
|
| 371 |
-
|
| 372 |
-
|
| 373 |
-
|
| 374 |
-
five rating datasets, and the per-level fits sum to between 0.92 and 1.00 (this probe reads the raw logits, T=1).
|
| 375 |
|
| 376 |
## Calibration
|
| 377 |
|
| 378 |
-
The
|
| 379 |
-
|
| 380 |
-
|
| 381 |
-
|
| 382 |
-
|
| 383 |
-
|
| 384 |
-
does not fit every task type. Outside the regression set v2 is better calibrated than v1 on most sets:
|
| 385 |
|
| 386 |
-
| set |
|
| 387 |
-
|---|---|---|---|---|
|
| 388 |
-
|
|
| 389 |
-
|
|
| 390 |
-
|
|
| 391 |
-
|
|
| 392 |
-
|
| 393 |
-
|
| 394 |
-
|
| 395 |
-
|
| 396 |
-
|
| 397 |
-
|
| 398 |
-
|
| 399 |
-
computed with the site's definition (benchmark-weighted pooled bins, ten bins; the same computation gives 0.027 on the 35B
|
| 400 |
-
against the site's published 0.031). v2 was read once at the candidate temperature 1.719 (ECE 0.090, sample index 48.3); the
|
| 401 |
-
chosen answers do not change with the temperature, and the calibration rows above are recomputed at 1.935 from the stored
|
| 402 |
-
probabilities. The JevBench hard-tier value is recomputed the same way, except its 6 Score items, which keep their T 1.719
|
| 403 |
-
probabilities. The sample index was not recomputed. Nothing was fitted on index rows. v2's confidence exceeds its accuracy by
|
| 404 |
-
0.053 on that sample (v1 0.085, the 2B 0.092; the 35B has no gap). The gap is per-benchmark heterogeneity, not a global scale: a
|
| 405 |
-
single temperature that removes it on the knowledge benchmarks would make the wide label sets underconfident. If you route on
|
| 406 |
-
confidence, calibrate on your own labels.
|
| 407 |
|
| 408 |
## Speed
|
| 409 |
|
| 410 |
-
v2 has the same architecture and size as v1
|
| 411 |
-
|
| 412 |
-
|
| 413 |
-
|
| 414 |
-
|
| 415 |
-
|
| 416 |
-
|
| 417 |
-
|
| 418 |
-
|
| 419 |
-
requests per second at a median of 13.3 ms with one client and 190 requests per second with 64 clients.
|
| 420 |
|
| 421 |
## Limitations
|
| 422 |
|
| 423 |
-
* No reinforcement-learning stage: stated beliefs about action outcomes were not trained against exact laws
|
| 424 |
-
|
| 425 |
-
|
| 426 |
-
|
| 427 |
-
|
| 428 |
-
|
| 429 |
-
|
| 430 |
-
|
| 431 |
-
|
| 432 |
-
|
| 433 |
-
|
| 434 |
-
* Below the 35B
|
| 435 |
-
|
| 436 |
-
* The abstention probe is 2.4 points under the 2B (0.582 against 0.606), with ECE 0.133.
|
| 437 |
-
* The JevBench public hard tier is 111 items; v2's gain there (+12.6 points over v1) is larger than its gain on our own held-out
|
| 438 |
-
hard sets, where it tied an earlier candidate. Do not read it as a gain of that size on hard items in general.
|
| 439 |
* The stage-2 LoRA was trained only in the plain state-first layout. The schema-first layout (the server's opt-in schema cache)
|
| 440 |
-
was not measured on v2, and `decider_config.json` does not mark
|
| 441 |
-
`DECIDER_SCHEMA_CACHE=1` does not turn the schema cache on for this model.
|
| 442 |
* Mixture v2's 26 additional public datasets and ten programmatic families, and stage 2's generators and document questions, are
|
| 443 |
described above but their builders are not in the public package; `scripts/train.sh full` reproduces the 60% of stage 1's data
|
| 444 |
-
that is the public mixture.
|
| 445 |
-
0.28, plans 0.44, probability 0.43, code 0.88, policy 0.99); this was not measured on v2.
|
| 446 |
* English is the main language; the multilingual rows (XNLI, PAWS-X, MASSIVE, Belebele, XCOPA) are a small share of the data
|
| 447 |
-
and were not measured beyond the mixture-v2 evaluation set.
|
| 448 |
* Everything else in the decider-2b card's limitations (packed questions see each other, long JSON arrays by position, full
|
| 449 |
label sets against sampled options, abstention wording, rules in the question) applies; those shapes were not re-measured at
|
| 450 |
this size.
|
|
@@ -453,7 +415,8 @@ requests per second at a median of 13.3 ms with one client and 190 requests per
|
|
| 453 |
|
| 454 |
| version | what changed |
|
| 455 |
|---|---|
|
| 456 |
-
| **v2** (2026-09-24, these weights) | v1 + a merged LoRA
|
|
|
|
| 457 |
| v1 (2026-09-22, Hub tag `v1`) | first release: one pass over mixture v2 on Qwen3.5-4B-Base with AdamW on bf16 parameters, no RL stage; temperature 1.05 |
|
| 458 |
|
| 459 |
The GitHub repository's [docs/CHANGELOG.md](https://github.com/Mapika/decider/blob/main/docs/CHANGELOG.md) lists every
|
|
@@ -464,10 +427,10 @@ decider release.
|
|
| 464 |
Code, data registry, training and evaluation scripts and the per-version history: https://github.com/Mapika/decider
|
| 465 |
(`docs/HISTORY.md`, section "decider-4b"). Stage 1 was trained with the data-parallel trainer of the architecture A/B study
|
| 466 |
(`arch_ab/train_dp_optvar.py` in the research repository, optimizer variant `bf16`); stage 2 with a LoRA trainer in the research
|
| 467 |
-
repository. Both were evaluated with the public `decider.evaluate` and the head-to-head tools
|
| 468 |
-
|
| 469 |
-
|
| 470 |
-
|
| 471 |
|
| 472 |
**Independence.** This is an independent project. It is not affiliated with or endorsed by TypeSafe AI. It is an open
|
| 473 |
reproduction of the "System One" model class (TypeSafe AI's Jev); nothing was distilled from Jev. The training data is public
|
|
|
|
| 17 |
attention and 24 with gated delta-net linear attention, hidden size 2,560. Two supervised stages. Stage 1 (decider-4b v1): one
|
| 18 |
pass of cross-entropy on the slot readout over mixture v2, the public decision mixture of
|
| 19 |
[decider-2b](https://huggingface.co/Mapika/decider-2b) plus 26 further public decision datasets and ten programmatically
|
| 20 |
+
generated families with verifiable gold (742M tokens). Stage 2 (v2.1): a LoRA of rank 64 on the attention and MLP weights,
|
| 21 |
+
trained for 2 epochs on 29,325 rows of harder decisions and replay, with the replay rows trained toward v1's own answer
|
| 22 |
+
distribution, then merged into the weights. There is no reinforcement-learning stage. **This repository holds v2.1**, the bf16
|
| 23 |
+
weights (8.4 GB), with one temperature per answer type. v2 stays available under the Hub tag `v2` and v1 under the tag `v1`
|
| 24 |
+
(see [Changes from v2](#changes-from-v2) for who should keep using them). The other sizes are listed under The decider family.
|
| 25 |
+
`decider/` in this repository is the inference subset of the GitHub package.
|
| 26 |
+
|
| 27 |
+
v2.1 keeps most of v2's gain on hard decisions and gets back most of what v2 lost in sampled play. Against v2 on the same rows:
|
| 28 |
+
bag-draw games in sampled play 52.0% against 37.9% wins (v1 56.6%), zero-shot games sampled 26.9% against 22.4% (v1 27.8%), live
|
| 29 |
+
browser tasks sampled 93.2% against 88.1% (v1 90.9%), CliffWalking −13 against −60, the regression set 0.831 / 0.784 against 0.824
|
| 30 |
+
/ 0.779 (v1 0.834 / 0.788); held-out generated decision families 0.556 against 0.560 (v1 0.469) and the JevBench public hard tier
|
| 31 |
+
0.649 against 0.676 (v1 0.550). It is less well calibrated than v2 on hard items (details under
|
| 32 |
+
[Calibration](#calibration)), still answers one of the two form-filling cases of issue #9 wrongly, and is worse than v1 on
|
| 33 |
+
BabyAI-GoTo and on greedy bag-draw play. Against decider-35b-a3b it is 2.4 and 2.6 points lower on the regression set.
|
| 34 |
+
|
| 35 |
+
**Contents:** [Changes from v2](#changes-from-v2) · [The decider family](#the-decider-family) · [Usage](#usage) · [How it works](#how-it-works) · [Training](#training) · [Evaluation](#evaluation) · [Calibration](#calibration) · [Speed](#speed) · [Limitations](#limitations) · [Changelog](#changelog) · [Reproduction](#reproduction)
|
| 36 |
+
|
| 37 |
+
## Changes from v2
|
| 38 |
+
|
| 39 |
+
v2.1 differs from v2 in two things:
|
| 40 |
+
|
| 41 |
+
* **Replay rows are trained toward v1's own distribution.** Stage 2 of v2 trained every row, including 6,700 replay rows from v1's
|
| 42 |
+
training mixture, on its hard label. v2.1 trains the replay rows toward v1's answer distribution instead, with the loss
|
| 43 |
+
KL(p_v1 ‖ p_model) over the options of each answer, and keeps hard labels on every other row. Everything else is v2's recipe:
|
| 44 |
+
the same rows (29,325 after removing 31 rows that shared a context with an evaluation file; v2 had 29,356), LoRA rank 64, learning
|
| 45 |
+
rate 1e-4, 2 epochs. Training the replay rows on hard labels had sharpened the logits everywhere; the fitted temperature then
|
| 46 |
+
rose to 1.935 and flattened every served answer, which is what cost v2 its sampled play. With the replay rows trained toward
|
| 47 |
+
v1, the fitted temperature is 1.099 (v1 1.05). On 1,000 held-out replay rows the mean KL(p_v1 ‖ p_model) at temperature 1 is 0.021
|
| 48 |
+
nats (v2: 0.112), and 96.3% of the argmax answers equal v1's (v2: 94.1%).
|
| 49 |
+
* **One temperature per answer type.** `decider_config.json` has `temperature` 1.099 and `temperature_by_type`
|
| 50 |
+
`{"choice": 1.110, "noul": 1.560, "score": 1.287}`, fitted by NLL with `decider.calibrate` (decider-ai 1.4.0). decider-ai 1.4.0
|
| 51 |
+
and later use the map. **decider-ai 1.3.0 and earlier ignore the map and serve every answer at 1.099**; a temperature does not
|
| 52 |
+
change which option is most probable, so the answers are the same either way (except exact ties between the level rows of an
|
| 53 |
+
isolated Score answer, where rounding can break the tie differently: 1 of 5,000 held-out generated-family rows and 2 of 447
|
| 54 |
+
document-question validation rows changed); the
|
| 55 |
+
probabilities differ, mostly on yes/no and Score answers. The map changes little in practice: sampled
|
| 56 |
+
play, the fixtures and the probes are the same within noise with and without it, and it lowers the calibration error on
|
| 57 |
+
held-out yes/no answers (0.146 to 0.101) and Score answers (0.140 to 0.108). It does not fix the overconfidence on hard
|
| 58 |
+
Choice answers (see Calibration).
|
| 59 |
+
|
| 60 |
+
All rows below are on identical inputs and seeds: v2.1 through decider-ai 1.4.0 with its map, v1 and v2 through decider-ai 1.3.0
|
| 61 |
+
at their stored temperatures (1.05 and 1.935), measured in the same session on 2026-09-24. Intervals are 95% paired bootstrap
|
| 62 |
+
intervals (rows for the fixtures, boards for the games, task-seed pairs for the browser). The regression set and the two held-out
|
| 63 |
+
sets were read from stored temperature-1 logits at each model's served temperatures. The JevBench files of v1 and v2 were read
|
| 64 |
+
earlier through decider-ai 1.2.1, v2's at the candidate temperature 1.719; accuracy does not depend on the temperature.
|
| 65 |
+
|
| 66 |
+
| set | v1 (T 1.05) | v2 (T 1.935) | v2.1 (map) | v2.1 minus v1 | v2.1 minus v2 |
|
| 67 |
+
|---|---|---|---|---|---|
|
| 68 |
+
| regression set, 67 in-task tasks, accuracy / NLL / ECE | 0.834 / 0.404 / 0.027 | 0.824 / 0.441 / 0.041 | 0.831 / 0.414 / 0.031 | −0.3 | +0.7 |
|
| 69 |
+
| regression set, 28 held-out tasks | 0.788 / 0.558 / 0.071 | 0.779 / 0.566 / 0.080 | 0.784 / 0.569 / 0.077 | −0.4 | +0.5 |
|
| 70 |
+
| 847 in-task validation rows, accuracy / NLL | 86.1% / 0.417 | 85.0% / 0.419 | 85.0% / 0.432 | −1.1 (−2.1 to +0.0); NLL +0.015 (+0.003 to +0.028) | 0.0 (−1.5 to +1.5); NLL +0.014 (−0.006 to +0.035) |
|
| 71 |
+
| OpenJev, 5,252 rows, accuracy / NLL | 63.9% / 0.893 | 66.7% / 0.789 | 66.0% / 0.846 | +2.1 (+1.2 to +2.9); NLL −0.047 (−0.060 to −0.035) | −0.7 (−1.6 to +0.2); NLL +0.057 (+0.044 to +0.070) |
|
| 72 |
+
| Mind2Web, 1,770 rows, accuracy / NLL | 88.4% / 0.366 | 87.4% / 0.393 | 87.5% / 0.375 | −0.8 (−1.7 to +0.0); NLL +0.009 (−0.004 to +0.023) | +0.1 (−0.8 to +1.1); NLL −0.018 (−0.032 to −0.005) |
|
| 73 |
+
| TypeSafe workflow decisions, 102 rows, accuracy / NLL | 81.4% / 0.611 | 86.3% / 0.410 | 84.3% / 0.441 | +2.9 (−2.9 to +8.8); NLL −0.170 (−0.324 to −0.036) | −2.0 (−7.8 to +3.9); NLL +0.031 (−0.069 to +0.134) |
|
| 74 |
+
| held-out generated families (heldout_jb), 5,000 rows, accuracy / ECE | 0.469 / 0.260 | 0.560 / 0.046 | 0.556 / 0.147 | +8.7 | −0.4 |
|
| 75 |
+
| held-out document questions (test_teacher2), 449 rows, accuracy / ECE | 0.726 / 0.110 | 0.826 / 0.053 | 0.820 / 0.043 | +9.4 | −0.7 |
|
| 76 |
+
| JevBench public items, easy / standard / hard accuracy | 1.000 / 0.958 / 0.550 | 1.000 / 0.986 / 0.676 | 1.000 / 0.986 / 0.649 | hard +11 items | hard −3 items |
|
| 77 |
+
| JevBench hard tier, top-label ECE (v1, v2: files read through 1.2.1, v2 at T 1.719) | 0.288 | 0.104 | 0.184 | | |
|
| 78 |
+
| Bespoke's public suite, macro / micro | 0.757 / 0.765 | 0.773 / 0.781 | 0.756 / 0.765 | −0.1 macro | −1.6 macro |
|
| 79 |
+
| live MiniWoB++, sampled, all 22 tasks | 90.9% | 88.1% | 93.2% | +2.3 (−1.7 to +6.2) | +5.1 (+0.6 to +9.7) |
|
| 80 |
+
| live MiniWoB++, sampled, 16 rewarded tasks | 96.1% | 93.0% | 95.3% | −0.8 (−3.9 to +2.3) | +2.3 (−2.3 to +7.0) |
|
| 81 |
+
| live MiniWoB++, sampled, 6 held-out tasks | 77.1% | 75.0% | 87.5% | +10.4 (0.0 to +20.8) | +12.5 (0.0 to +25.0) |
|
| 82 |
+
| live MiniWoB++, greedy, all 22 tasks | 91.5% | 92.6% | 93.8% | +2.3 (−0.6 to +5.7) | +1.1 (−1.7 to +4.5) |
|
| 83 |
+
| live MiniWoB++, greedy, 16 rewarded tasks | 97.7% | 96.9% | 96.1% | −1.6 (−3.9 to +0.0) | −0.8 (−2.3 to +0.0) |
|
| 84 |
+
| live MiniWoB++, greedy, 6 held-out tasks | 75.0% | 81.2% | 87.5% | +12.5 (+4.2 to +22.9) | +6.2 (−4.2 to +16.7) |
|
| 85 |
+
| zero-shot games, 234 boards, sampled, win rate | 27.8% | 22.4% | 26.9% | −0.9 (−2.6 to +1.0) | +4.5 (+2.5 to +6.6) |
|
| 86 |
+
| bag-draw games, 64 boards, sampled, win rate | 56.6% | 37.9% | 52.0% | −4.7 (−9.0 to −0.4) | +14.1 (+8.2 to +19.9) |
|
| 87 |
+
| slippery-grid games, 64 boards, sampled, win rate | 16.4% | 16.0% | 16.8% | +0.4 (−2.7 to +3.5) | +0.8 (−2.3 to +3.9) |
|
| 88 |
+
| zero-shot games, 234 boards, greedy, win rate | 29.1% | 27.8% | 25.2% | −3.8 (−7.3 to −0.4) | −2.6 (−6.4 to +1.3) |
|
| 89 |
+
| bag-draw games, 64 boards, greedy, win rate | 62.5% | 48.4% | 53.1% | −9.4 (−17.2 to −3.1) | +4.7 (−3.1 to +12.5) |
|
| 90 |
+
| slippery-grid games, 64 boards, greedy, win rate | 12.5% | 18.8% | 10.9% | −1.6 (−7.9 to +4.7) | −7.8 (−15.6 to +0.0) |
|
| 91 |
+
| ten text games, greedy: Pong / Breakout / CliffWalking / BabyAI-GoTo / Freeway / Blackjack | −21 / 14 / −13 / 0.54 / 0 / −0.6 | −21 / 12 / −60 / 0.35 / 1 / −0.6 | −5 / 35 / −13 / 0.19 / 0 / −0.6 | | |
|
| 92 |
+
| behaviour probes: model-router tier / needs-live-data (31 items) | 0.968 / 0.871 | 0.935 / 0.839 | 0.968 / 0.839 | 0 / −1 item | +1 / 0 items |
|
| 93 |
+
| behaviour probes: command risk / touches-outside-project (45 items) | 0.889 / 0.956 | 0.911 / 0.933 | 0.889 / 0.933 | 0 / −1 item | −1 / 0 items |
|
| 94 |
+
| behaviour probes: generic bucket / catch-all / abstention battery / browser element and action | 1.00 / 0.90 / 7 of 8 / 0.875 and 0.875 | 0.95 / 0.95 / 8 of 8 / 0.938 and 0.938 | 1.00 / 0.95 / 8 of 8 / 0.938 and 0.750 | | |
|
| 95 |
+
| issue #9 form cases c_1 / c_2 (probability of the gold option) | right (0.79) / right (0.98) | wrong (0.16) / right (0.28) | wrong (0.20) / right (0.41) | | |
|
| 96 |
|
| 97 |
**Regressions, stated plainly.**
|
| 98 |
+
* BabyAI-GoTo (greedy text game): 0.19 against v1's 0.54 and v2's 0.35.
|
| 99 |
+
* Greedy bag-draw play: 53.1% against v1's 62.5% wins (−9.4 points, interval −17.2 to −3.1); greedy zero-shot games overall 25.2%
|
| 100 |
+
against 29.1% (−3.8, interval −7.3 to −0.4). Sampled bag-draw play is 4.7 points under v1 (interval −9.0 to −0.4).
|
| 101 |
+
* Behaviour probes: needs-live-data 0.839 against v1's 0.871 and touches-outside-project 0.933 against 0.956, one item each (as
|
| 102 |
+
v2); the browser-agent action choice is 12 of 16 items (0.750), 2 fewer than v1 (0.875) and 3 fewer than v2
|
| 103 |
+
(0.938).
|
| 104 |
+
* Form filling: issue #9 case c_1 is still answered wrongly. The field is "Degree earned" and the document entity is "Studied:
|
| 105 |
+
Associate of Arts"; v2.1 chooses "skip" at 0.76 and gives the gold entity 0.20 (v1: gold 0.79; v2: gold 0.16). Case c_2 is
|
| 106 |
+
answered correctly at 0.41. For form filling we suggest v1.
|
| 107 |
+
* Calibration on hard items: on the held-out generated families the calibration error is 0.147 against v2's 0.046 (see
|
| 108 |
+
Calibration), and on the JevBench public hard tier 0.184 against v2's 0.104.
|
| 109 |
+
* JevBench public hard tier: 72 of 111 items (0.649) against v2's 75 (0.676).
|
| 110 |
+
* Against v2: TypeSafe −2.0 points and OpenJev −0.7 (intervals include zero), NLL higher on both; Bespoke's suite 0.756 against
|
| 111 |
+
0.773 macro (v1 0.757); greedy slippery-grid play 10.9% against 18.8%.
|
| 112 |
+
|
| 113 |
+
**How v2.1 was chosen, and why it is released although it did not pass.** The run had pre-registered release rules. v2.1 passed
|
| 114 |
+
the numeric items (held-out document questions 0.820 against the required 0.816, held-out generated families 0.556 against 0.550,
|
| 115 |
+
regression accuracy and calibration not worse than v2's, sampled bag-draw above the midpoint of v1 and v2, sampled zero-shot
|
| 116 |
+
games and browser play not shown to be below v1's (upper end of the 95% interval at least 0), each probe at most one item below v1) and failed the last item, which required both
|
| 117 |
+
issue #9 form cases to be right: c_1 is wrong. By the rule it was not recommended. A second pre-registered rule for the
|
| 118 |
+
temperature map required the calibration error on the held-out generated families to be at most 0.08; the map reaches 0.147
|
| 119 |
+
(0.163 without it), so the map did not pass either. v2.1 is released with the map on a decision made after reading the full comparison above: it is better than v2 on sampled play, CliffWalking, the model-router probe and the regression set, at v2's
|
| 120 |
+
level on our held-out hard sets, and its failures are listed in this section. The JevBench public items were read once for the
|
| 121 |
+
candidate at its global temperature and once with the map, after the rule decisions; they were not used for training,
|
| 122 |
+
selection or the temperatures.
|
| 123 |
+
|
| 124 |
+
**Which version to use.**
|
| 125 |
+
* v2.1 (this revision): the default. Sampled play (games, browser agents that sample actions), hard decisions.
|
| 126 |
+
* v2 (`revision="v2"`): if you rely on confidence values on hard multi-step items, where v2 is better calibrated (held-out
|
| 127 |
+
generated families 0.046, JevBench hard tier 0.104), or on its slightly higher JevBench hard tier and TypeSafe accuracy.
|
| 128 |
+
* v1 (`revision="v1"`): form filling (issue #9), BabyAI-GoTo-like grid navigation, greedy bag-draw play.
|
| 129 |
+
|
| 130 |
+
The package loads a local folder, so download the revision first:
|
| 131 |
|
| 132 |
```python
|
| 133 |
from huggingface_hub import snapshot_download
|
| 134 |
from decider.infer import Decider
|
| 135 |
+
d = Decider(snapshot_download("Mapika/decider-4b", revision="v2")) # or revision="v1"
|
| 136 |
```
|
| 137 |
|
| 138 |
For the HTTP server, set `DECIDER_MODEL` to the same downloaded folder.
|
| 139 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 140 |
## The decider family
|
| 141 |
|
| 142 |
All six repositories share one interface (`decider.infer.Decider`, `POST /v1/systemone` in TypeSafe's format) and one
|
|
|
|
| 144 |
|
| 145 |
| model | base | weights | use it for | numbers |
|
| 146 |
|---|---|---|---|---|
|
| 147 |
+
| [decider-2b](https://huggingface.co/Mapika/decider-2b) v11 | Qwen3.5-2B-Base | 3.8 GB bf16 | the default: routing, classification, judgments, browser agents; 4 ms per request with CUDA graphs on one GPU; v10 under the tag `v10` | regression set 0.802 in-task / 0.752 held-out; JevBench hard 0.577; live browser 90% sampled; Bespoke suite 0.706 |
|
| 148 |
+
| [decider-4b](https://huggingface.co/Mapika/decider-4b) v2.1 | Qwen3.5-4B-Base | 8.4 GB bf16 | the middle point: knowledge questions and hard judgments above the 2B in a dense 8.4 GB model; no RL stage; v2 and v1 under the tags `v2` and `v1` | 0.831 / 0.784; JevBench hard 0.649; live browser 93% sampled; Bespoke 0.756 |
|
| 149 |
| [decider-35b-a3b](https://huggingface.co/Mapika/decider-35b-a3b) v1 | Qwen3.5-35B-A3B-Base (3B active) | 65 GB bf16 | when accuracy is worth 3 to 4 times the cost per decision: knowledge and multi-step questions, long policies | 0.855 / 0.810, above the 2B on 93 of 95 tasks; JevBench hard 0.676; Bespoke 0.774; no RL stage |
|
| 150 |
| [decider-35b-a3b-nvfp4](https://huggingface.co/Mapika/decider-35b-a3b-nvfp4) | the 35B in NVFP4 | 19.6 GB | the 35B on Blackwell through vLLM or TensorRT-LLM | 1.0 to 1.5 points under bf16 on the measured fixtures |
|
| 151 |
| [decider-0.8b](https://huggingface.co/Mapika/decider-0.8b) | Qwen3.5-0.8B-Base | 1.4 GB bf16 | the smallest: routing, yes/no and short-state lookups within 1 to 4 points of the 2B, 1.5x faster | 0.776 / 0.707 on the single-run protocol (2B: 0.809 / 0.739) |
|
|
|
|
| 166 |
|
| 167 |
The API is the same as decider-2b's: `decide_batch` scores many states with many questions in one call, `abstain_below=t`
|
| 168 |
returns `None` under a confidence threshold, a question can have 2 to 255 options, and `system_one` / `decider.serve` accept
|
| 169 |
+
TypeSafe's `POST /v1/systemone` request shape. Every question and every Score level is scored in its own row. The state may be a
|
| 170 |
+
string, object or array of up to 32k tokens. See the decider-2b card for the full description of the request shape, field types
|
| 171 |
+
and the schema cache.
|
| 172 |
|
| 173 |
Requirements: `torch`, `transformers>=5`, and `flash-linear-attention` (Triton kernels for the Qwen3.5 linear-attention layers;
|
| 174 |
+
the model runs without it but several times slower). The weights take 8.4 GB in bf16. v2.1 uses the plain prompt layout, as v1 and
|
| 175 |
+
v2 do. The per-type temperatures need decider-ai 1.4.0 or later (or the `decider/` subset in this repository, taken from 1.4.0).
|
| 176 |
+
Checked on CUDA, eager path: decider-ai 1.3.0 loads v2.1 with the name `decider-4b-v2.1` and temperature 1.099 and gives exactly
|
| 177 |
+
the probabilities that 1.4.0 gives with the map switched off; 1.4.0 and the `decider/` subset in this repository load the map
|
| 178 |
+
and give the same probabilities as each other. For Choice and yes/no answers the 1.4.0 output with the map equals
|
| 179 |
+
softmax(log(p) · 1.099 / T_type) of the 1.3.0 output p (to 1e-7 on `decide`, to 8e-5 on the four-decimal `system_one` values);
|
| 180 |
+
an isolated-level Score answer is rescaled per level row (each level's yes/no pair at 1.287 instead of 1.099) before the
|
| 181 |
+
levels are normalised, so it cannot be recomputed from the combined Score probabilities. The 1.4.0 HTTP server with CUDA graphs was checked with `/v1/systemone` and `/decide` requests. Other package
|
| 182 |
+
versions, MPS and the FP8 path were not checked on v2.1. On Blackwell GPUs use 1.0.2 or later (1.0.0 and 1.0.1 have the cuDNN
|
| 183 |
+
attention fault fixed in 1.0.2). The model is dense, so the CUDA-graph engine, `torch.compile` and the FP8 path of the helper
|
| 184 |
+
package apply to it as to decider-2b (`use_graphs=False` selects eager PyTorch).
|
| 185 |
|
| 186 |
Without the helper package, the same computation in plain `transformers`:
|
| 187 |
|
|
|
|
| 195 |
with torch.no_grad():
|
| 196 |
logits = m(**ids).logits[0, -1]
|
| 197 |
letters = [tok.encode(L, add_special_tokens=False)[0] for L in "ABC"]
|
| 198 |
+
probs = torch.softmax(logits[letters].float() / 1.110, -1) # 1.110 is the stored choice temperature (yes/no 1.560, Score 1.287)
|
| 199 |
```
|
| 200 |
|
| 201 |
## How it works
|
|
|
|
| 203 |
The prompt is `Context: ...` followed by, for each question, the question text, the lettered options `(A) ... (B) ...` and an
|
| 204 |
answer slot `Answer k: (`. The hidden state at each slot is projected with the option-letter rows of the LM head and softmaxed
|
| 205 |
over the valid letters, divided by the temperature in `decider_config.json`. Letters are never generated, so all slots are read
|
| 206 |
+
from one pass. From decider-ai 1.4.0 the config may also hold `temperature_by_type`, one temperature per answer type
|
| 207 |
+
(`choice`, `noul`, `score`; a missing type uses `temperature`). This release's config has such a map: Choice answers use 1.110,
|
| 208 |
+
yes/no (noul) answers 1.560, Score answers 1.287 on each of their level rows. Package versions before 1.4.0 use `temperature`
|
| 209 |
+
(1.099) for every answer. `DECIDER_TEMPERATURE=T` or `Decider(path, temperature=T)` replaces `temperature` with T and switches the
|
| 210 |
+
map off, so every answer then uses T. Large label sets were sub-sampled to at most 10 options per
|
| 211 |
+
training example (gold always kept, order shuffled), so the model conditions on the supplied candidates rather than on a fixed
|
| 212 |
+
head.
|
| 213 |
|
| 214 |
## Training
|
| 215 |
|
|
|
|
| 242 |
through and moves the weights further from the base model; without it, updates below the bf16 resolution round away and more of
|
| 243 |
the base model's knowledge is kept.
|
| 244 |
|
| 245 |
+
**Stage 2 (v2.1).** A LoRA of rank 64 (alpha 128) on the attention and MLP weights of v1, trained for 2 epochs over 29,325 rows
|
| 246 |
+
in v1's plain state-first layout with isolated Score levels, then merged into the bf16 weights:
|
|
|
|
| 247 |
|
| 248 |
+
| source | rows | content | target |
|
| 249 |
+
|---|---|---|---|
|
| 250 |
+
| generated decision families | 8,000 | ten families (temporal and numeric decisions, subtle answer judgment, long policies, multi-hop lookup, abstention, probability, constrained trade-offs, safety judgment, paraphrase sensitivity, adversarial traps); the answers are computed by the generating code | the label |
|
| 251 |
+
| questions over business documents, written by Qwen3.6-27B with thinking on | 11,356 | in two rounds, one realistic business document (up to 34 domains and 22 document kinds) plus three or four typed questions per writer call; each question was answered twice more by the same model in fresh contexts, with shuffled options and without the writer's answer, and kept only when both answers agreed with the writer's (89% and 91% kept) | the label |
|
| 252 |
+
| human-labelled public sets (training halves) | 3,293 | MMLU, ARC, CommonsenseQA, BoolQ, MNLI, SNLI, Banking77, RACE, OpenBookQA, LogiQA 2, MedQA, Winogrande | the label |
|
| 253 |
+
| replay of mixture v2 | 6,676 | 100 rows from the training half of each of the 67 in-task regression tasks | v1's answer distribution: loss KL(p_v1 ‖ p_model) at temperature 1 |
|
| 254 |
+
|
| 255 |
+
These are v2's rows without 31 that the v2.1 checks found to share a context with an evaluation file (7 human-labelled rows and
|
| 256 |
+
24 replay rows).
|
| 257 |
|
| 258 |
| stage 2 | |
|
| 259 |
|---|---|
|
| 260 |
| trainable parameters | LoRA rank 64, alpha 128, on the attention and MLP projections; merged after training |
|
| 261 |
+
| loss | cross-entropy on the slot readout for labelled rows; KL(p_v1 ‖ p_model) over the options for replay rows, with p_v1 from the frozen v1 weights; the step loss is the sum over answers divided by the number of answers |
|
| 262 |
+
| schedule | learning rate 1e-4, 5% warm-up then cosine, 1,518 steps of 65,536 tokens (2 epochs), seed 0 |
|
| 263 |
+
| hardware | one NVIDIA B300, 94 minutes |
|
| 264 |
|
| 265 |
No JevBench item and no Decision Index item was used for training, for writing the generators or the document questions, for
|
| 266 |
+
selecting the checkpoint, or for the temperatures. The ten generated families and the skill list of the document questions were
|
| 267 |
written from the family names that JevBench publishes for its sealed set, not from its items. Every training row was checked
|
| 268 |
+
against every evaluation file used for selection, the canonical examples of the evaluation half of all 153 mixture-v2 evaluation
|
| 269 |
+
tasks, the four request fixtures and the issue #9 cases: there is no exact state-and-question overlap (a few contexts under 40
|
| 270 |
+
characters, such as a short utterance, occur under another task and question). The held-out sets used for selection are
|
| 271 |
+
generated families from held-out templates (heldout_jb) and document questions from business domains that are not in the
|
| 272 |
+
training data (test_teacher2). Two arms were trained to the end (B4 and C4, which differ in the size of the replay; two further arms were
|
| 273 |
+
stopped or cancelled at the user's request); the rule selected B4 after epoch 2.
|
| 274 |
+
|
| 275 |
+
**Temperatures.** `temperature` 1.099, fitted by NLL (all rows pooled) on the in-task half of the public regression set without
|
| 276 |
+
Banking77, CLINC-OOS, MMLU, ARC, Winogrande and HellaSwag: 61 tasks, 102,804 rows. `temperature_by_type` fitted by NLL per answer
|
| 277 |
+
type with `decider.calibrate.fit_by_type` on a pool of those regression rows and our own validation rows (the validation halves of
|
| 278 |
+
the generated families and document questions, human-labelled validation rows, held-out replay rows): 108,927 Choice answers,
|
| 279 |
+
1,677 yes/no answers and 535 Score answers. The Choice pool is 94% regression rows, so the Choice temperature (1.110) stays close
|
| 280 |
+
to the global one. v2's temperature (1.935) and v1's (1.05) were one value each.
|
| 281 |
|
| 282 |
## Evaluation
|
| 283 |
|
| 284 |
+
All v2.1 numbers are with the per-type map through decider-ai 1.4.0, except where a paragraph says otherwise. Greedy play and
|
| 285 |
+
accuracy do not depend on the temperature; sampled play and calibration do.
|
| 286 |
|
| 287 |
**Public regression set**, rebuilt on this machine (95 tasks: 67 in-task, 28 held-out; large label sets sub-sampled to 10
|
| 288 |
+
options; the temperatures fitted on in-task data). The rows are the same for every model. ECE is the expected calibration
|
| 289 |
+
error with 15 bins, per-task mean. The regression rows are all Choice answers, so v2.1 reads them at the Choice temperature 1.110.
|
| 290 |
|
| 291 |
| model | in-task acc / NLL / ECE (67 tasks) | held-out acc / NLL / ECE (28 tasks) |
|
| 292 |
|---|---|---|
|
| 293 |
+
| decider-2b v10, T=1.30 | 0.806 / 0.474 / 0.038 | 0.755 / 0.622 / 0.084 |
|
| 294 |
+
| decider-2b v11, per-type map | 0.802 / 0.481 / 0.038 | 0.752 / 0.626 / 0.083 |
|
| 295 |
| decider-4b v1, T=1.05 | 0.834 / 0.404 / 0.027 | 0.788 / 0.558 / 0.071 |
|
| 296 |
+
| decider-4b v2, T=1.935 | 0.824 / 0.441 / 0.041 | 0.779 / 0.566 / 0.080 |
|
| 297 |
+
| **decider-4b v2.1 (this repository), per-type map** | **0.831 / 0.414 / 0.031** | **0.784 / 0.569 / 0.077** |
|
| 298 |
| decider-35b-a3b v1, T=1.08 | 0.855 / 0.357 / 0.026 | 0.810 / 0.497 / 0.069 |
|
| 299 |
|
| 300 |
+
Per-task paired intervals were not recomputed for v2.1; the v2 card (tag `v2`) has them for v2 against v10 and the 35B, and v2.1
|
| 301 |
+
is within 0.7 points of v2 on both halves.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 302 |
|
| 303 |
+
**The same rows as v1 and v2** are in the table under [Changes from v2](#changes-from-v2): the four fixtures (847 in-task
|
| 304 |
+
validation rows, OpenJev, Mind2Web, TypeSafe), our held-out hard sets, JevBench, Bespoke, live browser tasks, zero-shot games,
|
| 305 |
+
text games, behaviour probes and the issue #9 cases.
|
| 306 |
|
| 307 |
+
**Ten text games, zero-shot** (`decider.games.play`, five episodes per game, greedy, eager; the 2B was trained on the first four,
|
| 308 |
+
the 4B on none): Pong −5 (v1 −21, v2 −21; the 2B 8), Breakout 35 (v1 14, v2 12; the 2B 22), CliffWalking −13 (v1 −13, v2 −60;
|
| 309 |
+
teacher −13), BabyAI-GoTo 0.19 (v1 0.54, v2 0.35), Freeway 0 (v2 1), Blackjack −0.6, FrozenLake, MiniGrid-Empty, LavaGap and DoorKey
|
| 310 |
+
0 as for v1 and v2. Greedy play takes the most probable option, so these results do not depend on the temperature.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 311 |
|
| 312 |
**JevBench public items** (231 items of [Benchmark Heaven](https://benchmarkheaven.com/jev-models); argmax over the exact
|
| 313 |
+
label set with the request the harness's TypeSafe adapter builds). Read once for v2.1 at its global temperature (1.3.0) and once
|
| 314 |
+
with the map (1.4.0), both after the rule decisions: easy 1.000, standard 0.986, hard 0.649 (72 of 111) in both reads. Top-label
|
| 315 |
+
ECE on the hard tier is 0.184 with the map (0.210 at the global temperature); by answer type with the map, Choice 0.156 (67 items),
|
| 316 |
+
yes/no 0.259 (38), Score 0.354 (6). Mean confidence on the hard tier is 0.799 at accuracy 0.649. For comparison on the same items:
|
| 317 |
+
v2 0.676 (ECE 0.104 as read at T 1.719; 0.071 recomputed at its release temperature 1.935 with 6 Score items kept at 1.719), v1
|
| 318 |
+
0.550 (ECE 0.288), decider-35b-a3b 0.676, Jev 1.13.0 0.730. The public hard tier is 111 items (95% interval about ±9 points).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 319 |
|
| 320 |
**Bespoke's public suite** (13 human-labelled subsets, 3,880 records in Jev's wire format, answered through `system_one` as
|
| 321 |
+
shipped). v1 and v2 through decider-ai 1.3.0, v2.1 through 1.4.0 with the map, same session. Nimble-9B (0.748 / 0.759 macro /
|
| 322 |
+
micro) and Jev 1.13.0 (0.760 / 0.773) are in Bespoke's report.
|
| 323 |
|
| 324 |
+
| subset (type) | decider-4b v1 | decider-4b v2 | decider-4b v2.1 |
|
| 325 |
+
|---|---|---|---|
|
| 326 |
+
| vitaminc-dev (choice) | 0.756 | 0.778 | 0.743 |
|
| 327 |
+
| massive-en-US (choice; trained) | 0.860 | 0.869 | 0.863 |
|
| 328 |
+
| massive-de-DE (choice) | 0.843 | 0.837 | 0.840 |
|
| 329 |
+
| boolq (noul; trained) | 0.873 | 0.860 | 0.867 |
|
| 330 |
+
| squad2 (noul) | 0.706 | 0.793 | 0.789 |
|
| 331 |
+
| paws (noul; trained) | 0.716 | 0.832 | 0.764 |
|
| 332 |
+
| multinli (choice; trained) | 0.933 | 0.926 | 0.933 |
|
| 333 |
+
| civil_comments (noul; trained) | 0.880 | 0.857 | 0.873 |
|
| 334 |
+
| aegis2 (noul) | 0.820 | 0.812 | 0.804 |
|
| 335 |
+
| helpsteer2 (score; trained) | 0.466 | 0.466 | 0.450 |
|
| 336 |
+
| summeval-relevance (score) | 0.425 | 0.463 | 0.396 |
|
| 337 |
+
| summeval-consistency (score) | 0.833 | 0.826 | 0.799 |
|
| 338 |
+
| pubmedqa (choice; trained) | 0.728 | 0.724 | 0.708 |
|
| 339 |
+
| **macro / micro** | 0.757 / 0.765 | 0.773 / 0.781 | 0.756 / 0.765 |
|
| 340 |
+
| macro over the subsets not trained on | 0.731 | 0.751 | 0.728 |
|
| 341 |
+
|
| 342 |
+
v2.1 is at v1's level on this suite and 1.6 points under v2 (macro). The largest differences to v2 are PAWS (0.764 against 0.832),
|
| 343 |
+
SummEval relevance (0.396 against 0.463) and VitaminC (0.743 against 0.778).
|
| 344 |
+
|
| 345 |
+
**Behaviour probes** (teacher-labelled, same probes as the other releases; v1 / v2 in brackets): generic-versus-specific bucket
|
| 346 |
+
choice 1.00 / 1.00 (1.00 / 1.00; 0.95 / 1.00), catch-all when nothing fits 0.95 (0.90; 0.95), abstention battery 8 of 8 (7 of 8;
|
| 347 |
+
8 of 8); model-router tier 0.968 (0.968; 0.935) and needs-live-data 0.839 (0.871; 0.839); command-risk classification 0.889
|
| 348 |
+
(0.889; 0.911) with no destructive command called safe, touches-outside-project 0.933 (0.956; 0.933); browser-agent element and
|
| 349 |
+
action choice 0.938 / 0.750 (0.875 / 0.875; 0.938 / 0.938). The yes/no Brier score on the probe batteries is 0.010 with the map
|
| 350 |
+
and 0.006 at the global temperature: the yes/no temperature 1.560 flattens easy yes/no answers there.
|
|
|
|
| 351 |
|
| 352 |
## Calibration
|
| 353 |
|
| 354 |
+
The per-type map was fitted on everyday regression rows and our own validation rows. It helps where it can: yes/no and Score
|
| 355 |
+
answers on the held-out sets. It does not change Choice answers much, because the Choice temperature fitted on a pool that is
|
| 356 |
+
94% everyday rows stays at 1.110. On the held-out generated families, whose answers are two-thirds Choice, v2.1 is
|
| 357 |
+
overconfident: calibration error 0.147 with the map, where our own release limit is 0.08. v2 is at 0.046 there, only because its
|
| 358 |
+
one temperature of 1.935 flattens every answer, which is also what cost it sampled play. ECE below is 10-bin top-label ECE on
|
| 359 |
+
stored temperature-1 logits read at the served temperatures.
|
|
|
|
| 360 |
|
| 361 |
+
| set | rows | accuracy | ECE at the global T / with the map | choice / noul / score ECE with the map |
|
| 362 |
+
|---|---|---|---|---|
|
| 363 |
+
| heldout_jb | 5,000 | 0.556 | 0.163 / 0.147 | 0.170 / 0.101 / 0.108 |
|
| 364 |
+
| test_teacher2 | 449 | 0.820 | 0.050 / 0.043 | 0.057 / 0.048 / 0.086 |
|
| 365 |
+
| guard | 2,994 | 0.819 | 0.019 / 0.017 | 0.017 / − / − |
|
| 366 |
+
| cal_human | 1,595 | 0.856 | 0.019 / 0.021 | 0.021 / − / − |
|
| 367 |
+
|
| 368 |
+
On the held-out generated families: v1 (T 1.05) 0.260, v2 (T 1.935) 0.046, v2.1 at T = 1 0.182. A map fitted without the
|
| 369 |
+
regression rows (Choice 1.259) would reach 0.127 there and would make the regression set less calibrated (in-task ECE 0.0375
|
| 370 |
+
against 0.0309); it was not used. On the fixtures (10-bin ECE, map): 847 validation rows 0.035 (v1 0.037, v2 0.031), TypeSafe
|
| 371 |
+
0.063 (v1 0.124, v2 0.072), OpenJev 0.154 (v1 0.158, v2 0.131), Mind2Web 0.017 (v1 0.016, v2 0.051). The Decision Index
|
| 372 |
+
sample was not read for v2.1. If you route on confidence, calibrate on your own labels; `python -m decider.calibrate` fits a
|
| 373 |
+
per-type map from your own answers read at temperature 1.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 374 |
|
| 375 |
## Speed
|
| 376 |
|
| 377 |
+
v2.1 has the same architecture and size as v1 and v2, and the temperature map is a division per answer; the speed was not
|
| 378 |
+
measured again. The measurements of v1 and v2 (the same weights layout): in one session on one unshared NVIDIA B300 (bf16), 34.6
|
| 379 |
+
ms (v2) and 32.4 ms (v1) median per decision over 200 game-state decisions of 156 tokens median, batch of one, eager PyTorch
|
| 380 |
+
without CUDA graphs or `torch.compile`, timed around the forward pass with `torch.cuda.synchronize()`. The eager path is
|
| 381 |
+
launch-bound, so host load changes it: v1's first measurement was 24.7 ms (10th to 90th percentile 24.6 to 43.5 ms), with
|
| 382 |
+
decider-2b at 17.9 ms and decider-35b-a3b at 41.4 ms on the same decisions and method. With the helper's CUDA graphs and
|
| 383 |
+
`torch.compile`, one support-ticket request (228 tokens, 3 questions) takes 5.2 ms (FP8 5.0 ms); a batch of 32 such states takes
|
| 384 |
+
81.5 ms, 1,178 decisions per second (FP8 71.2 ms, 1,349 per second). The HTTP server (`/decide`, bf16) answers 72.8 requests per
|
| 385 |
+
second at a median of 13.3 ms with one client and 190 requests per second with 64 clients.
|
|
|
|
| 386 |
|
| 387 |
## Limitations
|
| 388 |
|
| 389 |
+
* No reinforcement-learning stage: stated beliefs about action outcomes were not trained against exact laws. On live browser
|
| 390 |
+
tasks v2.1 is at 93.2% sampled and 87.5% on the six held-out tasks; decider-2b v10, which has the RL stage, is at 93.2% and 91.7%.
|
| 391 |
+
* Overconfident on hard multi-step items: calibration error 0.147 on our held-out generated families and 0.184 on the JevBench
|
| 392 |
+
public hard tier, against v2's 0.046 and 0.104. On those items, do not read a confidence of 0.8 as an 80% chance of being right.
|
| 393 |
+
* Issue #9 form case c_1 is answered wrongly (see Changes from v2); BabyAI-GoTo 0.19 against v1's 0.54; greedy bag-draw play 9.4
|
| 394 |
+
points under v1.
|
| 395 |
+
* About 0.3 points under v1 on the regression set and 1.1 points under it on the 847 validation rows; needs-live-data and
|
| 396 |
+
touches-outside-project probes one item under v1 each.
|
| 397 |
+
* The per-type map needs decider-ai 1.4.0 or later. With 1.3.0 or earlier, or with `DECIDER_TEMPERATURE` set, every answer uses
|
| 398 |
+
1.099: the answers are the same, yes/no answers are sharper (on the held-out generated families their ECE is 0.146 instead of
|
| 399 |
+
0.101) and Score answers are sharper.
|
| 400 |
+
* Below the 35B by 2.4 / 2.6 points on the regression set and 2.7 on the JevBench public hard tier (3 items).
|
| 401 |
+
* The JevBench public hard tier is 111 items; differences of a few items between versions are within its noise.
|
|
|
|
|
|
|
|
|
|
| 402 |
* The stage-2 LoRA was trained only in the plain state-first layout. The schema-first layout (the server's opt-in schema cache)
|
| 403 |
+
was not measured on v2.1, and `decider_config.json` does not mark it as trained for that layout (`schema_first_trained:
|
| 404 |
+
false`), so `DECIDER_SCHEMA_CACHE=1` does not turn the schema cache on for this model.
|
| 405 |
* Mixture v2's 26 additional public datasets and ten programmatic families, and stage 2's generators and document questions, are
|
| 406 |
described above but their builders are not in the public package; `scripts/train.sh full` reproduces the 60% of stage 1's data
|
| 407 |
+
that is the public mixture.
|
|
|
|
| 408 |
* English is the main language; the multilingual rows (XNLI, PAWS-X, MASSIVE, Belebele, XCOPA) are a small share of the data
|
| 409 |
+
and were not measured beyond the mixture-v2 evaluation set and Bespoke's German MASSIVE subset.
|
| 410 |
* Everything else in the decider-2b card's limitations (packed questions see each other, long JSON arrays by position, full
|
| 411 |
label sets against sampled options, abstention wording, rules in the question) applies; those shapes were not re-measured at
|
| 412 |
this size.
|
|
|
|
| 415 |
|
| 416 |
| version | what changed |
|
| 417 |
|---|---|
|
| 418 |
+
| **v2.1** (2026-09-24, these weights) | v1 + a merged LoRA on v2's rows (29,325 after removing 31 that overlapped evaluation files), with the replay rows trained toward v1's own answer distribution; temperature 1.099 and `temperature_by_type` {choice 1.110, noul 1.560, score 1.287} (decider-ai 1.4.0; older versions use 1.099). Against v2: sampled play recovered (bag-draw 52.0% against 37.9%, browser 93.2% against 88.1%), CliffWalking −13 against −60, regression set +0.7 / +0.5; hard sets at v2's level; less calibrated on hard items (0.147 against 0.046); issue #9 c_1 still wrong. Did not pass its pre-registered rules; released on the full comparison |
|
| 419 |
+
| v2 (2026-09-24, Hub tag `v2`) | v1 + a merged LoRA (rank 64, attention and MLP, 2 epochs, 29,356 rows: generated decision families, document questions written by Qwen3.6-27B and kept when two independent answers agreed, human-labelled public sets, replay of mixture v2); temperature 1.935. Better on hard judgments, TypeSafe, OpenJev and Bespoke's suite; worse on sampled play, some text games and by about 1 point on everyday tasks |
|
| 420 |
| v1 (2026-09-22, Hub tag `v1`) | first release: one pass over mixture v2 on Qwen3.5-4B-Base with AdamW on bf16 parameters, no RL stage; temperature 1.05 |
|
| 421 |
|
| 422 |
The GitHub repository's [docs/CHANGELOG.md](https://github.com/Mapika/decider/blob/main/docs/CHANGELOG.md) lists every
|
|
|
|
| 427 |
Code, data registry, training and evaluation scripts and the per-version history: https://github.com/Mapika/decider
|
| 428 |
(`docs/HISTORY.md`, section "decider-4b"). Stage 1 was trained with the data-parallel trainer of the architecture A/B study
|
| 429 |
(`arch_ab/train_dp_optvar.py` in the research repository, optimizer variant `bf16`); stage 2 with a LoRA trainer in the research
|
| 430 |
+
repository. Both were evaluated with the public `decider.evaluate` and the head-to-head tools. `eval_results.json` in this
|
| 431 |
+
repository has the regression metrics with the map and at the global temperature, our held-out sets by answer type, the
|
| 432 |
+
fixtures, games, browser and Bespoke results with the paired comparisons against v1 and v2, the text games, the behaviour
|
| 433 |
+
probes, the issue #9 cases and the JevBench public items.
|
| 434 |
|
| 435 |
**Independence.** This is an independent project. It is not affiliated with or endorsed by TypeSafe AI. It is an open
|
| 436 |
reproduction of the "System One" model class (TypeSafe AI's Jev); nothing was distilled from Jev. The training data is public
|
decider/calibrate.py
ADDED
|
@@ -0,0 +1,222 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Fit decider_config.json "temperature_by_type" by NLL, one temperature per answer type (decider.temperature).
|
| 2 |
+
|
| 3 |
+
python -m decider.calibrate records.jsonl [--min-rows 50]
|
| 4 |
+
-> {"temperature_by_type": {"choice": 1.48, "noul": 2.22, "score": 1.38}, "temperature": 1.6, "rows": {...}, "nll": {...}}
|
| 5 |
+
|
| 6 |
+
A record is one answer with its gold, read at temperature 1:
|
| 7 |
+
{"type": "choice" | "noul" | "score", "logits": [one value per option], "gold": option index}
|
| 8 |
+
{"type": "score", "level_logits": [[no, yes], one pair per level], "gold": level index} Score with isolated levels
|
| 9 |
+
"probs" / "level_probs" (probabilities at temperature 1) may stand in for "logits" / "level_logits": log p is the logit up to a
|
| 10 |
+
constant, which the softmax ignores. A noul gold is 1 for true, 0 for false. Every record is checked first (check_record):
|
| 11 |
+
a malformed one (gold out of range or not an integer, wrong shape, NaN, a row without a finite logit, an isolated Score whose
|
| 12 |
+
levels all have P(yes) = 0) raises ValueError naming its position.
|
| 13 |
+
|
| 14 |
+
An isolated-levels Score answer is fitted through the readout the server uses: every level row is softmax(row / T), and the
|
| 15 |
+
answer is P(yes) of each level divided by their sum (decider.systemone.combine_isolated). So the "score" temperature is fitted on
|
| 16 |
+
Score answers whichever readout the model serves them with; fit it on records collected with the same isolated_levels setting.
|
| 17 |
+
|
| 18 |
+
`collect(decider, examples)` produces records from a decider.infer.Decider and labelled /v1/systemone-shaped examples.
|
| 19 |
+
"temperature" in the output is one temperature fitted on all records together, for comparison; a type with fewer than
|
| 20 |
+
--min-rows records is left out of the map (it then uses "temperature" of the config).
|
| 21 |
+
"""
|
| 22 |
+
import json, math, sys
|
| 23 |
+
import numpy as np
|
| 24 |
+
|
| 25 |
+
from decider.temperature import TYPES
|
| 26 |
+
|
| 27 |
+
GRID = np.exp(np.linspace(math.log(0.05), math.log(20.0), 801)) # 0.05 .. 20, about 0.75 % apart
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
def _log(x):
|
| 31 |
+
x = np.asarray(x, dtype=np.float64)
|
| 32 |
+
with np.errstate(divide="ignore"):
|
| 33 |
+
return np.where(x > 0, np.log(np.clip(x, 1e-300, None)), -np.inf)
|
| 34 |
+
|
| 35 |
+
|
| 36 |
+
def _logits(rec, key):
|
| 37 |
+
if key in rec:
|
| 38 |
+
return np.asarray(rec[key], dtype=np.float64)
|
| 39 |
+
alt = {"logits": "probs", "level_logits": "level_probs"}[key]
|
| 40 |
+
return _log(rec[alt])
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
def _lse(z, axis=-1):
|
| 44 |
+
m = np.max(z, axis=axis, keepdims=True)
|
| 45 |
+
m = np.where(np.isfinite(m), m, 0.0)
|
| 46 |
+
with np.errstate(divide="ignore"): # an all -inf row gives -inf, handled by the callers
|
| 47 |
+
return (m + np.log(np.sum(np.exp(z - m), axis=axis, keepdims=True))).squeeze(axis)
|
| 48 |
+
|
| 49 |
+
|
| 50 |
+
def nll(rec, T):
|
| 51 |
+
"""Negative log-likelihood of one record's gold at temperature T."""
|
| 52 |
+
return float(_Batch([rec]).nll(T))
|
| 53 |
+
|
| 54 |
+
|
| 55 |
+
def _isolated(rec):
|
| 56 |
+
return "level_logits" in rec or "level_probs" in rec
|
| 57 |
+
|
| 58 |
+
|
| 59 |
+
def check_record(rec, i=None):
|
| 60 |
+
"""Raise ValueError unless `rec` is a well-formed record (module docstring): an integral gold within the options or levels,
|
| 61 |
+
a 1-D row of at least two options or an [levels, 2] array of at least two levels, logits that are not NaN or +inf,
|
| 62 |
+
probabilities in [0, 1] with positive mass, and isolated levels only on a Score record."""
|
| 63 |
+
where = f"record {i}" if i is not None else "record"
|
| 64 |
+
if not isinstance(rec, dict):
|
| 65 |
+
raise ValueError(f"{where}: not a JSON object")
|
| 66 |
+
if rec.get("type") not in TYPES:
|
| 67 |
+
raise ValueError(f"{where}: type {rec.get('type')!r}, expected one of {', '.join(TYPES)}")
|
| 68 |
+
iso = _isolated(rec)
|
| 69 |
+
keys = ("level_logits", "level_probs") if iso else ("logits", "probs")
|
| 70 |
+
if sum(k in rec for k in keys) != 1 or (iso and ("logits" in rec or "probs" in rec)):
|
| 71 |
+
raise ValueError(f"{where}: give exactly one of logits, probs, level_logits, level_probs")
|
| 72 |
+
if iso and rec["type"] != "score":
|
| 73 |
+
raise ValueError(f"{where}: isolated levels (level_logits / level_probs) belong to a score record, not {rec['type']!r}")
|
| 74 |
+
key = keys[0] if keys[0] in rec else keys[1]
|
| 75 |
+
try:
|
| 76 |
+
a = np.asarray(rec[key], dtype=np.float64)
|
| 77 |
+
except (TypeError, ValueError):
|
| 78 |
+
raise ValueError(f"{where}: {key} is not a numeric array") from None
|
| 79 |
+
if iso and (a.ndim != 2 or a.shape[1] != 2 or a.shape[0] < 2):
|
| 80 |
+
raise ValueError(f"{where}: {key} must be one [no, yes] pair per level, at least two levels; got shape {a.shape}")
|
| 81 |
+
if not iso and (a.ndim != 1 or a.shape[0] < 2):
|
| 82 |
+
raise ValueError(f"{where}: {key} must be one value per option, at least two options; got shape {a.shape}")
|
| 83 |
+
if key.endswith("probs"):
|
| 84 |
+
if not np.all(np.isfinite(a)) or np.any(a < 0) or np.any(a > 1) or np.any(a.sum(-1) <= 0):
|
| 85 |
+
raise ValueError(f"{where}: {key} must be probabilities in [0, 1] with positive mass per row")
|
| 86 |
+
elif np.any(np.isnan(a)) or np.any(a == np.inf) or not np.all(np.any(np.isfinite(a), axis=-1)):
|
| 87 |
+
raise ValueError(f"{where}: {key} contains NaN or +inf, or a row without a finite value")
|
| 88 |
+
if rec["type"] == "noul" and a.shape[0] != 2:
|
| 89 |
+
raise ValueError(f"{where}: a noul record has exactly two options (false, true); got {a.shape[0]}")
|
| 90 |
+
if iso and not np.any(a[:, 1] > 0 if key == "level_probs" else np.isfinite(a[:, 1])):
|
| 91 |
+
raise ValueError(f"{where}: no level has a yes probability above 0, so the served Score answer has no distribution")
|
| 92 |
+
g = rec.get("gold")
|
| 93 |
+
if isinstance(g, bool) or not isinstance(g, (int, np.integer)) or not 0 <= g < a.shape[0]:
|
| 94 |
+
raise ValueError(f"{where}: gold must be an integer index in 0..{a.shape[0] - 1}, got {g!r}")
|
| 95 |
+
return rec
|
| 96 |
+
|
| 97 |
+
|
| 98 |
+
class _Batch:
|
| 99 |
+
"""Records packed into padded arrays, so one temperature is scored over all of them at once."""
|
| 100 |
+
def __init__(self, records):
|
| 101 |
+
records = [check_record(r, i) for i, r in enumerate(records)]
|
| 102 |
+
lists = [r for r in records if not _isolated(r)]
|
| 103 |
+
isos = [r for r in records if _isolated(r)]
|
| 104 |
+
self.n = len(lists) + len(isos)
|
| 105 |
+
self.L = self.Lg = self.I = self.Im = self.Ig = None
|
| 106 |
+
if lists:
|
| 107 |
+
zs = [_logits(r, "logits") for r in lists]
|
| 108 |
+
K = max(len(z) for z in zs)
|
| 109 |
+
self.L = np.full((len(zs), K), -np.inf)
|
| 110 |
+
for i, z in enumerate(zs):
|
| 111 |
+
self.L[i, :len(z)] = z
|
| 112 |
+
self.Lg = np.array([int(r["gold"]) for r in lists])
|
| 113 |
+
if isos:
|
| 114 |
+
zs = [_logits(r, "level_logits").reshape(-1, 2) for r in isos]
|
| 115 |
+
n = max(len(z) for z in zs)
|
| 116 |
+
self.I = np.zeros((len(zs), n, 2)); self.Im = np.zeros((len(zs), n), dtype=bool)
|
| 117 |
+
for i, z in enumerate(zs):
|
| 118 |
+
self.I[i, :len(z)] = z; self.Im[i, :len(z)] = True
|
| 119 |
+
self.Ig = np.array([int(r["gold"]) for r in isos])
|
| 120 |
+
|
| 121 |
+
def nll(self, T):
|
| 122 |
+
"""Sum of the records' NLL at temperature T."""
|
| 123 |
+
tot = 0.0
|
| 124 |
+
if self.L is not None:
|
| 125 |
+
z = self.L / T
|
| 126 |
+
zg = z[np.arange(len(z)), self.Lg]
|
| 127 |
+
tot += float(np.sum(np.where(np.isfinite(zg), _lse(z) - zg, 690.0)))
|
| 128 |
+
if self.I is not None: # in log space: P(yes) of very confident rows underflows otherwise
|
| 129 |
+
z = self.I / T
|
| 130 |
+
with np.errstate(invalid="ignore"):
|
| 131 |
+
lp = np.where(self.Im, z[..., 1] - _lse(z), -np.inf) # log P(yes) per level, -inf on padding
|
| 132 |
+
norm = _lse(lp) # log of the summed P(yes)
|
| 133 |
+
lg = lp[np.arange(len(z)), self.Ig]
|
| 134 |
+
with np.errstate(invalid="ignore"):
|
| 135 |
+
per = np.where(np.isfinite(lg), norm - lg, 690.0) # check_record guarantees a finite norm
|
| 136 |
+
tot += float(np.sum(per))
|
| 137 |
+
return tot
|
| 138 |
+
|
| 139 |
+
|
| 140 |
+
def mean_nll(records, T):
|
| 141 |
+
b = records if isinstance(records, _Batch) else _Batch(records)
|
| 142 |
+
return b.nll(T) / b.n
|
| 143 |
+
|
| 144 |
+
|
| 145 |
+
def fit(records, grid=GRID):
|
| 146 |
+
"""The temperature with the lowest mean NLL on the grid, refined by a golden-section search between its neighbours."""
|
| 147 |
+
records = records if isinstance(records, _Batch) else _Batch(list(records))
|
| 148 |
+
if not records.n:
|
| 149 |
+
raise ValueError("no records to fit")
|
| 150 |
+
vals = [mean_nll(records, T) for T in grid]
|
| 151 |
+
i = int(np.argmin(vals))
|
| 152 |
+
lo, hi = math.log(grid[max(i - 1, 0)]), math.log(grid[min(i + 1, len(grid) - 1)])
|
| 153 |
+
r = (math.sqrt(5) - 1) / 2
|
| 154 |
+
a, b = hi - r * (hi - lo), lo + r * (hi - lo)
|
| 155 |
+
fa, fb = mean_nll(records, math.exp(a)), mean_nll(records, math.exp(b))
|
| 156 |
+
for _ in range(40):
|
| 157 |
+
if fa < fb:
|
| 158 |
+
hi, b, fb = b, a, fa; a = hi - r * (hi - lo); fa = mean_nll(records, math.exp(a))
|
| 159 |
+
else:
|
| 160 |
+
lo, a, fa = a, b, fb; b = lo + r * (hi - lo); fb = mean_nll(records, math.exp(b))
|
| 161 |
+
T = math.exp((lo + hi) / 2)
|
| 162 |
+
return T if mean_nll(records, T) <= vals[i] else float(grid[i])
|
| 163 |
+
|
| 164 |
+
|
| 165 |
+
def fit_by_type(records, min_rows=50):
|
| 166 |
+
"""-> {"temperature_by_type": {type: T}, "temperature": pooled T, "rows": {type: n}, "nll": {type: {"T=1": .., "fitted": ..}}}"""
|
| 167 |
+
records = [check_record(r, i) for i, r in enumerate(records)]
|
| 168 |
+
groups = {t: [] for t in TYPES}
|
| 169 |
+
for r in records:
|
| 170 |
+
groups[r["type"]].append(r)
|
| 171 |
+
out = {"temperature_by_type": {}, "temperature": round(fit(records), 3) if records else None,
|
| 172 |
+
"rows": {t: len(g) for t, g in groups.items()}, "nll": {}}
|
| 173 |
+
for t, g in groups.items():
|
| 174 |
+
if len(g) < max(1, min_rows):
|
| 175 |
+
continue
|
| 176 |
+
b = _Batch(g); T = fit(b)
|
| 177 |
+
out["temperature_by_type"][t] = round(T, 3)
|
| 178 |
+
out["nll"][t] = {"T=1": round(mean_nll(b, 1.0), 4), "fitted": round(mean_nll(b, T), 4)}
|
| 179 |
+
return out
|
| 180 |
+
|
| 181 |
+
|
| 182 |
+
def collect(decider, examples, independent=True):
|
| 183 |
+
"""Records for fit_by_type from a decider.infer.Decider, read at temperature 1 on the uncached state-first path.
|
| 184 |
+
examples: iterable of (state, questions, golds) with questions as /v1/systemone takes them and golds {question id: gold},
|
| 185 |
+
the gold being a Choice option name, a Noul true/false, or a Score level index. Questions without a gold are skipped.
|
| 186 |
+
Score questions are read with the Decider's isolated_levels setting, as system_one serves them."""
|
| 187 |
+
recs = []
|
| 188 |
+
for state, questions, golds in examples:
|
| 189 |
+
rqs, index, items = decider._system_one_items(state, questions, independent, layout="state_first")
|
| 190 |
+
rows = decider._system_one_probs(items, "state_first", temperature=1.0)
|
| 191 |
+
for k, kind, s, n in index:
|
| 192 |
+
if k not in golds:
|
| 193 |
+
continue
|
| 194 |
+
rq, g = rqs[k], golds[k]
|
| 195 |
+
if rq["type"] == "noul" and not (isinstance(g, bool) or g in (0, 1)):
|
| 196 |
+
raise ValueError(f"gold of noul question {k!r} must be true or false, got {g!r}")
|
| 197 |
+
if rq["type"] == "score" and (isinstance(g, bool) or not isinstance(g, int) or not 0 <= g < len(rq["names"])):
|
| 198 |
+
raise ValueError(f"gold of score question {k!r} must be a level index in 0..{len(rq['names']) - 1}, got {g!r}")
|
| 199 |
+
key = bool(g) if rq["type"] == "noul" else g
|
| 200 |
+
if key not in rq["names"]:
|
| 201 |
+
raise ValueError(f"gold of choice question {k!r} is not one of its options: {g!r}")
|
| 202 |
+
gold = rq["names"].index(key)
|
| 203 |
+
if kind == "iso":
|
| 204 |
+
recs.append({"type": rq["type"], "level_probs": [rows[s + j][:2] for j in range(n)], "gold": gold})
|
| 205 |
+
else:
|
| 206 |
+
recs.append({"type": rq["type"], "probs": rows[s][:len(rq["options"])], "gold": gold})
|
| 207 |
+
return recs
|
| 208 |
+
|
| 209 |
+
|
| 210 |
+
def main(argv=None):
|
| 211 |
+
import argparse
|
| 212 |
+
ap = argparse.ArgumentParser(description=__doc__.split("\n\n")[0])
|
| 213 |
+
ap.add_argument("records", help="JSON lines, one record per line (see the module docstring)")
|
| 214 |
+
ap.add_argument("--min-rows", type=int, default=50, help="fit a type only when it has at least this many records")
|
| 215 |
+
a = ap.parse_args(argv)
|
| 216 |
+
with open(a.records) as f:
|
| 217 |
+
recs = [json.loads(line) for line in f if line.strip()]
|
| 218 |
+
print(json.dumps(fit_by_type(recs, a.min_rows), indent=1))
|
| 219 |
+
|
| 220 |
+
|
| 221 |
+
if __name__ == "__main__":
|
| 222 |
+
main(sys.argv[1:])
|
decider/engine.py
CHANGED
|
@@ -7,6 +7,7 @@ option-letter logits for all positions [B, T, K]; slots are gathered outside.
|
|
| 7 |
import time, torch, torch._dynamo, torch.nn.functional as F
|
| 8 |
from decider.model import DecisionModel, collate
|
| 9 |
from decider.prompt import build, MAX_OPTIONS
|
|
|
|
| 10 |
|
| 11 |
T_BUCKETS = [64, 128, 192, 256, 320, 384, 512, 640, 768, 1024, 1280, 1536, 2048]
|
| 12 |
B_BUCKETS = [1, 2, 4, 8, 16, 32, 64]
|
|
@@ -46,11 +47,12 @@ def patch_conv():
|
|
| 46 |
|
| 47 |
def read_slots(out, rows, slots, nopts, temperature, n_per_item):
|
| 48 |
"""One gather + one softmax + one device-to-host copy for the whole batch (was: three small kernels and a sync per item).
|
| 49 |
-
out [B, T, K] logits; rows/slots/nopts: flat python lists, one entry per question; n_per_item: questions per item.
|
|
|
|
| 50 |
dev = out.device; idx = torch.tensor([rows, slots, nopts], dtype=torch.long).to(dev, non_blocking=True)
|
| 51 |
lg = out[idx[0], idx[1]] # [N, K]
|
| 52 |
lg = lg.masked_fill(torch.arange(lg.shape[1], device=dev)[None, :] >= idx[2][:, None], float("-inf"))
|
| 53 |
-
p =
|
| 54 |
return list(torch.split(p, n_per_item))
|
| 55 |
|
| 56 |
|
|
@@ -140,14 +142,15 @@ class Engine:
|
|
| 140 |
|
| 141 |
@torch.no_grad()
|
| 142 |
def score_items(self, items, temperature=1.0):
|
| 143 |
-
"""items: list of dicts from prompt.build. Returns list of [n_q, MAX_OPTIONS] prob tensors (cpu).
|
|
|
|
| 144 |
Tmax = max(len(it["ids"]) for it in items)
|
| 145 |
T = _bucket(Tmax, T_BUCKETS) or -(-Tmax // LONG_STEP) * LONG_STEP
|
| 146 |
B = (_bucket(len(items), B_BUCKETS) or len(items)) if T <= GRAPH_MAX_T else len(items)
|
| 147 |
ids = fill_ids([it["ids"] for it in items], B, T, self.tok.pad_token_id)
|
| 148 |
out = self.logits_all(ids.to(self.dev, non_blocking=True))
|
| 149 |
return read_slots(out, [b for b, it in enumerate(items) for _ in it["slots"]], [s for it in items for s in it["slots"]],
|
| 150 |
-
[n for it in items for n in it["nopts"]], temperature, [len(it["slots"]) for it in items])
|
| 151 |
|
| 152 |
@torch.no_grad()
|
| 153 |
def score_shared(self, items, temperature=1.0, min_prefix=192):
|
|
|
|
| 7 |
import time, torch, torch._dynamo, torch.nn.functional as F
|
| 8 |
from decider.model import DecisionModel, collate
|
| 9 |
from decider.prompt import build, MAX_OPTIONS
|
| 10 |
+
from decider.temperature import scaled_softmax, slot_temperatures
|
| 11 |
|
| 12 |
T_BUCKETS = [64, 128, 192, 256, 320, 384, 512, 640, 768, 1024, 1280, 1536, 2048]
|
| 13 |
B_BUCKETS = [1, 2, 4, 8, 16, 32, 64]
|
|
|
|
| 47 |
|
| 48 |
def read_slots(out, rows, slots, nopts, temperature, n_per_item):
|
| 49 |
"""One gather + one softmax + one device-to-host copy for the whole batch (was: three small kernels and a sync per item).
|
| 50 |
+
out [B, T, K] logits; rows/slots/nopts: flat python lists, one entry per question; n_per_item: questions per item.
|
| 51 |
+
temperature: a number for every question, or a flat list with one temperature per question (decider.temperature)."""
|
| 52 |
dev = out.device; idx = torch.tensor([rows, slots, nopts], dtype=torch.long).to(dev, non_blocking=True)
|
| 53 |
lg = out[idx[0], idx[1]] # [N, K]
|
| 54 |
lg = lg.masked_fill(torch.arange(lg.shape[1], device=dev)[None, :] >= idx[2][:, None], float("-inf"))
|
| 55 |
+
p = scaled_softmax(lg, temperature).cpu()
|
| 56 |
return list(torch.split(p, n_per_item))
|
| 57 |
|
| 58 |
|
|
|
|
| 142 |
|
| 143 |
@torch.no_grad()
|
| 144 |
def score_items(self, items, temperature=1.0):
|
| 145 |
+
"""items: list of dicts from prompt.build. Returns list of [n_q, MAX_OPTIONS] prob tensors (cpu).
|
| 146 |
+
temperature: a number, or one entry per item (a number or one number per slot; decider.temperature.for_items)."""
|
| 147 |
Tmax = max(len(it["ids"]) for it in items)
|
| 148 |
T = _bucket(Tmax, T_BUCKETS) or -(-Tmax // LONG_STEP) * LONG_STEP
|
| 149 |
B = (_bucket(len(items), B_BUCKETS) or len(items)) if T <= GRAPH_MAX_T else len(items)
|
| 150 |
ids = fill_ids([it["ids"] for it in items], B, T, self.tok.pad_token_id)
|
| 151 |
out = self.logits_all(ids.to(self.dev, non_blocking=True))
|
| 152 |
return read_slots(out, [b for b, it in enumerate(items) for _ in it["slots"]], [s for it in items for s in it["slots"]],
|
| 153 |
+
[n for it in items for n in it["nopts"]], slot_temperatures(temperature, items), [len(it["slots"]) for it in items])
|
| 154 |
|
| 155 |
@torch.no_grad()
|
| 156 |
def score_shared(self, items, temperature=1.0, min_prefix=192):
|
decider/engine_v2.py
CHANGED
|
@@ -19,6 +19,7 @@ import time, torch, torch.nn.functional as F
|
|
| 19 |
from decider import shared_prefix
|
| 20 |
from decider.engine import read_slots, fill_ids, patch_conv, set_attention_backend_policy
|
| 21 |
from decider.model import DecisionModel
|
|
|
|
| 22 |
|
| 23 |
T_BUCKETS = [64, 128, 192, 256, 320, 384, 512, 640, 768, 1024, 1280, 1536, 2048, 3072, 4096, 6144, 8192]
|
| 24 |
B_BUCKETS = [1, 2, 4, 8, 16, 32]
|
|
@@ -150,9 +151,11 @@ class EngineV2:
|
|
| 150 |
# ---- scoring -----------------------------------------------------------
|
| 151 |
@torch.no_grad()
|
| 152 |
def score_items(self, items, temperature=1.0):
|
| 153 |
-
"""items: dicts from prompt.build / build_rows. -> one [n_q, MAX_OPTIONS] cpu probability tensor per item.
|
|
|
|
| 154 |
if not items:
|
| 155 |
return []
|
|
|
|
| 156 |
Tmax = max(len(it["ids"]) for it in items)
|
| 157 |
T = self.t_bucket(Tmax)
|
| 158 |
if T is None:
|
|
@@ -168,7 +171,7 @@ class EngineV2:
|
|
| 168 |
lg = self.logits_all(ids.to(self.dev, non_blocking=True))
|
| 169 |
out += read_slots(lg, [b for b, it in enumerate(chunk) for _ in it["slots"]],
|
| 170 |
[s for it in chunk for s in it["slots"]], [n for it in chunk for n in it["nopts"]],
|
| 171 |
-
temperature, [len(it["slots"]) for it in chunk])
|
| 172 |
return out
|
| 173 |
|
| 174 |
@torch.no_grad()
|
|
|
|
| 19 |
from decider import shared_prefix
|
| 20 |
from decider.engine import read_slots, fill_ids, patch_conv, set_attention_backend_policy
|
| 21 |
from decider.model import DecisionModel
|
| 22 |
+
from decider.temperature import item_slice, slot_temperatures
|
| 23 |
|
| 24 |
T_BUCKETS = [64, 128, 192, 256, 320, 384, 512, 640, 768, 1024, 1280, 1536, 2048, 3072, 4096, 6144, 8192]
|
| 25 |
B_BUCKETS = [1, 2, 4, 8, 16, 32]
|
|
|
|
| 151 |
# ---- scoring -----------------------------------------------------------
|
| 152 |
@torch.no_grad()
|
| 153 |
def score_items(self, items, temperature=1.0):
|
| 154 |
+
"""items: dicts from prompt.build / build_rows. -> one [n_q, MAX_OPTIONS] cpu probability tensor per item.
|
| 155 |
+
temperature: a number, or one entry per item (a number or one number per slot; decider.temperature.for_items)."""
|
| 156 |
if not items:
|
| 157 |
return []
|
| 158 |
+
slot_temperatures(temperature, items) # a length mismatch fails before any forward
|
| 159 |
Tmax = max(len(it["ids"]) for it in items)
|
| 160 |
T = self.t_bucket(Tmax)
|
| 161 |
if T is None:
|
|
|
|
| 171 |
lg = self.logits_all(ids.to(self.dev, non_blocking=True))
|
| 172 |
out += read_slots(lg, [b for b, it in enumerate(chunk) for _ in it["slots"]],
|
| 173 |
[s for it in chunk for s in it["slots"]], [n for it in chunk for n in it["nopts"]],
|
| 174 |
+
slot_temperatures(item_slice(temperature, i - len(chunk), i), chunk), [len(it["slots"]) for it in chunk])
|
| 175 |
return out
|
| 176 |
|
| 177 |
@torch.no_grad()
|
decider/infer.py
CHANGED
|
@@ -12,6 +12,7 @@ import logging
|
|
| 12 |
import torch
|
| 13 |
from decider.model import DecisionModel, collate
|
| 14 |
from decider.prompt import build, MAX_OPTIONS, resolve_layout, chat_template
|
|
|
|
| 15 |
from dataclasses import dataclass
|
| 16 |
|
| 17 |
|
|
@@ -49,11 +50,15 @@ class _NoShuffle: # keep option order as given
|
|
| 49 |
|
| 50 |
|
| 51 |
class CompiledSchema:
|
| 52 |
-
def __init__(self, d, rqs, h, index):
|
|
|
|
|
|
|
|
|
|
| 53 |
|
| 54 |
def batch(self, states, max_state_tokens=32768):
|
| 55 |
from decider.systemone import render_state, assemble
|
| 56 |
-
|
|
|
|
| 57 |
return [{"model": self.d.name, "answers": assemble(self.rqs, self.index, [p.tolist() for p in pr])} for pr in probs]
|
| 58 |
|
| 59 |
def __call__(self, state, max_state_tokens=32768):
|
|
@@ -67,10 +72,15 @@ class Decider:
|
|
| 67 |
the optional MPS patch; CPU defaults to bfloat16. Set ``use_graphs=False``
|
| 68 |
for eager execution or debugging.
|
| 69 |
"""
|
| 70 |
-
def __init__(self, path, device=None, dtype=None, temperature=None, abstain_below=0.0, use_graphs=None):
|
| 71 |
"""The prompt layout comes from decider_config.json: "layout": "chat" (chat-trained checkpoints) wraps every prompt in the
|
| 72 |
tokenizer's chat template (decider.prompt.build_chat); no "layout" key is the plain layout of every earlier model.
|
| 73 |
-
An unknown layout raises ValueError before the weights are loaded.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 74 |
if device is None:
|
| 75 |
device = "cuda" if torch.cuda.is_available() else ("mps" if torch.backends.mps.is_available() else "cpu")
|
| 76 |
if dtype is None:
|
|
@@ -85,8 +95,7 @@ class Decider:
|
|
| 85 |
except Exception:
|
| 86 |
pass
|
| 87 |
self.layout = resolve_layout(cfg)
|
| 88 |
-
|
| 89 |
-
temperature = float(cfg.get("temperature", 1.0))
|
| 90 |
self.neutralize_none = bool(cfg.get("neutralize_none", True)) # v4 and earlier learned the literal string as an abstain signal
|
| 91 |
if use_graphs is None:
|
| 92 |
use_graphs = str(device).startswith("cuda")
|
|
@@ -102,7 +111,7 @@ class Decider:
|
|
| 102 |
self.dev = device; self.T = temperature; self.abstain_below = abstain_below
|
| 103 |
self.name = "decider-" + str(cfg.get("version", "dev"))
|
| 104 |
self.schema_first = bool(cfg.get("schema_first", False)) and self.eng is not None # default layout. Questions-first (the cacheable one) costs accuracy
|
| 105 |
-
self.T_schema =
|
| 106 |
self.isolated_levels = bool(cfg.get("isolated_levels", False)) # Score levels judged one per row (v8+)
|
| 107 |
self._se = None; self._schemas = {}
|
| 108 |
|
|
@@ -116,19 +125,22 @@ class Decider:
|
|
| 116 |
assert 2 <= len(q["options"]) <= MAX_OPTIONS, f"2..{MAX_OPTIONS} options required"
|
| 117 |
exs.append(Example(context, [Q(q["question"], list(q["options"]), 0) for q in qs], "infer"))
|
| 118 |
items = [build(e, self.m.tok, _NoShuffle(), max_options=MAX_OPTIONS, max_ctx_tokens=max_ctx_tokens, chat=self.chat) for e in exs]
|
|
|
|
|
|
|
| 119 |
return requests, items
|
| 120 |
|
| 121 |
@torch.no_grad()
|
| 122 |
def decide_batch(self, requests, max_ctx_tokens=1536):
|
| 123 |
"""requests: list of (context:str, questions:list[dict(question, options)]). One forward pass for everything."""
|
| 124 |
requests, items = self._decide_items(requests, max_ctx_tokens)
|
|
|
|
| 125 |
if self.eng is not None:
|
| 126 |
-
probs = torch.cat(self.eng.score_items(items, temperature=
|
| 127 |
else:
|
| 128 |
b = collate(items, self.m.tok.pad_token_id)
|
| 129 |
logits = self.m.slot_logits(b["input_ids"].to(self.dev), b["attention_mask"].to(self.dev), b["slot_idx"].to(self.dev),
|
| 130 |
b["slot_batch"].to(self.dev), b["nopts"].to(self.dev))
|
| 131 |
-
probs =
|
| 132 |
out, k = [], 0
|
| 133 |
for context, qs in requests:
|
| 134 |
res = []
|
|
@@ -176,27 +188,33 @@ class Decider:
|
|
| 176 |
(later questions can then see earlier question texts)."""
|
| 177 |
from decider.systemone import unique_tokens, assemble
|
| 178 |
rqs, index, items = self._system_one_items(state, questions, independent, max_state_tokens, layout, isolated)
|
| 179 |
-
|
| 180 |
-
if self.eng is not None and len(items) > 1 and layout == "state_first":
|
| 181 |
-
probs = self.eng.score_shared(items, temperature=self.T)
|
| 182 |
-
else:
|
| 183 |
-
probs = []; per = max(1, max_fwd_tokens // max(len(it["ids"]) for it in items))
|
| 184 |
-
for i in range(0, len(items), per):
|
| 185 |
-
if self.eng is not None:
|
| 186 |
-
probs += self.eng.score_items(items[i:i + per], temperature=self.T)
|
| 187 |
-
else:
|
| 188 |
-
bt = collate(items[i:i + per], self.m.tok.pad_token_id)
|
| 189 |
-
lg = self.m.slot_logits(*[bt[k].to(self.dev) for k in ("input_ids", "attention_mask", "slot_idx", "slot_batch", "nopts")])
|
| 190 |
-
pr = torch.softmax(lg / self.T, -1).cpu(); c = 0
|
| 191 |
-
for it in items[i:i + per]:
|
| 192 |
-
probs.append(pr[c:c + len(it["slots"])]); c += len(it["slots"])
|
| 193 |
-
flatp = [p.tolist() for ps in probs for p in ps]
|
| 194 |
return {"model": self.name, "answers": assemble(rqs, index, flatp),
|
| 195 |
"usage": {"input_tokens": unique_tokens(items), "output_tokens": 0}}
|
| 196 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 197 |
def _system_one_items(self, state, questions, independent=True, max_state_tokens=32768, layout=None, isolated=None):
|
| 198 |
"""The prompt rows of system_one's uncached path: -> (rendered questions, answer index, items)."""
|
| 199 |
-
from decider.systemone import render_state, render_question, plan_rows
|
| 200 |
layout = layout or ("schema_first" if self.schema_first else "state_first")
|
| 201 |
isolated = (self.isolated_levels if isolated is None else isolated) and independent
|
| 202 |
ctx = render_state(state); rqs = {k: render_question(v) for k, v in questions.items()}
|
|
@@ -205,6 +223,9 @@ class Decider:
|
|
| 205 |
rows = [[r] for r in flat] if independent else [flat]
|
| 206 |
items = [build(Example(ctx, [Q(r["question"], opts(r), 0) for r in row]), self.m.tok, _NoShuffle(), max_options=MAX_OPTIONS,
|
| 207 |
max_ctx_tokens=max_state_tokens, layout=layout, chat=self.chat) for row in rows]
|
|
|
|
|
|
|
|
|
|
| 208 |
return rqs, index, items
|
| 209 |
|
| 210 |
# ---- typed schema interface: {question: {"type": "bool"} | {"type": "choice", "options": [...]}
|
|
@@ -257,14 +278,14 @@ class Decider:
|
|
| 257 |
for qtext, spec in schema.items():
|
| 258 |
t = spec.get("type", "choice")
|
| 259 |
if t == "bool":
|
| 260 |
-
qs.append(dict(question=qtext, options=["no", "yes"]))
|
| 261 |
elif t == "choice":
|
| 262 |
-
qs.append(dict(question=qtext, options=list(spec["options"])))
|
| 263 |
elif t == "scale":
|
| 264 |
leg = spec["legend"]
|
| 265 |
keys = sorted(leg, key=lambda k: float(k)) if isinstance(leg, dict) else list(range(len(leg)))
|
| 266 |
labels = [f"{k}: {leg[k]}" if isinstance(leg, dict) else f"{i}: {leg[i]}" for i, k in enumerate(keys)]
|
| 267 |
-
qs.append(dict(question=qtext, options=labels, _keys=keys, _legend=leg))
|
| 268 |
else:
|
| 269 |
raise ValueError(f"unknown field type {t}")
|
| 270 |
return qs
|
|
|
|
| 12 |
import torch
|
| 13 |
from decider.model import DecisionModel, collate
|
| 14 |
from decider.prompt import build, MAX_OPTIONS, resolve_layout, chat_template
|
| 15 |
+
from decider import temperature as TT
|
| 16 |
from dataclasses import dataclass
|
| 17 |
|
| 18 |
|
|
|
|
| 50 |
|
| 51 |
|
| 52 |
class CompiledSchema:
|
| 53 |
+
def __init__(self, d, rqs, h, index):
|
| 54 |
+
from decider.systemone import row_types
|
| 55 |
+
self.d, self.rqs, self.h, self.index = d, rqs, h, index
|
| 56 |
+
self.types = row_types(rqs, index) # answer type of every schema row (decider.temperature)
|
| 57 |
|
| 58 |
def batch(self, states, max_state_tokens=32768):
|
| 59 |
from decider.systemone import render_state, assemble
|
| 60 |
+
T = TT.for_types(self.d.T_schema, self.d.T_schema_by_type, self.types)
|
| 61 |
+
probs = self.d._se.score(self.h, [render_state(s) for s in states], temperature=T, max_ctx_tokens=max_state_tokens)
|
| 62 |
return [{"model": self.d.name, "answers": assemble(self.rqs, self.index, [p.tolist() for p in pr])} for pr in probs]
|
| 63 |
|
| 64 |
def __call__(self, state, max_state_tokens=32768):
|
|
|
|
| 72 |
the optional MPS patch; CPU defaults to bfloat16. Set ``use_graphs=False``
|
| 73 |
for eager execution or debugging.
|
| 74 |
"""
|
| 75 |
+
def __init__(self, path, device=None, dtype=None, temperature=None, abstain_below=0.0, use_graphs=None, temperature_by_type=None):
|
| 76 |
"""The prompt layout comes from decider_config.json: "layout": "chat" (chat-trained checkpoints) wraps every prompt in the
|
| 77 |
tokenizer's chat template (decider.prompt.build_chat); no "layout" key is the plain layout of every earlier model.
|
| 78 |
+
An unknown layout raises ValueError before the weights are loaded.
|
| 79 |
+
|
| 80 |
+
Temperatures (decider.temperature): decider_config.json "temperature" and the optional "temperature_by_type"
|
| 81 |
+
{"choice": T, "noul": T, "score": T}. temperature= overrides "temperature" and switches the config's by-type map off;
|
| 82 |
+
temperature_by_type= sets the map explicitly. An invalid temperature or map key raises ValueError before the weights
|
| 83 |
+
are loaded."""
|
| 84 |
if device is None:
|
| 85 |
device = "cuda" if torch.cuda.is_available() else ("mps" if torch.backends.mps.is_available() else "cpu")
|
| 86 |
if dtype is None:
|
|
|
|
| 95 |
except Exception:
|
| 96 |
pass
|
| 97 |
self.layout = resolve_layout(cfg)
|
| 98 |
+
(temperature, self.T_by_type), (T_schema, self.T_schema_by_type) = TT.from_config(cfg, temperature, temperature_by_type)
|
|
|
|
| 99 |
self.neutralize_none = bool(cfg.get("neutralize_none", True)) # v4 and earlier learned the literal string as an abstain signal
|
| 100 |
if use_graphs is None:
|
| 101 |
use_graphs = str(device).startswith("cuda")
|
|
|
|
| 111 |
self.dev = device; self.T = temperature; self.abstain_below = abstain_below
|
| 112 |
self.name = "decider-" + str(cfg.get("version", "dev"))
|
| 113 |
self.schema_first = bool(cfg.get("schema_first", False)) and self.eng is not None # default layout. Questions-first (the cacheable one) costs accuracy
|
| 114 |
+
self.T_schema = T_schema # (about 1.5 points on fixed label sets, more elsewhere): opt in with schema()
|
| 115 |
self.isolated_levels = bool(cfg.get("isolated_levels", False)) # Score levels judged one per row (v8+)
|
| 116 |
self._se = None; self._schemas = {}
|
| 117 |
|
|
|
|
| 125 |
assert 2 <= len(q["options"]) <= MAX_OPTIONS, f"2..{MAX_OPTIONS} options required"
|
| 126 |
exs.append(Example(context, [Q(q["question"], list(q["options"]), 0) for q in qs], "infer"))
|
| 127 |
items = [build(e, self.m.tok, _NoShuffle(), max_options=MAX_OPTIONS, max_ctx_tokens=max_ctx_tokens, chat=self.chat) for e in exs]
|
| 128 |
+
for it, (_, qs) in zip(items, requests): # answer type per slot: a plain question is "choice"
|
| 129 |
+
it["types"] = [q.get("_type", "choice") for q in qs]
|
| 130 |
return requests, items
|
| 131 |
|
| 132 |
@torch.no_grad()
|
| 133 |
def decide_batch(self, requests, max_ctx_tokens=1536):
|
| 134 |
"""requests: list of (context:str, questions:list[dict(question, options)]). One forward pass for everything."""
|
| 135 |
requests, items = self._decide_items(requests, max_ctx_tokens)
|
| 136 |
+
T = TT.for_items(self.T, self.T_by_type, items) # the scalar self.T when there is no by-type map
|
| 137 |
if self.eng is not None:
|
| 138 |
+
probs = torch.cat(self.eng.score_items(items, temperature=T))
|
| 139 |
else:
|
| 140 |
b = collate(items, self.m.tok.pad_token_id)
|
| 141 |
logits = self.m.slot_logits(b["input_ids"].to(self.dev), b["attention_mask"].to(self.dev), b["slot_idx"].to(self.dev),
|
| 142 |
b["slot_batch"].to(self.dev), b["nopts"].to(self.dev))
|
| 143 |
+
probs = TT.scaled_softmax(logits, TT.slot_temperatures(T, items)).cpu()
|
| 144 |
out, k = [], 0
|
| 145 |
for context, qs in requests:
|
| 146 |
res = []
|
|
|
|
| 188 |
(later questions can then see earlier question texts)."""
|
| 189 |
from decider.systemone import unique_tokens, assemble
|
| 190 |
rqs, index, items = self._system_one_items(state, questions, independent, max_state_tokens, layout, isolated)
|
| 191 |
+
flatp = self._system_one_probs(items, layout, max_fwd_tokens, TT.for_items(self.T, self.T_by_type, items))
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 192 |
return {"model": self.name, "answers": assemble(rqs, index, flatp),
|
| 193 |
"usage": {"input_tokens": unique_tokens(items), "output_tokens": 0}}
|
| 194 |
|
| 195 |
+
@torch.no_grad()
|
| 196 |
+
def _system_one_probs(self, items, layout, max_fwd_tokens=65536, temperature=None):
|
| 197 |
+
"""Score system_one's uncached rows. temperature: as Engine.score_items takes it (a number, or one entry per item).
|
| 198 |
+
-> one probability list per question row, in row order."""
|
| 199 |
+
if self.eng is not None and len(items) > 1 and layout == "state_first":
|
| 200 |
+
probs = self.eng.score_shared(items, temperature=temperature)
|
| 201 |
+
else:
|
| 202 |
+
probs = []; per = max(1, max_fwd_tokens // max(len(it["ids"]) for it in items))
|
| 203 |
+
for i in range(0, len(items), per):
|
| 204 |
+
T = TT.item_slice(temperature, i, i + per)
|
| 205 |
+
if self.eng is not None:
|
| 206 |
+
probs += self.eng.score_items(items[i:i + per], temperature=T)
|
| 207 |
+
else:
|
| 208 |
+
bt = collate(items[i:i + per], self.m.tok.pad_token_id)
|
| 209 |
+
lg = self.m.slot_logits(*[bt[k].to(self.dev) for k in ("input_ids", "attention_mask", "slot_idx", "slot_batch", "nopts")])
|
| 210 |
+
pr = TT.scaled_softmax(lg, TT.slot_temperatures(T, items[i:i + per])).cpu(); c = 0
|
| 211 |
+
for it in items[i:i + per]:
|
| 212 |
+
probs.append(pr[c:c + len(it["slots"])]); c += len(it["slots"])
|
| 213 |
+
return [p.tolist() for ps in probs for p in ps]
|
| 214 |
+
|
| 215 |
def _system_one_items(self, state, questions, independent=True, max_state_tokens=32768, layout=None, isolated=None):
|
| 216 |
"""The prompt rows of system_one's uncached path: -> (rendered questions, answer index, items)."""
|
| 217 |
+
from decider.systemone import render_state, render_question, plan_rows, row_types
|
| 218 |
layout = layout or ("schema_first" if self.schema_first else "state_first")
|
| 219 |
isolated = (self.isolated_levels if isolated is None else isolated) and independent
|
| 220 |
ctx = render_state(state); rqs = {k: render_question(v) for k, v in questions.items()}
|
|
|
|
| 223 |
rows = [[r] for r in flat] if independent else [flat]
|
| 224 |
items = [build(Example(ctx, [Q(r["question"], opts(r), 0) for r in row]), self.m.tok, _NoShuffle(), max_options=MAX_OPTIONS,
|
| 225 |
max_ctx_tokens=max_state_tokens, layout=layout, chat=self.chat) for row in rows]
|
| 226 |
+
types = row_types(rqs, index) # answer type per row (decider.temperature)
|
| 227 |
+
for it, ts in zip(items, [[t] for t in types] if independent else [types]):
|
| 228 |
+
it["types"] = ts
|
| 229 |
return rqs, index, items
|
| 230 |
|
| 231 |
# ---- typed schema interface: {question: {"type": "bool"} | {"type": "choice", "options": [...]}
|
|
|
|
| 278 |
for qtext, spec in schema.items():
|
| 279 |
t = spec.get("type", "choice")
|
| 280 |
if t == "bool":
|
| 281 |
+
qs.append(dict(question=qtext, options=["no", "yes"], _type="noul"))
|
| 282 |
elif t == "choice":
|
| 283 |
+
qs.append(dict(question=qtext, options=list(spec["options"]), _type="choice"))
|
| 284 |
elif t == "scale":
|
| 285 |
leg = spec["legend"]
|
| 286 |
keys = sorted(leg, key=lambda k: float(k)) if isinstance(leg, dict) else list(range(len(leg)))
|
| 287 |
labels = [f"{k}: {leg[k]}" if isinstance(leg, dict) else f"{i}: {leg[i]}" for i, k in enumerate(keys)]
|
| 288 |
+
qs.append(dict(question=qtext, options=labels, _keys=keys, _legend=leg, _type="score"))
|
| 289 |
else:
|
| 290 |
raise ValueError(f"unknown field type {t}")
|
| 291 |
return qs
|
decider/schema_engine.py
CHANGED
|
@@ -124,12 +124,15 @@ class SchemaEngine:
|
|
| 124 |
return next((t for t in TS_BUCKETS if t >= n_tokens), -(-n_tokens // 256) * 256)
|
| 125 |
|
| 126 |
def score(self, h, contexts, temperature=1.0, max_ctx_tokens=1536):
|
| 127 |
-
"""-> one [n_questions, MAX_OPTIONS] probability tensor per context."""
|
| 128 |
return self.score_rows(h, [self.tokenize(h, c, max_ctx_tokens) for c in contexts], temperature)
|
| 129 |
|
| 130 |
@torch.no_grad()
|
| 131 |
def score_rows(self, h, rows, temperature=1.0):
|
| 132 |
-
"""rows: [(suffix ids, slots)] from tokenize().
|
|
|
|
|
|
|
|
|
|
| 133 |
Tmax = max(len(r[0]) for r in rows); Ts = next((t for t in TS_BUCKETS if t >= Tmax), None); n = len(rows)
|
| 134 |
R = next((b for b in B_BUCKETS if b >= n), n) if Ts else n; Ts = Ts or -(-Tmax // 256) * 256
|
| 135 |
ids = fill_ids([x for x, _ in rows for _ in range(h.P)], R * h.P, Ts, self.tok.pad_token_id).to(self.dev, non_blocking=True)
|
|
@@ -143,4 +146,5 @@ class SchemaEngine:
|
|
| 143 |
rws = [r for r in range(n) for _ in range(h.nq)]; sls = [x for _, sl in rows for x in sl]
|
| 144 |
else: # independent: one slot in each of the request's P rows
|
| 145 |
rws = [r * h.P + p for r in range(n) for p in range(h.P)]; sls = [sl[0] for _, sl in rows for _ in range(h.P)]
|
| 146 |
-
|
|
|
|
|
|
| 124 |
return next((t for t in TS_BUCKETS if t >= n_tokens), -(-n_tokens // 256) * 256)
|
| 125 |
|
| 126 |
def score(self, h, contexts, temperature=1.0, max_ctx_tokens=1536):
|
| 127 |
+
"""-> one [n_questions, MAX_OPTIONS] probability tensor per context. temperature: as in score_rows."""
|
| 128 |
return self.score_rows(h, [self.tokenize(h, c, max_ctx_tokens) for c in contexts], temperature)
|
| 129 |
|
| 130 |
@torch.no_grad()
|
| 131 |
def score_rows(self, h, rows, temperature=1.0):
|
| 132 |
+
"""rows: [(suffix ids, slots)] from tokenize(). temperature: a number, or a list with one temperature per schema row
|
| 133 |
+
(h.nq values, in the order of prepare's questions), applied to every request."""
|
| 134 |
+
if isinstance(temperature, (list, tuple)) and len(temperature) != h.nq:
|
| 135 |
+
raise ValueError(f"temperature: {len(temperature)} values for a schema with {h.nq} rows")
|
| 136 |
Tmax = max(len(r[0]) for r in rows); Ts = next((t for t in TS_BUCKETS if t >= Tmax), None); n = len(rows)
|
| 137 |
R = next((b for b in B_BUCKETS if b >= n), n) if Ts else n; Ts = Ts or -(-Tmax // 256) * 256
|
| 138 |
ids = fill_ids([x for x, _ in rows for _ in range(h.P)], R * h.P, Ts, self.tok.pad_token_id).to(self.dev, non_blocking=True)
|
|
|
|
| 146 |
rws = [r for r in range(n) for _ in range(h.nq)]; sls = [x for _, sl in rows for x in sl]
|
| 147 |
else: # independent: one slot in each of the request's P rows
|
| 148 |
rws = [r * h.P + p for r in range(n) for p in range(h.P)]; sls = [sl[0] for _, sl in rows for _ in range(h.P)]
|
| 149 |
+
temps = list(temperature) * n if isinstance(temperature, (list, tuple)) else temperature
|
| 150 |
+
return read_slots(out, rws, sls, h.nopts * n, temps, [h.nq] * n)
|
decider/serve.py
CHANGED
|
@@ -40,6 +40,12 @@ Variables (default): DECIDER_DEVICE (auto: cuda, else mps, else cpu) DECIDER_C
|
|
| 40 |
DECIDER_T_BUCKETS DECIDER_B_BUCKETS DECIDER_GRAPH_TOKEN_BUDGET (32768) DECIDER_WARMUP (1) DECIDER_TOKENIZE_THREADS (8)
|
| 41 |
DECIDER_MAX_ROWS (1024) DECIDER_MAX_ROW_TOKENS (DECIDER_MAX_STATE_TOKENS + 4096) DECIDER_MAX_REQUEST_TOKENS (1048576)
|
| 42 |
DECIDER_MAX_QUEUE_ROWS (4096) DECIDER_TEMPERATURE DECIDER_SCHEMA_CACHE (0) DECIDER_SCHEMA_MIN_SEEN (2) DECIDER_SCHEMAS
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
"""
|
| 44 |
import asyncio, json, os, time
|
| 45 |
from concurrent.futures import ThreadPoolExecutor
|
|
@@ -47,6 +53,7 @@ from contextlib import asynccontextmanager
|
|
| 47 |
from fastapi import FastAPI, HTTPException
|
| 48 |
from pydantic import BaseModel
|
| 49 |
from decider import systemone as S1
|
|
|
|
| 50 |
from decider.batching import DEFAULT_MERGE_OVERHEAD_TOKENS, plan_batches
|
| 51 |
from decider.prompt import build, MAX_OPTIONS, resolve_layout, chat_template
|
| 52 |
from decider.prompt_fast import build_rows, unique_tokens
|
|
@@ -78,7 +85,7 @@ MAX_REQUEST_TOKENS = _env_int("DECIDER_MAX_REQUEST_TOKENS", 1 << 20) # sum
|
|
| 78 |
MAX_QUEUE_ROWS = _env_int("DECIDER_MAX_QUEUE_ROWS", 4096) # rows admitted and not yet scored, over all requests
|
| 79 |
DECIDE_MAX_CTX_TOKENS = 1536 # /decide context cap, unchanged from 1.0.x
|
| 80 |
|
| 81 |
-
MODEL_NAME = "decider"; TEMP = 1.0; TEMP_SCHEMA = 1.0; RELEASE_DATE = "2026-09-17"; ISOLATED = False; NEUTRALIZE_NONE = True
|
| 82 |
SCHEMA_FIRST = False; LAYOUT = "plain"; CHAT = None; se = None; squeue = None; schemas = {}; seen = {}
|
| 83 |
eng = None; queue = None; gpu = None; cpu = None; batcher_task = None; schema_task = None
|
| 84 |
outstanding = 0; REQ_SEQ = 0
|
|
@@ -118,6 +125,9 @@ def prepare(tok, state, questions, independent, isolated=False, max_state_tokens
|
|
| 118 |
pairs = [(r["question"], list(r["options"])) for r in flat]
|
| 119 |
rows = [[p] for p in pairs] if independent else [pairs]
|
| 120 |
items, ctx_len = build_rows(tok, ctx, rows, max_ctx_tokens=max_state_tokens, chat=chat)
|
|
|
|
|
|
|
|
|
|
| 121 |
return rqs, index, items, ctx_len
|
| 122 |
|
| 123 |
|
|
@@ -133,6 +143,7 @@ def _prepare_decide(context, schema):
|
|
| 133 |
q["options"], q["_back"] = neutralize_options(q["options"])
|
| 134 |
ex = Example(context, [Q(q["question"], list(q["options"]), 0) for q in qs])
|
| 135 |
it = build(ex, eng.tok, _NoShuffle(), max_options=MAX_OPTIONS, max_ctx_tokens=DECIDE_MAX_CTX_TOKENS, chat=CHAT)
|
|
|
|
| 136 |
return qs, it
|
| 137 |
|
| 138 |
|
|
@@ -188,11 +199,11 @@ def _release(n):
|
|
| 188 |
|
| 189 |
# ---- GPU work ----------------------------------------------------------------
|
| 190 |
def _score_items(items):
|
| 191 |
-
return eng.score_items(items, temperature=TEMP)
|
| 192 |
|
| 193 |
|
| 194 |
def _score_shared(items):
|
| 195 |
-
return eng.score_shared(items, temperature=TEMP)
|
| 196 |
|
| 197 |
|
| 198 |
async def _collect(q, wait_ms=None, adaptive_ms=0.0, idle_reset_ms=None):
|
|
@@ -294,6 +305,7 @@ def _schema_handle(questions, independent, compile=False):
|
|
| 294 |
old = next(iter(schemas)); hid = schemas.pop(old)[1].id
|
| 295 |
for k in [k for k in se.graphs if k[0] == hid]: del se.graphs[k]
|
| 296 |
h = se.prepare(rows, independent=independent, compile=compile)
|
|
|
|
| 297 |
schemas[key] = (rqs, h, index)
|
| 298 |
return schemas[key]
|
| 299 |
|
|
@@ -308,7 +320,7 @@ def _worth_caching(questions, independent):
|
|
| 308 |
|
| 309 |
|
| 310 |
def _score_schema(h, rows):
|
| 311 |
-
return se.score_rows(h, rows, temperature=TEMP_SCHEMA)
|
| 312 |
|
| 313 |
|
| 314 |
async def schema_batcher():
|
|
@@ -333,12 +345,13 @@ async def schema_batcher():
|
|
| 333 |
|
| 334 |
# ---- start-up / shutdown -------------------------------------------------------
|
| 335 |
def apply_config(cfg):
|
| 336 |
-
global MODEL_NAME, TEMP, TEMP_SCHEMA, RELEASE_DATE, ISOLATED, NEUTRALIZE_NONE, SCHEMA_FIRST, LAYOUT
|
| 337 |
LAYOUT = resolve_layout(cfg) # ValueError for an unknown layout, before the engine is built
|
| 338 |
NEUTRALIZE_NONE = bool(cfg.get("neutralize_none", True))
|
| 339 |
MODEL_NAME = "decider-" + str(cfg.get("version", "dev"))
|
| 340 |
-
|
| 341 |
-
|
|
|
|
| 342 |
RELEASE_DATE = str(cfg.get("release_date", RELEASE_DATE))
|
| 343 |
ISOLATED = bool(cfg.get("isolated_levels", False))
|
| 344 |
trained = bool(cfg.get("schema_first", False) or cfg.get("schema_first_trained", False))
|
|
@@ -403,7 +416,9 @@ async def _start():
|
|
| 403 |
t = await loop.run_in_executor(gpu, lambda: eng.warmup(log=lambda s: print(s, flush=True)))
|
| 404 |
print(f"[serve] captured {len(eng.graphs)} graphs in {t:.0f}s", flush=True)
|
| 405 |
eng.seal()
|
| 406 |
-
print("[serve] ready", json.dumps(dict(model=MODEL_NAME, layout=LAYOUT, temperature=TEMP,
|
|
|
|
|
|
|
| 407 |
shared=SHARED, graphs=len(eng.graphs), limits=dict(
|
| 408 |
max_rows=MAX_ROWS, max_row_tokens=MAX_ROW_TOKENS, max_request_tokens=MAX_REQUEST_TOKENS,
|
| 409 |
max_queue_rows=MAX_QUEUE_ROWS))), flush=True)
|
|
@@ -535,7 +550,9 @@ async def models():
|
|
| 535 |
@app.get("/health")
|
| 536 |
async def health():
|
| 537 |
return {"ok": eng is not None and bool(getattr(eng, "sealed", True)) and _alive(), "model": MODEL,
|
| 538 |
-
"device": str(getattr(eng, "dev", "")) if eng is not None else None, "layout": LAYOUT
|
|
|
|
|
|
|
| 539 |
|
| 540 |
|
| 541 |
@app.get("/stats")
|
|
|
|
| 40 |
DECIDER_T_BUCKETS DECIDER_B_BUCKETS DECIDER_GRAPH_TOKEN_BUDGET (32768) DECIDER_WARMUP (1) DECIDER_TOKENIZE_THREADS (8)
|
| 41 |
DECIDER_MAX_ROWS (1024) DECIDER_MAX_ROW_TOKENS (DECIDER_MAX_STATE_TOKENS + 4096) DECIDER_MAX_REQUEST_TOKENS (1048576)
|
| 42 |
DECIDER_MAX_QUEUE_ROWS (4096) DECIDER_TEMPERATURE DECIDER_SCHEMA_CACHE (0) DECIDER_SCHEMA_MIN_SEEN (2) DECIDER_SCHEMAS
|
| 43 |
+
|
| 44 |
+
Temperatures (1.4.0, decider.temperature): decider_config.json "temperature", and optionally "temperature_by_type"
|
| 45 |
+
{"choice": T, "noul": T, "score": T} (a /decide "bool" field is "noul", a "scale" field is "score"; a missing type uses
|
| 46 |
+
"temperature"). Every row carries the answer type of each of its slots, so a batch that mixes requests and types still
|
| 47 |
+
applies each answer's own temperature. DECIDER_TEMPERATURE replaces "temperature" and switches the by-type map off.
|
| 48 |
+
/health and the ready line report the temperature every answer type gets.
|
| 49 |
"""
|
| 50 |
import asyncio, json, os, time
|
| 51 |
from concurrent.futures import ThreadPoolExecutor
|
|
|
|
| 53 |
from fastapi import FastAPI, HTTPException
|
| 54 |
from pydantic import BaseModel
|
| 55 |
from decider import systemone as S1
|
| 56 |
+
from decider import temperature as TT
|
| 57 |
from decider.batching import DEFAULT_MERGE_OVERHEAD_TOKENS, plan_batches
|
| 58 |
from decider.prompt import build, MAX_OPTIONS, resolve_layout, chat_template
|
| 59 |
from decider.prompt_fast import build_rows, unique_tokens
|
|
|
|
| 85 |
MAX_QUEUE_ROWS = _env_int("DECIDER_MAX_QUEUE_ROWS", 4096) # rows admitted and not yet scored, over all requests
|
| 86 |
DECIDE_MAX_CTX_TOKENS = 1536 # /decide context cap, unchanged from 1.0.x
|
| 87 |
|
| 88 |
+
MODEL_NAME = "decider"; TEMP = 1.0; TEMP_SCHEMA = 1.0; TEMP_BY_TYPE = {}; TEMP_SCHEMA_BY_TYPE = {}; RELEASE_DATE = "2026-09-17"; ISOLATED = False; NEUTRALIZE_NONE = True
|
| 89 |
SCHEMA_FIRST = False; LAYOUT = "plain"; CHAT = None; se = None; squeue = None; schemas = {}; seen = {}
|
| 90 |
eng = None; queue = None; gpu = None; cpu = None; batcher_task = None; schema_task = None
|
| 91 |
outstanding = 0; REQ_SEQ = 0
|
|
|
|
| 125 |
pairs = [(r["question"], list(r["options"])) for r in flat]
|
| 126 |
rows = [[p] for p in pairs] if independent else [pairs]
|
| 127 |
items, ctx_len = build_rows(tok, ctx, rows, max_ctx_tokens=max_state_tokens, chat=chat)
|
| 128 |
+
types = S1.row_types(rqs, index) # answer type per slot (decider.temperature)
|
| 129 |
+
for it, ts in zip(items, [[t] for t in types] if independent else [types]):
|
| 130 |
+
it["types"] = ts
|
| 131 |
return rqs, index, items, ctx_len
|
| 132 |
|
| 133 |
|
|
|
|
| 143 |
q["options"], q["_back"] = neutralize_options(q["options"])
|
| 144 |
ex = Example(context, [Q(q["question"], list(q["options"]), 0) for q in qs])
|
| 145 |
it = build(ex, eng.tok, _NoShuffle(), max_options=MAX_OPTIONS, max_ctx_tokens=DECIDE_MAX_CTX_TOKENS, chat=CHAT)
|
| 146 |
+
it["types"] = [q["_type"] for q in qs] # bool -> noul, scale -> score (decider.temperature)
|
| 147 |
return qs, it
|
| 148 |
|
| 149 |
|
|
|
|
| 199 |
|
| 200 |
# ---- GPU work ----------------------------------------------------------------
|
| 201 |
def _score_items(items):
|
| 202 |
+
return eng.score_items(items, temperature=TT.for_items(TEMP, TEMP_BY_TYPE, items)) # the scalar TEMP without a by-type map
|
| 203 |
|
| 204 |
|
| 205 |
def _score_shared(items):
|
| 206 |
+
return eng.score_shared(items, temperature=TT.for_items(TEMP, TEMP_BY_TYPE, items))
|
| 207 |
|
| 208 |
|
| 209 |
async def _collect(q, wait_ms=None, adaptive_ms=0.0, idle_reset_ms=None):
|
|
|
|
| 305 |
old = next(iter(schemas)); hid = schemas.pop(old)[1].id
|
| 306 |
for k in [k for k in se.graphs if k[0] == hid]: del se.graphs[k]
|
| 307 |
h = se.prepare(rows, independent=independent, compile=compile)
|
| 308 |
+
h.types = S1.row_types(rqs, index) # answer type per schema row (decider.temperature)
|
| 309 |
schemas[key] = (rqs, h, index)
|
| 310 |
return schemas[key]
|
| 311 |
|
|
|
|
| 320 |
|
| 321 |
|
| 322 |
def _score_schema(h, rows):
|
| 323 |
+
return se.score_rows(h, rows, temperature=TT.for_types(TEMP_SCHEMA, TEMP_SCHEMA_BY_TYPE, h.types))
|
| 324 |
|
| 325 |
|
| 326 |
async def schema_batcher():
|
|
|
|
| 345 |
|
| 346 |
# ---- start-up / shutdown -------------------------------------------------------
|
| 347 |
def apply_config(cfg):
|
| 348 |
+
global MODEL_NAME, TEMP, TEMP_SCHEMA, TEMP_BY_TYPE, TEMP_SCHEMA_BY_TYPE, RELEASE_DATE, ISOLATED, NEUTRALIZE_NONE, SCHEMA_FIRST, LAYOUT
|
| 349 |
LAYOUT = resolve_layout(cfg) # ValueError for an unknown layout, before the engine is built
|
| 350 |
NEUTRALIZE_NONE = bool(cfg.get("neutralize_none", True))
|
| 351 |
MODEL_NAME = "decider-" + str(cfg.get("version", "dev"))
|
| 352 |
+
# ValueError for a temperature that is not finite and > 0 or an unknown by-type key, before the engine is built.
|
| 353 |
+
# DECIDER_TEMPERATURE replaces "temperature" and switches "temperature_by_type" off (decider.temperature).
|
| 354 |
+
(TEMP, TEMP_BY_TYPE), (TEMP_SCHEMA, TEMP_SCHEMA_BY_TYPE) = TT.from_config(cfg, os.environ.get("DECIDER_TEMPERATURE"))
|
| 355 |
RELEASE_DATE = str(cfg.get("release_date", RELEASE_DATE))
|
| 356 |
ISOLATED = bool(cfg.get("isolated_levels", False))
|
| 357 |
trained = bool(cfg.get("schema_first", False) or cfg.get("schema_first_trained", False))
|
|
|
|
| 416 |
t = await loop.run_in_executor(gpu, lambda: eng.warmup(log=lambda s: print(s, flush=True)))
|
| 417 |
print(f"[serve] captured {len(eng.graphs)} graphs in {t:.0f}s", flush=True)
|
| 418 |
eng.seal()
|
| 419 |
+
print("[serve] ready", json.dumps(dict(model=MODEL_NAME, layout=LAYOUT, temperature=TEMP, temperature_by_type=TT.effective(TEMP, TEMP_BY_TYPE),
|
| 420 |
+
**({"temperature_schema_first_by_type": TT.effective(TEMP_SCHEMA, TEMP_SCHEMA_BY_TYPE)} if SCHEMA_FIRST else {}),
|
| 421 |
+
isolated_levels=ISOLATED, schema_first=SCHEMA_FIRST,
|
| 422 |
shared=SHARED, graphs=len(eng.graphs), limits=dict(
|
| 423 |
max_rows=MAX_ROWS, max_row_tokens=MAX_ROW_TOKENS, max_request_tokens=MAX_REQUEST_TOKENS,
|
| 424 |
max_queue_rows=MAX_QUEUE_ROWS))), flush=True)
|
|
|
|
| 550 |
@app.get("/health")
|
| 551 |
async def health():
|
| 552 |
return {"ok": eng is not None and bool(getattr(eng, "sealed", True)) and _alive(), "model": MODEL,
|
| 553 |
+
"device": str(getattr(eng, "dev", "")) if eng is not None else None, "layout": LAYOUT,
|
| 554 |
+
"temperature": TEMP, "temperature_by_type": TT.effective(TEMP, TEMP_BY_TYPE),
|
| 555 |
+
**({"temperature_schema_first_by_type": TT.effective(TEMP_SCHEMA, TEMP_SCHEMA_BY_TYPE)} if SCHEMA_FIRST else {})}
|
| 556 |
|
| 557 |
|
| 558 |
@app.get("/stats")
|
decider/shared_prefix.py
CHANGED
|
@@ -33,6 +33,7 @@ import torch
|
|
| 33 |
import torch.nn.functional as F
|
| 34 |
|
| 35 |
from decider.engine import fill_ids, read_slots
|
|
|
|
| 36 |
|
| 37 |
DEFAULT_FORK_GB = 8.0
|
| 38 |
STATE_NAMES = ("keys", "values", "indexer_keys", "conv_states", "recurrent_states")
|
|
@@ -159,6 +160,7 @@ def score_shared(engine, items, temperature=1.0, min_prefix=192, budget_bytes=No
|
|
| 159 |
request does not qualify (fewer than two rows, or a common prefix below `min_prefix`) and the caller should use
|
| 160 |
`score_items`.
|
| 161 |
|
|
|
|
| 162 |
`rows_per_fork` forces the chunk size; it exists for the tests that compare chunked against unchunked answers."""
|
| 163 |
ids = [it["ids"] for it in items]
|
| 164 |
n = len(ids)
|
|
@@ -167,6 +169,7 @@ def score_shared(engine, items, temperature=1.0, min_prefix=192, budget_bytes=No
|
|
| 167 |
lcp = common_prefix_len(ids)
|
| 168 |
if lcp < min_prefix:
|
| 169 |
return None
|
|
|
|
| 170 |
core, W, dev, pad = engine.core, engine.W, engine.dev, engine.tok.pad_token_id
|
| 171 |
pre = torch.tensor(ids[0][:lcp], device=dev)[None]
|
| 172 |
cache = core(input_ids=pre, use_cache=True).past_key_values
|
|
@@ -191,6 +194,7 @@ def score_shared(engine, items, temperature=1.0, min_prefix=192, budget_bytes=No
|
|
| 191 |
sl = [s - lcp for it in part for s in it["slots"]]
|
| 192 |
idx = torch.tensor([rows, sl], device=dev)
|
| 193 |
out += read_slots(F.linear(h[idx[0], idx[1]], W).float()[:, None, :], list(range(len(rows))), [0] * len(rows),
|
| 194 |
-
[k for it in part for k in it["nopts"]],
|
|
|
|
| 195 |
del fork, h, suf # drop this chunk's fork before the next one is built
|
| 196 |
return out
|
|
|
|
| 33 |
import torch.nn.functional as F
|
| 34 |
|
| 35 |
from decider.engine import fill_ids, read_slots
|
| 36 |
+
from decider.temperature import item_slice, slot_temperatures
|
| 37 |
|
| 38 |
DEFAULT_FORK_GB = 8.0
|
| 39 |
STATE_NAMES = ("keys", "values", "indexer_keys", "conv_states", "recurrent_states")
|
|
|
|
| 160 |
request does not qualify (fewer than two rows, or a common prefix below `min_prefix`) and the caller should use
|
| 161 |
`score_items`.
|
| 162 |
|
| 163 |
+
`temperature`: a number, or one entry per item (decider.temperature.slot_temperatures).
|
| 164 |
`rows_per_fork` forces the chunk size; it exists for the tests that compare chunked against unchunked answers."""
|
| 165 |
ids = [it["ids"] for it in items]
|
| 166 |
n = len(ids)
|
|
|
|
| 169 |
lcp = common_prefix_len(ids)
|
| 170 |
if lcp < min_prefix:
|
| 171 |
return None
|
| 172 |
+
slot_temperatures(temperature, items) # a length mismatch fails before any forward
|
| 173 |
core, W, dev, pad = engine.core, engine.W, engine.dev, engine.tok.pad_token_id
|
| 174 |
pre = torch.tensor(ids[0][:lcp], device=dev)[None]
|
| 175 |
cache = core(input_ids=pre, use_cache=True).past_key_values
|
|
|
|
| 194 |
sl = [s - lcp for it in part for s in it["slots"]]
|
| 195 |
idx = torch.tensor([rows, sl], device=dev)
|
| 196 |
out += read_slots(F.linear(h[idx[0], idx[1]], W).float()[:, None, :], list(range(len(rows))), [0] * len(rows),
|
| 197 |
+
[k for it in part for k in it["nopts"]],
|
| 198 |
+
slot_temperatures(item_slice(temperature, i, i + b), part), [len(it["slots"]) for it in part])
|
| 199 |
del fork, h, suf # drop this chunk's fork before the next one is built
|
| 200 |
return out
|
decider/systemone.py
CHANGED
|
@@ -3,7 +3,8 @@
|
|
| 3 |
state str | dict | list JSON state is serialised compactly; questions may name a part by path (`ticket.messages[0].text`)
|
| 4 |
questions {id: {"type": "choice", "instructions": ..., "criteria": {name: description | {...} | [...] | None}} up to 255 options
|
| 5 |
{"type": "score", "instructions": ..., "criteria": [level 0 description, level 1 description, ...]} 2..10 levels
|
| 6 |
-
{"type": "noul", "instructions": ..., "criteria": {"true": ..., "false": ...} (optional)}}
|
|
|
|
| 7 |
ids are never shown to the model. `instructions` and every description may be a string or any JSON value.
|
| 8 |
"""
|
| 9 |
import json, math
|
|
@@ -38,7 +39,15 @@ def render_state(state, index_arrays=True):
|
|
| 38 |
|
| 39 |
def render_question(spec):
|
| 40 |
"""-> dict(question=str, options=[str], type=..., names=[...]) (names: what the answer reports for each option)"""
|
| 41 |
-
t = spec.get("type", "choice");
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 42 |
if not ins:
|
| 43 |
raise ValueError("question without instructions")
|
| 44 |
if t == "choice":
|
|
@@ -65,6 +74,12 @@ def render_question(spec):
|
|
| 65 |
isolated=bool(spec.get("isolated", True)))
|
| 66 |
|
| 67 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 68 |
# ---- isolated levels: every Score level is judged in its own row, without its number or its neighbours
|
| 69 |
ISOLATED = "{q}\nProposed answer: {level}\nDoes the proposed answer fit?"
|
| 70 |
_NUM = None
|
|
@@ -100,6 +115,15 @@ def plan_rows(rqs, isolated=True):
|
|
| 100 |
return rows, index
|
| 101 |
|
| 102 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 103 |
def assemble(rqs, index, probs):
|
| 104 |
"""probs: one probability list per row (plan_rows order) -> {id: answer}."""
|
| 105 |
out = {}
|
|
@@ -118,16 +142,50 @@ def certainty(p):
|
|
| 118 |
return max(0.0, 1.0 - h / math.log(len(p))) if len(p) > 1 else 1.0
|
| 119 |
|
| 120 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 121 |
def format_answer(rq, p, nd=4):
|
| 122 |
-
"""rq: render_question output; p: probabilities in option order.
|
|
|
|
|
|
|
| 123 |
p = [float(x) for x in p[:len(rq["options"])]]; s = sum(p) or 1.0; p = [x / s for x in p]
|
| 124 |
j = max(range(len(p)), key=p.__getitem__)
|
| 125 |
if rq["type"] == "noul":
|
| 126 |
return {"type": "noul", "noul": round(p[1], nd)}
|
| 127 |
if rq["type"] == "choice":
|
| 128 |
-
return {"type": "choice", "choice": rq["names"][j], "confidence": round(p
|
| 129 |
-
"probabilities": {n: round(x, nd) for n, x in zip(rq["names"], p)}}
|
| 130 |
-
return {"type": "score", "score": round(sum(i * x for i, x in enumerate(p)), 2), "confidence": round(p
|
|
|
|
| 131 |
"legend": {str(i): d for i, d in enumerate(rq["legend"])}, "probabilities": {str(i): round(x, nd) for i, x in enumerate(p)}}
|
| 132 |
|
| 133 |
|
|
|
|
| 3 |
state str | dict | list JSON state is serialised compactly; questions may name a part by path (`ticket.messages[0].text`)
|
| 4 |
questions {id: {"type": "choice", "instructions": ..., "criteria": {name: description | {...} | [...] | None}} up to 255 options
|
| 5 |
{"type": "score", "instructions": ..., "criteria": [level 0 description, level 1 description, ...]} 2..10 levels
|
| 6 |
+
{"type": "noul", "instructions": ... (optional), "criteria": {"true": ..., "false": ...} (optional)}}
|
| 7 |
+
a noul question needs instructions or at least one true/false description.
|
| 8 |
ids are never shown to the model. `instructions` and every description may be a string or any JSON value.
|
| 9 |
"""
|
| 10 |
import json, math
|
|
|
|
| 39 |
|
| 40 |
def render_question(spec):
|
| 41 |
"""-> dict(question=str, options=[str], type=..., names=[...]) (names: what the answer reports for each option)"""
|
| 42 |
+
t = spec.get("type", "choice"); crit = spec.get("criteria", spec.get("options"))
|
| 43 |
+
raw = spec.get("instructions", spec.get("question", ""))
|
| 44 |
+
if t in ("noul", "bool") and raw in (None, ""): # the criteria carry the question (NOUL_WITHOUT_INSTRUCTIONS)
|
| 45 |
+
ins = NOUL_WITHOUT_INSTRUCTIONS
|
| 46 |
+
described = isinstance(crit, dict) and any(d not in (None, "") for d in (crit.get("true", crit.get(True)), crit.get("false", crit.get(False))))
|
| 47 |
+
if not described and (crit is None or isinstance(crit, dict)): # a non-map is rejected below with the criteria message
|
| 48 |
+
raise ValueError("noul question without instructions: criteria must describe true or false")
|
| 49 |
+
else:
|
| 50 |
+
ins = _txt(raw)
|
| 51 |
if not ins:
|
| 52 |
raise ValueError("question without instructions")
|
| 53 |
if t == "choice":
|
|
|
|
| 74 |
isolated=bool(spec.get("isolated", True)))
|
| 75 |
|
| 76 |
|
| 77 |
+
# A noul question may omit `instructions` (TypeSafe's OpenAPI file marks it optional). The question id is never shown to the
|
| 78 |
+
# model, so the question text is this fixed sentence and the true/false descriptions, rendered as the options "no: ..." and
|
| 79 |
+
# "yes: ...", say what is being asked. A request that gives instructions is rendered exactly as before.
|
| 80 |
+
NOUL_WITHOUT_INSTRUCTIONS = "Which answer fits the context?"
|
| 81 |
+
|
| 82 |
+
|
| 83 |
# ---- isolated levels: every Score level is judged in its own row, without its number or its neighbours
|
| 84 |
ISOLATED = "{q}\nProposed answer: {level}\nDoes the proposed answer fit?"
|
| 85 |
_NUM = None
|
|
|
|
| 115 |
return rows, index
|
| 116 |
|
| 117 |
|
| 118 |
+
def row_types(rqs, index):
|
| 119 |
+
"""The answer type ("choice", "noul" or "score") of every plan_rows row, in row order. An isolated Score question's
|
| 120 |
+
yes/no level rows carry "score": together they are one Score answer (decider.temperature)."""
|
| 121 |
+
types = [None] * sum(n for _, _, _, n in index)
|
| 122 |
+
for k, _, s, n in index:
|
| 123 |
+
types[s:s + n] = [rqs[k]["type"]] * n
|
| 124 |
+
return types
|
| 125 |
+
|
| 126 |
+
|
| 127 |
def assemble(rqs, index, probs):
|
| 128 |
"""probs: one probability list per row (plan_rows order) -> {id: answer}."""
|
| 129 |
out = {}
|
|
|
|
| 142 |
return max(0.0, 1.0 - h / math.log(len(p))) if len(p) > 1 else 1.0
|
| 143 |
|
| 144 |
|
| 145 |
+
def _clip01(x):
|
| 146 |
+
return min(1.0, max(0.0, x))
|
| 147 |
+
|
| 148 |
+
|
| 149 |
+
def _normalised(p):
|
| 150 |
+
"""As the adapter's _normalize: a distribution with zero total counts as uniform."""
|
| 151 |
+
tot = sum(p)
|
| 152 |
+
return [1.0 / len(p)] * len(p) if tot == 0 else [x / tot for x in p]
|
| 153 |
+
|
| 154 |
+
|
| 155 |
+
def choice_confidence(p):
|
| 156 |
+
"""TypeSafe's Choice confidence: the largest probability rescaled so that a uniform distribution gives 0 and all mass on one
|
| 157 |
+
option gives 1, (n * p_max - 1) / (n - 1); 1 for a single option."""
|
| 158 |
+
n = len(p); p = _normalised(p)
|
| 159 |
+
return 1.0 if n <= 1 else _clip01((n * max(p) - 1) / (n - 1))
|
| 160 |
+
|
| 161 |
+
|
| 162 |
+
def score_confidence(p):
|
| 163 |
+
"""TypeSafe's Score confidence (system-one-adapter-python, confidence_metrics.score_confidence): 1 minus the expected distance
|
| 164 |
+
from the most likely level, divided by D = mean over levels i of |i - (n - 1)/2| (the mean distance of the levels from the
|
| 165 |
+
middle of the scale), floored at 0; 1 for a single level.
|
| 166 |
+
With two levels it equals choice_confidence."""
|
| 167 |
+
n = len(p); p = _normalised(p)
|
| 168 |
+
if n <= 1:
|
| 169 |
+
return 1.0
|
| 170 |
+
k = max(range(n), key=p.__getitem__)
|
| 171 |
+
spread = sum(x * abs(i - k) for i, x in enumerate(p))
|
| 172 |
+
uniform = sum(abs(i - (n - 1) / 2) for i in range(n)) / n
|
| 173 |
+
return _clip01(1.0 - spread / uniform)
|
| 174 |
+
|
| 175 |
+
|
| 176 |
def format_answer(rq, p, nd=4):
|
| 177 |
+
"""rq: render_question output; p: probabilities in option order.
|
| 178 |
+
`confidence` is TypeSafe's (choice_confidence / score_confidence); `x_p_max` is the largest probability, which was
|
| 179 |
+
`confidence` before 1.3.0."""
|
| 180 |
p = [float(x) for x in p[:len(rq["options"])]]; s = sum(p) or 1.0; p = [x / s for x in p]
|
| 181 |
j = max(range(len(p)), key=p.__getitem__)
|
| 182 |
if rq["type"] == "noul":
|
| 183 |
return {"type": "noul", "noul": round(p[1], nd)}
|
| 184 |
if rq["type"] == "choice":
|
| 185 |
+
return {"type": "choice", "choice": rq["names"][j], "confidence": round(choice_confidence(p), nd), "x_p_max": round(p[j], nd),
|
| 186 |
+
"certainty": round(certainty(p), nd), "probabilities": {n: round(x, nd) for n, x in zip(rq["names"], p)}}
|
| 187 |
+
return {"type": "score", "score": round(sum(i * x for i, x in enumerate(p)), 2), "confidence": round(score_confidence(p), nd), "x_p_max": round(p[j], nd),
|
| 188 |
+
"certainty": round(certainty(p), nd),
|
| 189 |
"legend": {str(i): d for i, d in enumerate(rq["legend"])}, "probabilities": {str(i): round(x, nd) for i, x in enumerate(p)}}
|
| 190 |
|
| 191 |
|
decider/temperature.py
ADDED
|
@@ -0,0 +1,144 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Answer temperatures from decider_config.json (1.4.0).
|
| 2 |
+
|
| 3 |
+
Every answer is softmax(logits / T) over its option letters. decider_config.json sets T:
|
| 4 |
+
|
| 5 |
+
"temperature": 1.3 one value for every answer (the only form before 1.4.0)
|
| 6 |
+
"temperature_by_type": {"choice": 1.48, "noul": 2.22, "score": 1.38}
|
| 7 |
+
optional; one value per answer type, a missing type uses "temperature"
|
| 8 |
+
"temperature_schema_first": 1.18 optional; the schema cache (questions-first layout), as before
|
| 9 |
+
"temperature_schema_first_by_type": {...} optional; per answer type on the schema cache
|
| 10 |
+
|
| 11 |
+
The answer types are the /v1/systemone question types. A /decide field maps onto them: "choice" -> "choice", "bool" -> "noul",
|
| 12 |
+
"scale" -> "score". Plain `Decider.decide` questions (a question and its options, no type) are "choice". A Score question
|
| 13 |
+
read with isolated levels (one yes/no row per level) uses the "score" temperature on every one of its level rows: the rows form
|
| 14 |
+
one Score answer, and a temperature fitted on Score answers is fitted through that same readout (decider.calibrate).
|
| 15 |
+
|
| 16 |
+
On the state-first layout the temperature of an answer of type t is temperature_by_type[t], else temperature. On the schema
|
| 17 |
+
cache it is temperature_schema_first_by_type[t], else temperature_schema_first, else (no schema-first value at all) the
|
| 18 |
+
state-first temperature of t. An explicit override (Decider(temperature=...), DECIDER_TEMPERATURE) replaces the state-first
|
| 19 |
+
"temperature" and switches temperature_by_type off, so the override is the one temperature of every state-first answer, as it
|
| 20 |
+
was before 1.4.0.
|
| 21 |
+
|
| 22 |
+
A config without a by-type map gives every path the single number it gave in 1.3.0, and that number reaches the engines as the
|
| 23 |
+
same Python float, so the probabilities are bit-identical to 1.3.0.
|
| 24 |
+
"""
|
| 25 |
+
import math
|
| 26 |
+
|
| 27 |
+
TYPES = ("choice", "noul", "score")
|
| 28 |
+
FIELD_TYPES = {"choice": "choice", "bool": "noul", "scale": "score"} # /decide field type -> answer type
|
| 29 |
+
_KEYS_TEXT = ('the keys are "choice", "noul" and "score" (a /decide "bool" field is "noul", a "scale" field is "score"; '
|
| 30 |
+
'a /v1/systemone "bool" question is "noul")')
|
| 31 |
+
|
| 32 |
+
|
| 33 |
+
def positive(value, where):
|
| 34 |
+
"""A temperature: a number (or a numeric string, as float() reads it, which 1.3.0 accepted for "temperature") that is finite
|
| 35 |
+
and > 0. Raises ValueError naming `where`."""
|
| 36 |
+
if isinstance(value, bool):
|
| 37 |
+
raise ValueError(f"{where} must be a finite number > 0, got {value!r}")
|
| 38 |
+
try:
|
| 39 |
+
v = float(value)
|
| 40 |
+
except (TypeError, ValueError):
|
| 41 |
+
raise ValueError(f"{where} must be a finite number > 0, got {value!r}") from None
|
| 42 |
+
if not math.isfinite(v) or v <= 0:
|
| 43 |
+
raise ValueError(f"{where} must be a finite number > 0, got {value!r}")
|
| 44 |
+
return v
|
| 45 |
+
|
| 46 |
+
|
| 47 |
+
def by_type(m, where):
|
| 48 |
+
"""Validate a {answer type: temperature} map. None -> {}. Unknown keys, non-numbers and values that are not finite and > 0
|
| 49 |
+
raise ValueError."""
|
| 50 |
+
if m is None:
|
| 51 |
+
return {}
|
| 52 |
+
if not isinstance(m, dict):
|
| 53 |
+
raise ValueError(f"{where} must be a map {{answer type: temperature}}, got {type(m).__name__}; " + _KEYS_TEXT)
|
| 54 |
+
out = {}
|
| 55 |
+
for k, v in m.items():
|
| 56 |
+
if k not in TYPES:
|
| 57 |
+
raise ValueError(f"{where} has the unknown key {k!r}; " + _KEYS_TEXT)
|
| 58 |
+
if isinstance(v, bool) or not isinstance(v, (int, float)):
|
| 59 |
+
raise ValueError(f"{where}[{k!r}] must be a finite number > 0, got {v!r}")
|
| 60 |
+
out[k] = positive(v, f"{where}[{k!r}]")
|
| 61 |
+
return out
|
| 62 |
+
|
| 63 |
+
|
| 64 |
+
def from_config(cfg, temperature=None, temperature_by_type=None):
|
| 65 |
+
"""-> ((T, by_type) for the state-first layout, (T, by_type) for the schema cache).
|
| 66 |
+
|
| 67 |
+
temperature / temperature_by_type: explicit overrides (Decider arguments, DECIDER_TEMPERATURE). An explicit temperature
|
| 68 |
+
without an explicit map switches the config's map off (see the module docstring)."""
|
| 69 |
+
cfg = cfg or {}
|
| 70 |
+
where = "decider_config.json"
|
| 71 |
+
if temperature is not None:
|
| 72 |
+
T = positive(temperature, "temperature")
|
| 73 |
+
m = by_type(temperature_by_type, "temperature_by_type") if temperature_by_type is not None else {}
|
| 74 |
+
else:
|
| 75 |
+
T = positive(cfg.get("temperature", 1.0), f'{where} "temperature"')
|
| 76 |
+
m = (by_type(temperature_by_type, "temperature_by_type") if temperature_by_type is not None
|
| 77 |
+
else by_type(cfg.get("temperature_by_type"), f'{where} "temperature_by_type"'))
|
| 78 |
+
ms = by_type(cfg.get("temperature_schema_first_by_type"), f'{where} "temperature_schema_first_by_type"')
|
| 79 |
+
if "temperature_schema_first" in cfg:
|
| 80 |
+
schema = (positive(cfg["temperature_schema_first"], f'{where} "temperature_schema_first"'), ms)
|
| 81 |
+
else:
|
| 82 |
+
schema = (T, {**m, **ms})
|
| 83 |
+
return (T, m), schema
|
| 84 |
+
|
| 85 |
+
|
| 86 |
+
def effective(T, m):
|
| 87 |
+
"""{answer type: the temperature it gets} for reporting (/health, the ready line)."""
|
| 88 |
+
return {t: m.get(t, T) for t in TYPES}
|
| 89 |
+
|
| 90 |
+
|
| 91 |
+
def for_types(T, m, types):
|
| 92 |
+
"""One temperature per slot, or the scalar T itself when there is no map (the 1.3.0 call)."""
|
| 93 |
+
if not m:
|
| 94 |
+
return T
|
| 95 |
+
return [m.get(t, T) for t in types]
|
| 96 |
+
|
| 97 |
+
|
| 98 |
+
def item_types(it):
|
| 99 |
+
"""The answer type of every slot of a prompt item ("types", set where the item is built); None for an item without it."""
|
| 100 |
+
ts = it.get("types")
|
| 101 |
+
return list(ts) if ts is not None else [None] * len(it["slots"])
|
| 102 |
+
|
| 103 |
+
|
| 104 |
+
def for_items(T, m, items):
|
| 105 |
+
"""The `temperature` argument of Engine.score_items / score_shared: the scalar T when there is no map (the 1.3.0 call),
|
| 106 |
+
else one list of per-slot temperatures per item."""
|
| 107 |
+
if not m:
|
| 108 |
+
return T
|
| 109 |
+
return [[m.get(t, T) for t in item_types(it)] for it in items]
|
| 110 |
+
|
| 111 |
+
|
| 112 |
+
def slot_temperatures(temperature, items):
|
| 113 |
+
"""Engine side. temperature: a scalar (returned as it is), or one entry per item, each a scalar or one value per slot of
|
| 114 |
+
that item. -> the scalar, or a flat list with one temperature per slot in item order."""
|
| 115 |
+
if not isinstance(temperature, (list, tuple)):
|
| 116 |
+
return temperature
|
| 117 |
+
if len(temperature) != len(items):
|
| 118 |
+
raise ValueError(f"temperature: {len(temperature)} entries for {len(items)} items")
|
| 119 |
+
flat = []
|
| 120 |
+
for t, it in zip(temperature, items):
|
| 121 |
+
n = len(it["slots"])
|
| 122 |
+
if isinstance(t, (list, tuple)):
|
| 123 |
+
if len(t) != n:
|
| 124 |
+
raise ValueError(f"temperature: {len(t)} values for an item with {n} slots")
|
| 125 |
+
flat += list(t)
|
| 126 |
+
else:
|
| 127 |
+
flat += [t] * n
|
| 128 |
+
return flat
|
| 129 |
+
|
| 130 |
+
|
| 131 |
+
def item_slice(temperature, lo, hi):
|
| 132 |
+
"""The per-item temperature entries of items[lo:hi] (a scalar is shared by every item)."""
|
| 133 |
+
return temperature[lo:hi] if isinstance(temperature, (list, tuple)) else temperature
|
| 134 |
+
|
| 135 |
+
|
| 136 |
+
def scaled_softmax(lg, temperature):
|
| 137 |
+
"""softmax(lg / T) over the last axis. A scalar T is the 1.3.0 expression unchanged; a list gives one T per row of lg."""
|
| 138 |
+
import torch
|
| 139 |
+
if isinstance(temperature, (list, tuple)):
|
| 140 |
+
if len(temperature) != lg.shape[0]:
|
| 141 |
+
raise ValueError(f"temperature: {len(temperature)} values for {lg.shape[0]} slots")
|
| 142 |
+
t = torch.tensor(temperature, dtype=lg.dtype).to(lg.device, non_blocking=True)[:, None]
|
| 143 |
+
return torch.softmax(lg / t, -1)
|
| 144 |
+
return torch.softmax(lg / temperature, -1)
|
decider_config.json
CHANGED
|
@@ -1,7 +1,12 @@
|
|
| 1 |
{
|
| 2 |
-
"temperature": 1.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
"neutralize_none": false,
|
| 4 |
-
"version": "4b-v2",
|
| 5 |
"base": "Mapika/decider-4b v1 + LoRA (merged); v1 is Qwen/Qwen3.5-4B-Base + one supervised pass over mixture v2",
|
| 6 |
"layout": "plain",
|
| 7 |
"max_options": 255,
|
|
@@ -10,5 +15,6 @@
|
|
| 10 |
"schema_first_trained": false,
|
| 11 |
"isolated_levels": true,
|
| 12 |
"release_date": "2026-09-24",
|
| 13 |
-
"
|
|
|
|
| 14 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"temperature": 1.099,
|
| 3 |
+
"temperature_by_type": {
|
| 4 |
+
"choice": 1.11,
|
| 5 |
+
"noul": 1.56,
|
| 6 |
+
"score": 1.287
|
| 7 |
+
},
|
| 8 |
"neutralize_none": false,
|
| 9 |
+
"version": "4b-v2.1",
|
| 10 |
"base": "Mapika/decider-4b v1 + LoRA (merged); v1 is Qwen/Qwen3.5-4B-Base + one supervised pass over mixture v2",
|
| 11 |
"layout": "plain",
|
| 12 |
"max_options": 255,
|
|
|
|
| 15 |
"schema_first_trained": false,
|
| 16 |
"isolated_levels": true,
|
| 17 |
"release_date": "2026-09-24",
|
| 18 |
+
"requires": "decider-ai>=1.4.0 for temperature_by_type; older versions serve every answer at temperature",
|
| 19 |
+
"stage": "decider-4b v1 + LoRA rank 64 (alpha 128) on attention and MLP, LR 1e-4, 2 epochs (1,518 steps of 65,536 tokens) over v2's 29,325-row mix in the plain state-first layout (generated decision families with code-computed answers, Qwen3.6-27B-written document questions kept when two independent answers agreed, human-labelled public sets, replay of v1's mixture v2), with the replay rows trained toward v1's own answer distribution (KL to v1) instead of their labels, merged into the bf16 weights; no RL stage; temperature fitted by NLL on 61 in-task regression tasks (the 67 in-task tasks without banking77, clinc_oos, mmlu, arc, winogrande, hellaswag); temperature_by_type fitted with decider.calibrate.fit_by_type on the same regression rows plus our own validation rows (choice, noul and score answers)"
|
| 20 |
}
|
eval_results.json
CHANGED
|
@@ -1,1796 +1,1549 @@
|
|
| 1 |
{
|
| 2 |
-
"model": "decider-4b-v2",
|
| 3 |
-
"temperature": 1.
|
| 4 |
-
"
|
| 5 |
-
"
|
| 6 |
-
"
|
| 7 |
-
"
|
| 8 |
-
"excluded_in_task_tasks": [
|
| 9 |
-
"banking77",
|
| 10 |
-
"clinc_oos",
|
| 11 |
-
"mmlu",
|
| 12 |
-
"arc",
|
| 13 |
-
"winogrande",
|
| 14 |
-
"hellaswag"
|
| 15 |
-
],
|
| 16 |
-
"fitted": 1.9349278889417365,
|
| 17 |
-
"grid_fit_temperature": 1.9420684917099538,
|
| 18 |
-
"stored": 1.935,
|
| 19 |
-
"note": "fitted on the in-task half of the public regression set without banking77, clinc_oos, mmlu, arc, winogrande and hellaswag (61 tasks); no JevBench or Decision Index item was used for training, selection or temperature"
|
| 20 |
},
|
| 21 |
-
"
|
| 22 |
-
|
| 23 |
-
"
|
| 24 |
-
|
| 25 |
-
"
|
| 26 |
-
"
|
| 27 |
-
"
|
| 28 |
-
},
|
| 29 |
-
"
|
| 30 |
-
"acc": 0.7787589834608305,
|
| 31 |
-
"nll": 0.5663163406508309,
|
| 32 |
-
"ece": 0.08033093216064069,
|
| 33 |
-
"tasks": 28
|
| 34 |
-
}
|
| 35 |
},
|
| 36 |
-
"
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
"
|
| 40 |
-
"
|
| 41 |
-
"
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
"
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
"
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
"
|
| 51 |
-
"
|
| 52 |
-
"
|
| 53 |
-
"ece": 0.04695633757114412
|
| 54 |
-
},
|
| 55 |
-
"agenttraj": {
|
| 56 |
-
"heldout": false,
|
| 57 |
-
"acc": 0.8973333333333333,
|
| 58 |
-
"nll": 0.2499454766511917,
|
| 59 |
-
"ece": 0.03513214365641278
|
| 60 |
-
},
|
| 61 |
-
"amazon_stars": {
|
| 62 |
-
"heldout": false,
|
| 63 |
-
"acc": 0.602,
|
| 64 |
-
"nll": 0.9151002764701843,
|
| 65 |
-
"ece": 0.035494098405043265
|
| 66 |
-
},
|
| 67 |
-
"arc": {
|
| 68 |
-
"heldout": false,
|
| 69 |
-
"acc": 0.939419795221843,
|
| 70 |
-
"nll": 0.1868872344493866,
|
| 71 |
-
"ece": 0.0222670579577062
|
| 72 |
-
},
|
| 73 |
-
"arena_pref": {
|
| 74 |
-
"heldout": true,
|
| 75 |
-
"acc": 0.5033333333333333,
|
| 76 |
-
"nll": 1.2348569631576538,
|
| 77 |
-
"ece": 0.21849741025765737
|
| 78 |
-
},
|
| 79 |
-
"banking77": {
|
| 80 |
-
"heldout": false,
|
| 81 |
-
"acc": 0.972,
|
| 82 |
-
"nll": 0.10399508476257324,
|
| 83 |
-
"ece": 0.0268547142942746
|
| 84 |
-
},
|
| 85 |
-
"bbc_news": {
|
| 86 |
-
"heldout": true,
|
| 87 |
-
"acc": 0.951,
|
| 88 |
-
"nll": 0.18327508866786957,
|
| 89 |
-
"ece": 0.06801495689153671
|
| 90 |
-
},
|
| 91 |
-
"bias_in_bios": {
|
| 92 |
-
"heldout": false,
|
| 93 |
-
"acc": 0.9486666666666667,
|
| 94 |
-
"nll": 0.15249282121658325,
|
| 95 |
-
"ece": 0.013438965007662774
|
| 96 |
-
},
|
| 97 |
-
"bitext_support": {
|
| 98 |
-
"heldout": false,
|
| 99 |
-
"acc": 0.998,
|
| 100 |
-
"nll": 0.016967715695500374,
|
| 101 |
-
"ece": 0.01042288645108536
|
| 102 |
-
},
|
| 103 |
-
"boolq": {
|
| 104 |
-
"heldout": false,
|
| 105 |
-
"acc": 0.9033333333333333,
|
| 106 |
-
"nll": 0.24676379561424255,
|
| 107 |
-
"ece": 0.029159295320510863
|
| 108 |
-
},
|
| 109 |
-
"cb": {
|
| 110 |
-
"heldout": true,
|
| 111 |
-
"acc": 0.8392857142857143,
|
| 112 |
-
"nll": 0.4322512745857239,
|
| 113 |
-
"ece": 0.11792815795966557
|
| 114 |
-
},
|
| 115 |
-
"civil_comments": {
|
| 116 |
-
"heldout": false,
|
| 117 |
-
"acc": 0.9328,
|
| 118 |
-
"nll": 0.1662478744983673,
|
| 119 |
-
"ece": 0.010289057048161809
|
| 120 |
-
},
|
| 121 |
-
"clinc_oos": {
|
| 122 |
-
"heldout": false,
|
| 123 |
-
"acc": 0.9826666666666667,
|
| 124 |
-
"nll": 0.08267337828874588,
|
| 125 |
-
"ece": 0.024905774275461832
|
| 126 |
-
},
|
| 127 |
-
"cola": {
|
| 128 |
-
"heldout": false,
|
| 129 |
-
"acc": 0.8331735378715245,
|
| 130 |
-
"nll": 0.3790046274662018,
|
| 131 |
-
"ece": 0.02507039284546103
|
| 132 |
-
},
|
| 133 |
-
"commonsense_qa": {
|
| 134 |
-
"heldout": false,
|
| 135 |
-
"acc": 0.823914823914824,
|
| 136 |
-
"nll": 0.4847744107246399,
|
| 137 |
-
"ece": 0.030845976944930447
|
| 138 |
-
},
|
| 139 |
-
"copa": {
|
| 140 |
-
"heldout": false,
|
| 141 |
-
"acc": 1.0,
|
| 142 |
-
"nll": 0.030766163021326065,
|
| 143 |
-
"ece": 0.028578150868415832
|
| 144 |
-
},
|
| 145 |
-
"counterfactual": {
|
| 146 |
-
"heldout": false,
|
| 147 |
-
"acc": 0.9573333333333334,
|
| 148 |
-
"nll": 0.13998979330062866,
|
| 149 |
-
"ece": 0.012234268069267251
|
| 150 |
-
},
|
| 151 |
-
"cr_reviews": {
|
| 152 |
-
"heldout": true,
|
| 153 |
-
"acc": 0.897742363877822,
|
| 154 |
-
"nll": 0.24005697667598724,
|
| 155 |
-
"ece": 0.045616757347289316
|
| 156 |
-
},
|
| 157 |
-
"dbpedia": {
|
| 158 |
-
"heldout": false,
|
| 159 |
-
"acc": 0.9933333333333333,
|
| 160 |
-
"nll": 0.027618318796157837,
|
| 161 |
-
"ece": 0.013420193096001943
|
| 162 |
-
},
|
| 163 |
-
"dbpedia_l2": {
|
| 164 |
-
"heldout": true,
|
| 165 |
-
"acc": 0.9506666666666667,
|
| 166 |
-
"nll": 0.1401711255311966,
|
| 167 |
-
"ece": 0.017278132518132504
|
| 168 |
-
},
|
| 169 |
-
"dbpedia_l3": {
|
| 170 |
-
"heldout": true,
|
| 171 |
-
"acc": 0.9913333333333333,
|
| 172 |
-
"nll": 0.03723839297890663,
|
| 173 |
-
"ece": 0.011898971815903936
|
| 174 |
-
},
|
| 175 |
-
"dolly_category": {
|
| 176 |
-
"heldout": true,
|
| 177 |
-
"acc": 0.374,
|
| 178 |
-
"nll": 1.7009543180465698,
|
| 179 |
-
"ece": 0.1419213018019994
|
| 180 |
-
},
|
| 181 |
-
"emotion": {
|
| 182 |
-
"heldout": false,
|
| 183 |
-
"acc": 0.8333333333333334,
|
| 184 |
-
"nll": 0.4790882170200348,
|
| 185 |
-
"ece": 0.025491972724596655
|
| 186 |
-
},
|
| 187 |
-
"enron_spam": {
|
| 188 |
-
"heldout": false,
|
| 189 |
-
"acc": 0.988,
|
| 190 |
-
"nll": 0.03346330299973488,
|
| 191 |
-
"ece": 0.006207752108573955
|
| 192 |
-
},
|
| 193 |
-
"fever": {
|
| 194 |
-
"heldout": false,
|
| 195 |
-
"acc": 0.8806666666666667,
|
| 196 |
-
"nll": 0.32262682914733887,
|
| 197 |
-
"ece": 0.01808753836154938
|
| 198 |
-
},
|
| 199 |
-
"fin_phrasebank": {
|
| 200 |
-
"heldout": true,
|
| 201 |
-
"acc": 0.6989690721649484,
|
| 202 |
-
"nll": 0.6234146952629089,
|
| 203 |
-
"ece": 0.07664692620026696
|
| 204 |
-
},
|
| 205 |
-
"fin_sentiment": {
|
| 206 |
-
"heldout": true,
|
| 207 |
-
"acc": 0.826,
|
| 208 |
-
"nll": 0.42305630445480347,
|
| 209 |
-
"ece": 0.051281696399052924
|
| 210 |
-
},
|
| 211 |
-
"glaive_tools": {
|
| 212 |
-
"heldout": false,
|
| 213 |
-
"acc": 0.9369436201780416,
|
| 214 |
-
"nll": 0.2514188289642334,
|
| 215 |
-
"ece": 0.10150587457047022
|
| 216 |
-
},
|
| 217 |
-
"go_emotions": {
|
| 218 |
-
"heldout": false,
|
| 219 |
-
"acc": 0.776,
|
| 220 |
-
"nll": 0.6853362917900085,
|
| 221 |
-
"ece": 0.03606782659888266
|
| 222 |
-
},
|
| 223 |
-
"hate_offensive": {
|
| 224 |
-
"heldout": false,
|
| 225 |
-
"acc": 0.9133333333333333,
|
| 226 |
-
"nll": 0.26428458094596863,
|
| 227 |
-
"ece": 0.014751521805922176
|
| 228 |
-
},
|
| 229 |
-
"hate_speech_scales": {
|
| 230 |
-
"heldout": false,
|
| 231 |
-
"acc": 0.5582666666666667,
|
| 232 |
-
"nll": 0.9945353269577026,
|
| 233 |
-
"ece": 0.03258252822558084
|
| 234 |
-
},
|
| 235 |
-
"hellaswag": {
|
| 236 |
-
"heldout": false,
|
| 237 |
-
"acc": 0.9346666666666666,
|
| 238 |
-
"nll": 0.1996963769197464,
|
| 239 |
-
"ece": 0.015657478451728855
|
| 240 |
-
},
|
| 241 |
-
"helpsteer2": {
|
| 242 |
-
"heldout": false,
|
| 243 |
-
"acc": 0.5980732177263969,
|
| 244 |
-
"nll": 0.9590670466423035,
|
| 245 |
-
"ece": 0.02256552288183126
|
| 246 |
-
},
|
| 247 |
-
"helpsteer3_pref": {
|
| 248 |
-
"heldout": false,
|
| 249 |
-
"acc": 0.38605442176870747,
|
| 250 |
-
"nll": 1.5417133569717407,
|
| 251 |
-
"ece": 0.061995073291314706
|
| 252 |
-
},
|
| 253 |
-
"hermes_tools": {
|
| 254 |
-
"heldout": true,
|
| 255 |
-
"acc": 0.7373333333333333,
|
| 256 |
-
"nll": 0.5897114872932434,
|
| 257 |
-
"ece": 0.1240579075018565
|
| 258 |
-
},
|
| 259 |
-
"hh_rlhf": {
|
| 260 |
-
"heldout": false,
|
| 261 |
-
"acc": 0.6451612903225806,
|
| 262 |
-
"nll": 0.680734395980835,
|
| 263 |
-
"ece": 0.11346792193349971
|
| 264 |
-
},
|
| 265 |
-
"hwu64": {
|
| 266 |
-
"heldout": true,
|
| 267 |
-
"acc": 0.9628252788104089,
|
| 268 |
-
"nll": 0.12855161726474762,
|
| 269 |
-
"ece": 0.03735405791093864
|
| 270 |
-
},
|
| 271 |
-
"imdb": {
|
| 272 |
-
"heldout": false,
|
| 273 |
-
"acc": 0.968,
|
| 274 |
-
"nll": 0.09611714631319046,
|
| 275 |
-
"ece": 0.01006534445285793
|
| 276 |
-
},
|
| 277 |
-
"insincere_questions": {
|
| 278 |
-
"heldout": false,
|
| 279 |
-
"acc": 0.9553333333333334,
|
| 280 |
-
"nll": 0.12364762276411057,
|
| 281 |
-
"ece": 0.01763951281706497
|
| 282 |
-
},
|
| 283 |
-
"liar2": {
|
| 284 |
-
"heldout": false,
|
| 285 |
-
"acc": 0.35,
|
| 286 |
-
"nll": 1.513845682144165,
|
| 287 |
-
"ece": 0.06293645719687144
|
| 288 |
-
},
|
| 289 |
-
"massive_intent": {
|
| 290 |
-
"heldout": false,
|
| 291 |
-
"acc": 0.9606666666666667,
|
| 292 |
-
"nll": 0.13908076286315918,
|
| 293 |
-
"ece": 0.02024710043271384
|
| 294 |
-
},
|
| 295 |
-
"massive_scenario": {
|
| 296 |
-
"heldout": true,
|
| 297 |
-
"acc": 0.7933333333333333,
|
| 298 |
-
"nll": 0.5642920732498169,
|
| 299 |
-
"ece": 0.05245566362142563
|
| 300 |
-
},
|
| 301 |
-
"medmcqa": {
|
| 302 |
-
"heldout": false,
|
| 303 |
-
"acc": 0.6233333333333333,
|
| 304 |
-
"nll": 0.9161791801452637,
|
| 305 |
-
"ece": 0.07326398479938506
|
| 306 |
-
},
|
| 307 |
-
"medqa": {
|
| 308 |
-
"heldout": false,
|
| 309 |
-
"acc": 0.6779261586802828,
|
| 310 |
-
"nll": 0.8010168075561523,
|
| 311 |
-
"ece": 0.0796375350744258
|
| 312 |
-
},
|
| 313 |
-
"mind2web": {
|
| 314 |
-
"heldout": false,
|
| 315 |
-
"acc": 0.822,
|
| 316 |
-
"nll": 0.48613396286964417,
|
| 317 |
-
"ece": 0.04646803086996078
|
| 318 |
-
},
|
| 319 |
-
"mmlu": {
|
| 320 |
-
"heldout": false,
|
| 321 |
-
"acc": 0.742,
|
| 322 |
-
"nll": 0.6732416749000549,
|
| 323 |
-
"ece": 0.04000420778989793
|
| 324 |
-
},
|
| 325 |
-
"mnli": {
|
| 326 |
-
"heldout": false,
|
| 327 |
-
"acc": 0.9026666666666666,
|
| 328 |
-
"nll": 0.28677940368652344,
|
| 329 |
-
"ece": 0.019106337686379753
|
| 330 |
-
},
|
| 331 |
-
"mrpc": {
|
| 332 |
-
"heldout": false,
|
| 333 |
-
"acc": 0.8431372549019608,
|
| 334 |
-
"nll": 0.35424306988716125,
|
| 335 |
-
"ece": 0.031691053334404445
|
| 336 |
-
},
|
| 337 |
-
"msmarco_rel": {
|
| 338 |
-
"heldout": false,
|
| 339 |
-
"acc": 0.6874135546334716,
|
| 340 |
-
"nll": 0.586560845375061,
|
| 341 |
-
"ece": 0.035539872212693564
|
| 342 |
-
},
|
| 343 |
-
"multirc": {
|
| 344 |
-
"heldout": false,
|
| 345 |
-
"acc": 0.9053333333333333,
|
| 346 |
-
"nll": 0.2743209898471832,
|
| 347 |
-
"ece": 0.025939866781234736
|
| 348 |
-
},
|
| 349 |
-
"newsgroups": {
|
| 350 |
-
"heldout": false,
|
| 351 |
-
"acc": 0.820109439124487,
|
| 352 |
-
"nll": 0.5166534781455994,
|
| 353 |
-
"ece": 0.0274887290539885
|
| 354 |
-
},
|
| 355 |
-
"offtopic_probe": {
|
| 356 |
-
"heldout": true,
|
| 357 |
-
"acc": 0.8234567901234567,
|
| 358 |
-
"nll": 0.5037797093391418,
|
| 359 |
-
"ece": 0.043301358119941055
|
| 360 |
-
},
|
| 361 |
-
"openbookqa": {
|
| 362 |
-
"heldout": false,
|
| 363 |
-
"acc": 0.904,
|
| 364 |
-
"nll": 0.2716519832611084,
|
| 365 |
-
"ece": 0.033541144907474535
|
| 366 |
-
},
|
| 367 |
-
"paws": {
|
| 368 |
-
"heldout": true,
|
| 369 |
-
"acc": 0.794,
|
| 370 |
-
"nll": 0.5527914762496948,
|
| 371 |
-
"ece": 0.12406627035140992
|
| 372 |
-
},
|
| 373 |
-
"piqa": {
|
| 374 |
-
"heldout": false,
|
| 375 |
-
"acc": 0.89,
|
| 376 |
-
"nll": 0.2665766477584839,
|
| 377 |
-
"ece": 0.026692868391672748
|
| 378 |
-
},
|
| 379 |
-
"prosocial_safety": {
|
| 380 |
-
"heldout": false,
|
| 381 |
-
"acc": 0.512,
|
| 382 |
-
"nll": 1.1915110349655151,
|
| 383 |
-
"ece": 0.030761619885762537
|
| 384 |
-
},
|
| 385 |
-
"pubmedqa": {
|
| 386 |
-
"heldout": true,
|
| 387 |
-
"acc": 0.758,
|
| 388 |
-
"nll": 0.6072925329208374,
|
| 389 |
-
"ece": 0.05524836438894271
|
| 390 |
-
},
|
| 391 |
-
"qasc": {
|
| 392 |
-
"heldout": false,
|
| 393 |
-
"acc": 0.8596112311015118,
|
| 394 |
-
"nll": 0.47038528323173523,
|
| 395 |
-
"ece": 0.06783940297094844
|
| 396 |
-
},
|
| 397 |
-
"qnli": {
|
| 398 |
-
"heldout": false,
|
| 399 |
-
"acc": 0.946,
|
| 400 |
-
"nll": 0.14609812200069427,
|
| 401 |
-
"ece": 0.01228535489241284
|
| 402 |
-
},
|
| 403 |
-
"qqp": {
|
| 404 |
-
"heldout": false,
|
| 405 |
-
"acc": 0.848,
|
| 406 |
-
"nll": 0.3371011018753052,
|
| 407 |
-
"ece": 0.03070641124248504
|
| 408 |
-
},
|
| 409 |
-
"quality": {
|
| 410 |
-
"heldout": true,
|
| 411 |
-
"acc": 0.5606666666666666,
|
| 412 |
-
"nll": 1.1722065210342407,
|
| 413 |
-
"ece": 0.17895789907375972
|
| 414 |
-
},
|
| 415 |
-
"quality_full": {
|
| 416 |
-
"heldout": true,
|
| 417 |
-
"acc": 0.505,
|
| 418 |
-
"nll": 1.2171955108642578,
|
| 419 |
-
"ece": 0.1881948218246301
|
| 420 |
-
},
|
| 421 |
-
"race": {
|
| 422 |
-
"heldout": false,
|
| 423 |
-
"acc": 0.8953333333333333,
|
| 424 |
-
"nll": 0.33984363079071045,
|
| 425 |
-
"ece": 0.02000871032476428
|
| 426 |
-
},
|
| 427 |
-
"reward_bench": {
|
| 428 |
-
"heldout": true,
|
| 429 |
-
"acc": 0.8526666666666667,
|
| 430 |
-
"nll": 0.34169474244117737,
|
| 431 |
-
"ece": 0.03198828514417014
|
| 432 |
-
},
|
| 433 |
-
"rte": {
|
| 434 |
-
"heldout": false,
|
| 435 |
-
"acc": 0.9205776173285198,
|
| 436 |
-
"nll": 0.1940610110759735,
|
| 437 |
-
"ece": 0.04438359793342838
|
| 438 |
-
},
|
| 439 |
-
"sciq": {
|
| 440 |
-
"heldout": true,
|
| 441 |
-
"acc": 0.991,
|
| 442 |
-
"nll": 0.044908951967954636,
|
| 443 |
-
"ece": 0.02056253519654279
|
| 444 |
-
},
|
| 445 |
-
"shp": {
|
| 446 |
-
"heldout": false,
|
| 447 |
-
"acc": 0.7353333333333333,
|
| 448 |
-
"nll": 0.594969630241394,
|
| 449 |
-
"ece": 0.1105338017543157
|
| 450 |
-
},
|
| 451 |
-
"sms_spam": {
|
| 452 |
-
"heldout": false,
|
| 453 |
-
"acc": 0.987,
|
| 454 |
-
"nll": 0.055126290768384933,
|
| 455 |
-
"ece": 0.006505105614662178
|
| 456 |
-
},
|
| 457 |
-
"snli": {
|
| 458 |
-
"heldout": false,
|
| 459 |
-
"acc": 0.9227642276422764,
|
| 460 |
-
"nll": 0.24105508625507355,
|
| 461 |
-
"ece": 0.03949038551105716
|
| 462 |
-
},
|
| 463 |
-
"social_iqa": {
|
| 464 |
-
"heldout": true,
|
| 465 |
-
"acc": 0.7593333333333333,
|
| 466 |
-
"nll": 0.5866184830665588,
|
| 467 |
-
"ece": 0.06180198764801027
|
| 468 |
-
},
|
| 469 |
-
"sst2": {
|
| 470 |
-
"heldout": false,
|
| 471 |
-
"acc": 0.948394495412844,
|
| 472 |
-
"nll": 0.14050182700157166,
|
| 473 |
-
"ece": 0.01752297238472409
|
| 474 |
-
},
|
| 475 |
-
"sst5": {
|
| 476 |
-
"heldout": false,
|
| 477 |
-
"acc": 0.5826666666666667,
|
| 478 |
-
"nll": 0.9461587071418762,
|
| 479 |
-
"ece": 0.0451480305393537
|
| 480 |
-
},
|
| 481 |
-
"strategyqa": {
|
| 482 |
-
"heldout": true,
|
| 483 |
-
"acc": 0.6317321688500728,
|
| 484 |
-
"nll": 0.6970252990722656,
|
| 485 |
-
"ece": 0.1567338213129335
|
| 486 |
-
},
|
| 487 |
-
"student_questions": {
|
| 488 |
-
"heldout": true,
|
| 489 |
-
"acc": 0.94,
|
| 490 |
-
"nll": 0.20459797978401184,
|
| 491 |
-
"ece": 0.05185317244132365
|
| 492 |
-
},
|
| 493 |
-
"subjectivity": {
|
| 494 |
-
"heldout": false,
|
| 495 |
-
"acc": 0.9566666666666667,
|
| 496 |
-
"nll": 0.11102409660816193,
|
| 497 |
-
"ece": 0.015209745287895248
|
| 498 |
-
},
|
| 499 |
-
"support_tickets": {
|
| 500 |
-
"heldout": false,
|
| 501 |
-
"acc": 0.5586666666666666,
|
| 502 |
-
"nll": 1.0297108888626099,
|
| 503 |
-
"ece": 0.024460139827595817
|
| 504 |
-
},
|
| 505 |
-
"synth": {
|
| 506 |
-
"heldout": false,
|
| 507 |
-
"acc": 0.8066666666666666,
|
| 508 |
-
"nll": 0.4917465150356293,
|
| 509 |
-
"ece": 0.04993985931078593
|
| 510 |
-
},
|
| 511 |
-
"toolace": {
|
| 512 |
-
"heldout": false,
|
| 513 |
-
"acc": 0.931,
|
| 514 |
-
"nll": 0.2593076229095459,
|
| 515 |
-
"ece": 0.06617412987351419
|
| 516 |
-
},
|
| 517 |
-
"toxic_chat": {
|
| 518 |
-
"heldout": false,
|
| 519 |
-
"acc": 0.977,
|
| 520 |
-
"nll": 0.06245163083076477,
|
| 521 |
-
"ece": 0.00981356755892436
|
| 522 |
-
},
|
| 523 |
-
"trec": {
|
| 524 |
-
"heldout": true,
|
| 525 |
-
"acc": 0.842,
|
| 526 |
-
"nll": 0.5022116899490356,
|
| 527 |
-
"ece": 0.02651941215991974
|
| 528 |
-
},
|
| 529 |
-
"truthfulqa": {
|
| 530 |
-
"heldout": true,
|
| 531 |
-
"acc": 0.653610771113831,
|
| 532 |
-
"nll": 1.0122824907302856,
|
| 533 |
-
"ece": 0.06321512999849775
|
| 534 |
-
},
|
| 535 |
-
"tweet_emotion": {
|
| 536 |
-
"heldout": false,
|
| 537 |
-
"acc": 0.8578465869106263,
|
| 538 |
-
"nll": 0.41329097747802734,
|
| 539 |
-
"ece": 0.039397706623063994
|
| 540 |
-
},
|
| 541 |
-
"tweet_hate": {
|
| 542 |
-
"heldout": false,
|
| 543 |
-
"acc": 0.5126666666666667,
|
| 544 |
-
"nll": 1.136211633682251,
|
| 545 |
-
"ece": 0.34860085622469583
|
| 546 |
-
},
|
| 547 |
-
"tweet_irony": {
|
| 548 |
-
"heldout": true,
|
| 549 |
-
"acc": 0.826530612244898,
|
| 550 |
-
"nll": 0.39445212483406067,
|
| 551 |
-
"ece": 0.027230415189144572
|
| 552 |
-
},
|
| 553 |
-
"tweet_offensive": {
|
| 554 |
-
"heldout": false,
|
| 555 |
-
"acc": 0.8523255813953489,
|
| 556 |
-
"nll": 0.3407946228981018,
|
| 557 |
-
"ece": 0.021613172181817006
|
| 558 |
-
},
|
| 559 |
-
"tweet_sentiment": {
|
| 560 |
-
"heldout": false,
|
| 561 |
-
"acc": 0.7206666666666667,
|
| 562 |
-
"nll": 0.610216498374939,
|
| 563 |
-
"ece": 0.028348857482274364
|
| 564 |
-
},
|
| 565 |
-
"ultrafeedback_pref": {
|
| 566 |
-
"heldout": false,
|
| 567 |
-
"acc": 0.7713333333333333,
|
| 568 |
-
"nll": 0.4769253432750702,
|
| 569 |
-
"ece": 0.06639519715309146
|
| 570 |
-
},
|
| 571 |
-
"wic": {
|
| 572 |
-
"heldout": false,
|
| 573 |
-
"acc": 0.7429467084639498,
|
| 574 |
-
"nll": 0.5646101832389832,
|
| 575 |
-
"ece": 0.09329287749846528
|
| 576 |
-
},
|
| 577 |
-
"wiki_qa": {
|
| 578 |
-
"heldout": false,
|
| 579 |
-
"acc": 0.9089874857792947,
|
| 580 |
-
"nll": 0.2449624091386795,
|
| 581 |
-
"ece": 0.020241137725365718
|
| 582 |
-
},
|
| 583 |
-
"winogrande": {
|
| 584 |
-
"heldout": false,
|
| 585 |
-
"acc": 0.8089976322020521,
|
| 586 |
-
"nll": 0.48967406153678894,
|
| 587 |
-
"ece": 0.08959193594736103
|
| 588 |
-
},
|
| 589 |
-
"xstory_cloze": {
|
| 590 |
-
"heldout": true,
|
| 591 |
-
"acc": 0.9706666666666667,
|
| 592 |
-
"nll": 0.07805304229259491,
|
| 593 |
-
"ece": 0.01724668137232463
|
| 594 |
-
},
|
| 595 |
-
"yahoo_topics": {
|
| 596 |
-
"heldout": false,
|
| 597 |
-
"acc": 0.7613333333333333,
|
| 598 |
-
"nll": 0.7648204565048218,
|
| 599 |
-
"ece": 0.0657656339406967
|
| 600 |
-
},
|
| 601 |
-
"yelp": {
|
| 602 |
-
"heldout": false,
|
| 603 |
-
"acc": 0.7033333333333334,
|
| 604 |
-
"nll": 0.7055407166481018,
|
| 605 |
-
"ece": 0.04224825153748195
|
| 606 |
}
|
| 607 |
},
|
| 608 |
-
"
|
| 609 |
-
"
|
| 610 |
-
"
|
| 611 |
-
|
| 612 |
-
|
| 613 |
-
|
| 614 |
-
|
| 615 |
-
|
| 616 |
-
|
| 617 |
-
|
| 618 |
-
|
| 619 |
-
|
| 620 |
-
|
| 621 |
-
|
| 622 |
-
|
| 623 |
-
|
| 624 |
-
|
| 625 |
-
|
| 626 |
-
|
| 627 |
-
|
| 628 |
-
|
| 629 |
-
|
| 630 |
-
|
| 631 |
-
|
| 632 |
-
|
| 633 |
-
|
| 634 |
-
|
| 635 |
-
|
| 636 |
-
|
| 637 |
-
|
| 638 |
-
|
| 639 |
-
|
| 640 |
-
|
| 641 |
-
"hard_conf_ge_0_9_coverage": 0.3063,
|
| 642 |
-
"hard_conf_ge_0_9_accuracy": 0.9412
|
| 643 |
-
},
|
| 644 |
-
"bespoke_public_suite": {
|
| 645 |
-
"macro": 0.7725,
|
| 646 |
-
"micro": 0.7809,
|
| 647 |
-
"macro_untrained": 0.7514,
|
| 648 |
-
"per_subset": {
|
| 649 |
-
"vitaminc-dev": 0.778,
|
| 650 |
-
"massive-en-US": 0.8686,
|
| 651 |
-
"massive-de-DE": 0.8371,
|
| 652 |
-
"boolq": 0.86,
|
| 653 |
-
"squad2": 0.7926,
|
| 654 |
-
"paws": 0.832,
|
| 655 |
-
"multinli": 0.9264,
|
| 656 |
-
"civil_comments": 0.8567,
|
| 657 |
-
"aegis2": 0.812,
|
| 658 |
-
"helpsteer2": 0.4659,
|
| 659 |
-
"summeval-relevance": 0.4625,
|
| 660 |
-
"summeval-consistency": 0.8264,
|
| 661 |
-
"pubmedqa": 0.724
|
| 662 |
-
}
|
| 663 |
-
},
|
| 664 |
-
"zero_shot_games_win_rate": {
|
| 665 |
-
"sampled_all": 0.2244,
|
| 666 |
-
"greedy_all": 0.2778,
|
| 667 |
-
"sampled_bags": 0.3789,
|
| 668 |
-
"greedy_bags": 0.4844
|
| 669 |
-
},
|
| 670 |
-
"text_games_zero_shot_5_episodes": {
|
| 671 |
-
"pong": {
|
| 672 |
-
"model": -21.0,
|
| 673 |
-
"teacher": 8.0,
|
| 674 |
-
"random": -20.4
|
| 675 |
-
},
|
| 676 |
-
"freeway": {
|
| 677 |
-
"model": 1.0,
|
| 678 |
-
"teacher": 5.0,
|
| 679 |
-
"random": 0.0
|
| 680 |
-
},
|
| 681 |
-
"breakout": {
|
| 682 |
-
"model": 12.0,
|
| 683 |
-
"teacher": 22.0,
|
| 684 |
-
"random": 0.8
|
| 685 |
},
|
| 686 |
-
"
|
| 687 |
-
"
|
| 688 |
-
|
| 689 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 690 |
},
|
| 691 |
-
"
|
| 692 |
-
"
|
| 693 |
-
|
| 694 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 695 |
},
|
| 696 |
-
"
|
| 697 |
-
"
|
| 698 |
-
|
| 699 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 700 |
},
|
| 701 |
-
"
|
| 702 |
-
"
|
| 703 |
-
|
| 704 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 705 |
},
|
| 706 |
-
"
|
| 707 |
-
"
|
| 708 |
-
|
| 709 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 710 |
},
|
| 711 |
-
"
|
| 712 |
-
"
|
| 713 |
-
|
| 714 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 715 |
},
|
| 716 |
-
"
|
| 717 |
-
"
|
| 718 |
-
|
| 719 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 720 |
}
|
| 721 |
},
|
| 722 |
-
"
|
| 723 |
-
"
|
| 724 |
-
|
| 725 |
-
|
| 726 |
-
|
| 727 |
-
|
| 728 |
-
|
| 729 |
-
|
| 730 |
-
|
| 731 |
-
|
| 732 |
-
|
| 733 |
-
|
| 734 |
-
|
| 735 |
-
|
| 736 |
-
|
| 737 |
-
|
| 738 |
-
|
| 739 |
-
|
| 740 |
-
|
| 741 |
-
|
| 742 |
-
|
| 743 |
-
|
| 744 |
-
|
| 745 |
-
|
| 746 |
-
|
| 747 |
-
|
| 748 |
-
|
| 749 |
-
|
| 750 |
-
|
| 751 |
-
|
| 752 |
-
|
| 753 |
-
"isolated_minus_listwise_max_abs_points": 2.0,
|
| 754 |
-
"per_level_fit_sum_range": [
|
| 755 |
-
0.92,
|
| 756 |
-
1.0
|
| 757 |
-
],
|
| 758 |
-
"listwise_ece_five_rating_sets_T1": [
|
| 759 |
-
0.127,
|
| 760 |
-
0.265
|
| 761 |
-
]
|
| 762 |
-
},
|
| 763 |
-
"speed": {
|
| 764 |
-
"note": "same weights, measured on 2026-09-24 in the candidate readout session (server config T 1.719; the temperature does not change speed); one unshared B300, bf16",
|
| 765 |
-
"eager_game_state_median_ms": 34.6,
|
| 766 |
-
"v1_same_session_ms": 32.4,
|
| 767 |
-
"graphs_compile_single_request_ms": 5.2,
|
| 768 |
-
"fp8_single_request_ms": 5.0,
|
| 769 |
-
"batch32_decisions_per_s": 1178,
|
| 770 |
-
"batch32_fp8_decisions_per_s": 1349,
|
| 771 |
-
"http_decide_1_client_req_per_s": 72.8,
|
| 772 |
-
"http_decide_1_client_p50_ms": 13.3,
|
| 773 |
-
"http_decide_64_clients_req_per_s": 189.7
|
| 774 |
-
}
|
| 775 |
-
},
|
| 776 |
-
"paired": {
|
| 777 |
-
"regression_set_per_task_bootstrap": {
|
| 778 |
-
"v2-v1 T_1.05|in_task": {
|
| 779 |
-
"acc": {
|
| 780 |
-
"tasks": 67,
|
| 781 |
-
"delta": -0.00963326990956669,
|
| 782 |
-
"ci95": [
|
| 783 |
-
-0.012258285462742567,
|
| 784 |
-
-0.007113672993330958
|
| 785 |
-
],
|
| 786 |
-
"higher": 6,
|
| 787 |
-
"lower": 54,
|
| 788 |
-
"equal": 7
|
| 789 |
-
},
|
| 790 |
-
"nll": {
|
| 791 |
-
"tasks": 67,
|
| 792 |
-
"delta": 0.0367151916395428,
|
| 793 |
-
"ci95": [
|
| 794 |
-
0.02853969128380196,
|
| 795 |
-
0.04618769902040932
|
| 796 |
-
],
|
| 797 |
-
"higher": 64,
|
| 798 |
-
"lower": 3,
|
| 799 |
-
"equal": 0
|
| 800 |
-
},
|
| 801 |
-
"ece": {
|
| 802 |
-
"tasks": 67,
|
| 803 |
-
"delta": 0.014469173095775816,
|
| 804 |
-
"ci95": [
|
| 805 |
-
0.008371143739489972,
|
| 806 |
-
0.021221732726272202
|
| 807 |
-
],
|
| 808 |
-
"higher": 49,
|
| 809 |
-
"lower": 18,
|
| 810 |
-
"equal": 0
|
| 811 |
}
|
| 812 |
},
|
| 813 |
-
"
|
| 814 |
-
"
|
| 815 |
-
"
|
| 816 |
-
"
|
| 817 |
-
"
|
| 818 |
-
|
| 819 |
-
|
| 820 |
-
|
| 821 |
-
|
| 822 |
-
"
|
| 823 |
-
|
| 824 |
-
|
| 825 |
-
|
| 826 |
-
|
| 827 |
-
|
| 828 |
-
|
| 829 |
-
|
| 830 |
-
|
| 831 |
-
|
| 832 |
-
|
| 833 |
-
|
| 834 |
-
|
| 835 |
-
|
| 836 |
-
|
| 837 |
-
|
| 838 |
-
|
| 839 |
-
|
| 840 |
-
|
| 841 |
-
0.
|
| 842 |
-
|
| 843 |
-
"higher": 20,
|
| 844 |
-
"lower": 8,
|
| 845 |
-
"equal": 0
|
| 846 |
}
|
| 847 |
},
|
| 848 |
-
"
|
| 849 |
-
"
|
| 850 |
-
"
|
| 851 |
-
"
|
| 852 |
-
"
|
| 853 |
-
|
| 854 |
-
|
| 855 |
-
|
| 856 |
-
|
| 857 |
-
"
|
| 858 |
-
|
| 859 |
-
|
| 860 |
-
|
| 861 |
-
|
| 862 |
-
|
| 863 |
-
|
| 864 |
-
|
| 865 |
-
|
| 866 |
-
],
|
| 867 |
-
"higher": 26,
|
| 868 |
-
"lower": 41,
|
| 869 |
-
"equal": 0
|
| 870 |
-
},
|
| 871 |
-
"ece": {
|
| 872 |
-
"tasks": 67,
|
| 873 |
-
"delta": 0.004320193883140634,
|
| 874 |
-
"ci95": [
|
| 875 |
-
-0.0023582249916685393,
|
| 876 |
-
0.011625304997016925
|
| 877 |
-
],
|
| 878 |
-
"higher": 34,
|
| 879 |
-
"lower": 33,
|
| 880 |
-
"equal": 0
|
| 881 |
}
|
| 882 |
},
|
| 883 |
-
"
|
| 884 |
-
"
|
| 885 |
-
"
|
| 886 |
-
"
|
| 887 |
-
"
|
| 888 |
-
|
| 889 |
-
|
| 890 |
-
|
| 891 |
-
|
| 892 |
-
"
|
| 893 |
-
|
| 894 |
-
|
| 895 |
-
|
| 896 |
-
|
| 897 |
-
|
| 898 |
-
|
| 899 |
-
|
| 900 |
-
|
| 901 |
-
|
| 902 |
-
|
| 903 |
-
|
| 904 |
-
|
| 905 |
-
|
| 906 |
-
|
| 907 |
-
|
| 908 |
-
|
| 909 |
-
|
| 910 |
-
|
| 911 |
-
0.
|
| 912 |
-
|
| 913 |
-
"higher": 13,
|
| 914 |
-
"lower": 15,
|
| 915 |
-
"equal": 0
|
| 916 |
}
|
| 917 |
},
|
| 918 |
-
"
|
| 919 |
-
"
|
| 920 |
-
"
|
| 921 |
-
"
|
| 922 |
-
"
|
| 923 |
-
|
| 924 |
-
|
| 925 |
-
|
| 926 |
-
|
| 927 |
-
"
|
| 928 |
-
|
| 929 |
-
|
| 930 |
-
|
| 931 |
-
|
| 932 |
-
|
| 933 |
-
|
| 934 |
-
|
| 935 |
-
|
| 936 |
-
|
| 937 |
-
|
| 938 |
-
|
| 939 |
-
|
| 940 |
-
|
| 941 |
-
|
| 942 |
-
|
| 943 |
-
|
| 944 |
-
|
| 945 |
-
0.
|
| 946 |
-
0.
|
| 947 |
-
|
| 948 |
-
"higher": 55,
|
| 949 |
-
"lower": 12,
|
| 950 |
-
"equal": 0
|
| 951 |
}
|
| 952 |
},
|
| 953 |
-
"
|
| 954 |
-
"
|
| 955 |
-
"
|
| 956 |
-
"
|
| 957 |
-
"
|
| 958 |
-
|
| 959 |
-
|
| 960 |
-
|
| 961 |
-
|
| 962 |
-
"
|
| 963 |
-
|
| 964 |
-
|
| 965 |
-
|
| 966 |
-
|
| 967 |
-
|
| 968 |
-
|
| 969 |
-
|
| 970 |
-
|
| 971 |
-
|
| 972 |
-
|
| 973 |
-
|
| 974 |
-
|
| 975 |
-
|
| 976 |
-
|
| 977 |
-
|
| 978 |
-
|
| 979 |
-
|
| 980 |
-
|
| 981 |
-
0.
|
| 982 |
-
|
| 983 |
-
"higher": 20,
|
| 984 |
-
"lower": 8,
|
| 985 |
-
"equal": 0
|
| 986 |
-
}
|
| 987 |
-
}
|
| 988 |
-
},
|
| 989 |
-
"fixtures_row_bootstrap_minus_v1": {
|
| 990 |
-
"general_validation": {
|
| 991 |
-
"accuracy": {
|
| 992 |
-
"delta": -0.010625737898465172,
|
| 993 |
-
"ci95": [
|
| 994 |
-
-0.025974025974025976,
|
| 995 |
-
0.004722550177095631
|
| 996 |
-
]
|
| 997 |
-
},
|
| 998 |
-
"nll": {
|
| 999 |
-
"delta": 0.0015418035959596215,
|
| 1000 |
-
"ci95": [
|
| 1001 |
-
-0.018830652980342588,
|
| 1002 |
-
0.02145239419471463
|
| 1003 |
-
]
|
| 1004 |
}
|
| 1005 |
},
|
| 1006 |
-
"
|
| 1007 |
-
"
|
| 1008 |
-
"
|
| 1009 |
-
"
|
| 1010 |
-
|
| 1011 |
-
|
| 1012 |
-
|
| 1013 |
-
},
|
| 1014 |
-
"
|
| 1015 |
-
"
|
| 1016 |
-
|
| 1017 |
-
|
| 1018 |
-
|
| 1019 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1020 |
}
|
| 1021 |
},
|
| 1022 |
-
"
|
| 1023 |
-
"
|
| 1024 |
-
"
|
| 1025 |
-
"
|
| 1026 |
-
|
| 1027 |
-
|
| 1028 |
-
|
| 1029 |
-
},
|
| 1030 |
-
"
|
| 1031 |
-
"
|
| 1032 |
-
|
| 1033 |
-
|
| 1034 |
-
|
| 1035 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1036 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1037 |
},
|
| 1038 |
-
"
|
| 1039 |
-
"
|
| 1040 |
-
"
|
| 1041 |
-
|
| 1042 |
-
|
| 1043 |
-
|
| 1044 |
-
|
| 1045 |
-
|
| 1046 |
-
|
| 1047 |
-
"
|
| 1048 |
-
|
| 1049 |
-
|
| 1050 |
-
|
| 1051 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1052 |
}
|
| 1053 |
}
|
| 1054 |
},
|
| 1055 |
-
"
|
| 1056 |
-
"
|
| 1057 |
-
"
|
| 1058 |
-
|
| 1059 |
-
|
| 1060 |
-
|
| 1061 |
-
|
| 1062 |
-
|
| 1063 |
-
|
| 1064 |
-
|
| 1065 |
-
|
| 1066 |
-
},
|
| 1067 |
-
"v2-v1|sampled|bag_game": {
|
| 1068 |
-
"n": 64,
|
| 1069 |
-
"a": 0.37890625,
|
| 1070 |
-
"b": 0.56640625,
|
| 1071 |
-
"delta": -0.1875,
|
| 1072 |
-
"ci95": [
|
| 1073 |
-
-0.24609375,
|
| 1074 |
-
-0.12890625
|
| 1075 |
-
]
|
| 1076 |
-
},
|
| 1077 |
-
"v2-v1|sampled|stochastic_grid": {
|
| 1078 |
-
"n": 64,
|
| 1079 |
-
"a": 0.16015625,
|
| 1080 |
-
"b": 0.1640625,
|
| 1081 |
-
"delta": -0.00390625,
|
| 1082 |
-
"ci95": [
|
| 1083 |
-
-0.02734375,
|
| 1084 |
-
0.01953125
|
| 1085 |
-
]
|
| 1086 |
-
},
|
| 1087 |
-
"v2-v1|sampled|tic_tac_toe": {
|
| 1088 |
-
"n": 74,
|
| 1089 |
-
"a": 0.24324324324324326,
|
| 1090 |
-
"b": 0.24324324324324326,
|
| 1091 |
-
"delta": 0.0,
|
| 1092 |
-
"ci95": [
|
| 1093 |
-
-0.02702702702702703,
|
| 1094 |
-
0.02702702702702703
|
| 1095 |
-
]
|
| 1096 |
-
},
|
| 1097 |
-
"v2-v1|sampled|exact_minesweeper": {
|
| 1098 |
-
"n": 32,
|
| 1099 |
-
"a": 0.0,
|
| 1100 |
-
"b": 0.0078125,
|
| 1101 |
-
"delta": -0.0078125,
|
| 1102 |
-
"ci95": [
|
| 1103 |
-
-0.0234375,
|
| 1104 |
-
0.0
|
| 1105 |
-
]
|
| 1106 |
-
},
|
| 1107 |
-
"v2-v1|greedy|all": {
|
| 1108 |
-
"n": 234,
|
| 1109 |
-
"a": 0.2777777777777778,
|
| 1110 |
-
"b": 0.2905982905982906,
|
| 1111 |
-
"delta": -0.01282051282051282,
|
| 1112 |
-
"ci95": [
|
| 1113 |
-
-0.05555555555555555,
|
| 1114 |
-
0.02564102564102564
|
| 1115 |
-
]
|
| 1116 |
-
},
|
| 1117 |
-
"v2-v1|greedy|bag_game": {
|
| 1118 |
-
"n": 64,
|
| 1119 |
-
"a": 0.484375,
|
| 1120 |
-
"b": 0.625,
|
| 1121 |
-
"delta": -0.140625,
|
| 1122 |
-
"ci95": [
|
| 1123 |
-
-0.234375,
|
| 1124 |
-
-0.0625
|
| 1125 |
-
]
|
| 1126 |
-
},
|
| 1127 |
-
"v2-v1|greedy|stochastic_grid": {
|
| 1128 |
-
"n": 64,
|
| 1129 |
-
"a": 0.1875,
|
| 1130 |
-
"b": 0.125,
|
| 1131 |
-
"delta": 0.0625,
|
| 1132 |
-
"ci95": [
|
| 1133 |
-
0.0,
|
| 1134 |
-
0.140625
|
| 1135 |
-
]
|
| 1136 |
-
},
|
| 1137 |
-
"v2-v1|greedy|tic_tac_toe": {
|
| 1138 |
-
"n": 74,
|
| 1139 |
-
"a": 0.2972972972972973,
|
| 1140 |
-
"b": 0.2702702702702703,
|
| 1141 |
-
"delta": 0.02702702702702703,
|
| 1142 |
-
"ci95": [
|
| 1143 |
-
-0.04054054054054054,
|
| 1144 |
-
0.0945945945945946
|
| 1145 |
-
]
|
| 1146 |
-
},
|
| 1147 |
-
"v2-v1|greedy|exact_minesweeper": {
|
| 1148 |
-
"n": 32,
|
| 1149 |
-
"a": 0.0,
|
| 1150 |
-
"b": 0.0,
|
| 1151 |
-
"delta": 0.0,
|
| 1152 |
-
"ci95": [
|
| 1153 |
-
0.0,
|
| 1154 |
-
0.0
|
| 1155 |
-
]
|
| 1156 |
-
},
|
| 1157 |
-
"v2-v2_T1719|sampled|all": {
|
| 1158 |
-
"n": 234,
|
| 1159 |
-
"a": 0.22435897435897437,
|
| 1160 |
-
"b": 0.23397435897435898,
|
| 1161 |
-
"delta": -0.009615384615384616,
|
| 1162 |
-
"ci95": [
|
| 1163 |
-
-0.021367521367521368,
|
| 1164 |
-
0.0010683760683760685
|
| 1165 |
-
]
|
| 1166 |
-
},
|
| 1167 |
-
"v2-v2_T1719|sampled|bag_game": {
|
| 1168 |
-
"n": 64,
|
| 1169 |
-
"a": 0.37890625,
|
| 1170 |
-
"b": 0.390625,
|
| 1171 |
-
"delta": -0.01171875,
|
| 1172 |
-
"ci95": [
|
| 1173 |
-
-0.04296875,
|
| 1174 |
-
0.01953125
|
| 1175 |
-
]
|
| 1176 |
-
},
|
| 1177 |
-
"v2-v2_T1719|sampled|stochastic_grid": {
|
| 1178 |
-
"n": 64,
|
| 1179 |
-
"a": 0.16015625,
|
| 1180 |
-
"b": 0.1796875,
|
| 1181 |
-
"delta": -0.01953125,
|
| 1182 |
-
"ci95": [
|
| 1183 |
-
-0.04296875,
|
| 1184 |
-
0.0
|
| 1185 |
-
]
|
| 1186 |
-
},
|
| 1187 |
-
"v2-v2_T1719|sampled|tic_tac_toe": {
|
| 1188 |
-
"n": 74,
|
| 1189 |
-
"a": 0.24324324324324326,
|
| 1190 |
-
"b": 0.24662162162162163,
|
| 1191 |
-
"delta": -0.0033783783783783786,
|
| 1192 |
-
"ci95": [
|
| 1193 |
-
-0.016891891891891893,
|
| 1194 |
-
0.006756756756756757
|
| 1195 |
-
]
|
| 1196 |
-
},
|
| 1197 |
-
"v2-v2_T1719|sampled|exact_minesweeper": {
|
| 1198 |
-
"n": 32,
|
| 1199 |
-
"a": 0.0,
|
| 1200 |
-
"b": 0.0,
|
| 1201 |
-
"delta": 0.0,
|
| 1202 |
-
"ci95": [
|
| 1203 |
-
0.0,
|
| 1204 |
-
0.0
|
| 1205 |
-
]
|
| 1206 |
-
},
|
| 1207 |
-
"v2-v2_T1719|greedy|all": {
|
| 1208 |
-
"n": 234,
|
| 1209 |
-
"a": 0.2777777777777778,
|
| 1210 |
-
"b": 0.2777777777777778,
|
| 1211 |
-
"delta": 0.0,
|
| 1212 |
-
"ci95": [
|
| 1213 |
-
0.0,
|
| 1214 |
-
0.0
|
| 1215 |
-
]
|
| 1216 |
-
},
|
| 1217 |
-
"v2-v2_T1719|greedy|bag_game": {
|
| 1218 |
-
"n": 64,
|
| 1219 |
-
"a": 0.484375,
|
| 1220 |
-
"b": 0.484375,
|
| 1221 |
-
"delta": 0.0,
|
| 1222 |
-
"ci95": [
|
| 1223 |
-
0.0,
|
| 1224 |
-
0.0
|
| 1225 |
-
]
|
| 1226 |
-
},
|
| 1227 |
-
"v2-v2_T1719|greedy|stochastic_grid": {
|
| 1228 |
-
"n": 64,
|
| 1229 |
-
"a": 0.1875,
|
| 1230 |
-
"b": 0.1875,
|
| 1231 |
-
"delta": 0.0,
|
| 1232 |
-
"ci95": [
|
| 1233 |
-
0.0,
|
| 1234 |
-
0.0
|
| 1235 |
-
]
|
| 1236 |
-
},
|
| 1237 |
-
"v2-v2_T1719|greedy|tic_tac_toe": {
|
| 1238 |
-
"n": 74,
|
| 1239 |
-
"a": 0.2972972972972973,
|
| 1240 |
-
"b": 0.2972972972972973,
|
| 1241 |
-
"delta": 0.0,
|
| 1242 |
-
"ci95": [
|
| 1243 |
-
0.0,
|
| 1244 |
-
0.0
|
| 1245 |
-
]
|
| 1246 |
-
},
|
| 1247 |
-
"v2-v2_T1719|greedy|exact_minesweeper": {
|
| 1248 |
-
"n": 32,
|
| 1249 |
-
"a": 0.0,
|
| 1250 |
-
"b": 0.0,
|
| 1251 |
-
"delta": 0.0,
|
| 1252 |
-
"ci95": [
|
| 1253 |
-
0.0,
|
| 1254 |
-
0.0
|
| 1255 |
-
]
|
| 1256 |
-
},
|
| 1257 |
-
"v2-v10|sampled|all": {
|
| 1258 |
-
"n": 234,
|
| 1259 |
-
"a": 0.22435897435897437,
|
| 1260 |
-
"b": 0.23717948717948717,
|
| 1261 |
-
"delta": -0.01282051282051282,
|
| 1262 |
-
"ci95": [
|
| 1263 |
-
-0.041666666666666664,
|
| 1264 |
-
0.014957264957264958
|
| 1265 |
-
]
|
| 1266 |
-
},
|
| 1267 |
-
"v2-v10|sampled|bag_game": {
|
| 1268 |
-
"n": 64,
|
| 1269 |
-
"a": 0.37890625,
|
| 1270 |
-
"b": 0.4140625,
|
| 1271 |
-
"delta": -0.03515625,
|
| 1272 |
-
"ci95": [
|
| 1273 |
-
-0.08984375,
|
| 1274 |
-
0.01953125
|
| 1275 |
-
]
|
| 1276 |
-
},
|
| 1277 |
-
"v2-v10|sampled|stochastic_grid": {
|
| 1278 |
-
"n": 64,
|
| 1279 |
-
"a": 0.16015625,
|
| 1280 |
-
"b": 0.1875,
|
| 1281 |
-
"delta": -0.02734375,
|
| 1282 |
-
"ci95": [
|
| 1283 |
-
-0.10546875,
|
| 1284 |
-
0.04296875
|
| 1285 |
-
]
|
| 1286 |
-
},
|
| 1287 |
-
"v2-v10|sampled|tic_tac_toe": {
|
| 1288 |
-
"n": 74,
|
| 1289 |
-
"a": 0.24324324324324326,
|
| 1290 |
-
"b": 0.22972972972972974,
|
| 1291 |
-
"delta": 0.013513513513513514,
|
| 1292 |
-
"ci95": [
|
| 1293 |
-
-0.02364864864864865,
|
| 1294 |
-
0.05405405405405406
|
| 1295 |
-
]
|
| 1296 |
-
},
|
| 1297 |
-
"v2-v10|sampled|exact_minesweeper": {
|
| 1298 |
-
"n": 32,
|
| 1299 |
-
"a": 0.0,
|
| 1300 |
-
"b": 0.0,
|
| 1301 |
-
"delta": 0.0,
|
| 1302 |
-
"ci95": [
|
| 1303 |
-
0.0,
|
| 1304 |
-
0.0
|
| 1305 |
-
]
|
| 1306 |
-
},
|
| 1307 |
-
"v2-v10|greedy|all": {
|
| 1308 |
-
"n": 234,
|
| 1309 |
-
"a": 0.2777777777777778,
|
| 1310 |
-
"b": 0.26495726495726496,
|
| 1311 |
-
"delta": 0.01282051282051282,
|
| 1312 |
-
"ci95": [
|
| 1313 |
-
-0.042735042735042736,
|
| 1314 |
-
0.06837606837606838
|
| 1315 |
-
]
|
| 1316 |
-
},
|
| 1317 |
-
"v2-v10|greedy|bag_game": {
|
| 1318 |
-
"n": 64,
|
| 1319 |
-
"a": 0.484375,
|
| 1320 |
-
"b": 0.578125,
|
| 1321 |
-
"delta": -0.09375,
|
| 1322 |
-
"ci95": [
|
| 1323 |
-
-0.203125,
|
| 1324 |
-
0.015625
|
| 1325 |
-
]
|
| 1326 |
-
},
|
| 1327 |
-
"v2-v10|greedy|stochastic_grid": {
|
| 1328 |
-
"n": 64,
|
| 1329 |
-
"a": 0.1875,
|
| 1330 |
-
"b": 0.15625,
|
| 1331 |
-
"delta": 0.03125,
|
| 1332 |
-
"ci95": [
|
| 1333 |
-
-0.078125,
|
| 1334 |
-
0.140625
|
| 1335 |
-
]
|
| 1336 |
-
},
|
| 1337 |
-
"v2-v10|greedy|tic_tac_toe": {
|
| 1338 |
-
"n": 74,
|
| 1339 |
-
"a": 0.2972972972972973,
|
| 1340 |
-
"b": 0.20270270270270271,
|
| 1341 |
-
"delta": 0.0945945945945946,
|
| 1342 |
-
"ci95": [
|
| 1343 |
-
-0.013513513513513514,
|
| 1344 |
-
0.20270270270270271
|
| 1345 |
-
]
|
| 1346 |
-
},
|
| 1347 |
-
"v2-v10|greedy|exact_minesweeper": {
|
| 1348 |
-
"n": 32,
|
| 1349 |
-
"a": 0.0,
|
| 1350 |
-
"b": 0.0,
|
| 1351 |
-
"delta": 0.0,
|
| 1352 |
-
"ci95": [
|
| 1353 |
-
0.0,
|
| 1354 |
-
0.0
|
| 1355 |
-
]
|
| 1356 |
-
},
|
| 1357 |
-
"v2-35b|sampled|all": {
|
| 1358 |
-
"n": 234,
|
| 1359 |
-
"a": 0.22435897435897437,
|
| 1360 |
-
"b": 0.24145299145299146,
|
| 1361 |
-
"delta": -0.017094017094017096,
|
| 1362 |
-
"ci95": [
|
| 1363 |
-
-0.03739316239316239,
|
| 1364 |
-
0.002136752136752137
|
| 1365 |
-
]
|
| 1366 |
-
},
|
| 1367 |
-
"v2-35b|sampled|bag_game": {
|
| 1368 |
-
"n": 64,
|
| 1369 |
-
"a": 0.37890625,
|
| 1370 |
-
"b": 0.41796875,
|
| 1371 |
-
"delta": -0.0390625,
|
| 1372 |
-
"ci95": [
|
| 1373 |
-
-0.0859375,
|
| 1374 |
-
0.0078125
|
| 1375 |
-
]
|
| 1376 |
-
},
|
| 1377 |
-
"v2-35b|sampled|stochastic_grid": {
|
| 1378 |
-
"n": 64,
|
| 1379 |
-
"a": 0.16015625,
|
| 1380 |
-
"b": 0.16796875,
|
| 1381 |
-
"delta": -0.0078125,
|
| 1382 |
-
"ci95": [
|
| 1383 |
-
-0.04296875,
|
| 1384 |
-
0.02734375
|
| 1385 |
-
]
|
| 1386 |
-
},
|
| 1387 |
-
"v2-35b|sampled|tic_tac_toe": {
|
| 1388 |
-
"n": 74,
|
| 1389 |
-
"a": 0.24324324324324326,
|
| 1390 |
-
"b": 0.25675675675675674,
|
| 1391 |
-
"delta": -0.013513513513513514,
|
| 1392 |
-
"ci95": [
|
| 1393 |
-
-0.05405405405405406,
|
| 1394 |
-
0.02027027027027027
|
| 1395 |
-
]
|
| 1396 |
-
},
|
| 1397 |
-
"v2-35b|sampled|exact_minesweeper": {
|
| 1398 |
-
"n": 32,
|
| 1399 |
-
"a": 0.0,
|
| 1400 |
-
"b": 0.0,
|
| 1401 |
-
"delta": 0.0,
|
| 1402 |
-
"ci95": [
|
| 1403 |
-
0.0,
|
| 1404 |
-
0.0
|
| 1405 |
-
]
|
| 1406 |
-
},
|
| 1407 |
-
"v2-35b|greedy|all": {
|
| 1408 |
-
"n": 234,
|
| 1409 |
-
"a": 0.2777777777777778,
|
| 1410 |
-
"b": 0.3717948717948718,
|
| 1411 |
-
"delta": -0.09401709401709402,
|
| 1412 |
-
"ci95": [
|
| 1413 |
-
-0.14957264957264957,
|
| 1414 |
-
-0.038461538461538464
|
| 1415 |
-
]
|
| 1416 |
-
},
|
| 1417 |
-
"v2-35b|greedy|bag_game": {
|
| 1418 |
-
"n": 64,
|
| 1419 |
-
"a": 0.484375,
|
| 1420 |
-
"b": 0.625,
|
| 1421 |
-
"delta": -0.140625,
|
| 1422 |
-
"ci95": [
|
| 1423 |
-
-0.234375,
|
| 1424 |
-
-0.0625
|
| 1425 |
-
]
|
| 1426 |
-
},
|
| 1427 |
-
"v2-35b|greedy|stochastic_grid": {
|
| 1428 |
-
"n": 64,
|
| 1429 |
-
"a": 0.1875,
|
| 1430 |
-
"b": 0.328125,
|
| 1431 |
-
"delta": -0.140625,
|
| 1432 |
-
"ci95": [
|
| 1433 |
-
-0.25,
|
| 1434 |
-
-0.03125
|
| 1435 |
-
]
|
| 1436 |
-
},
|
| 1437 |
-
"v2-35b|greedy|tic_tac_toe": {
|
| 1438 |
-
"n": 74,
|
| 1439 |
-
"a": 0.2972972972972973,
|
| 1440 |
-
"b": 0.33783783783783783,
|
| 1441 |
-
"delta": -0.04054054054054054,
|
| 1442 |
-
"ci95": [
|
| 1443 |
-
-0.16216216216216217,
|
| 1444 |
-
0.08108108108108109
|
| 1445 |
-
]
|
| 1446 |
-
},
|
| 1447 |
-
"v2-35b|greedy|exact_minesweeper": {
|
| 1448 |
-
"n": 32,
|
| 1449 |
-
"a": 0.0,
|
| 1450 |
-
"b": 0.03125,
|
| 1451 |
-
"delta": -0.03125,
|
| 1452 |
-
"ci95": [
|
| 1453 |
-
-0.09375,
|
| 1454 |
-
0.0
|
| 1455 |
-
]
|
| 1456 |
-
}
|
| 1457 |
},
|
| 1458 |
-
"
|
| 1459 |
-
"
|
| 1460 |
-
"
|
| 1461 |
-
|
| 1462 |
-
|
| 1463 |
-
|
| 1464 |
-
|
| 1465 |
-
|
| 1466 |
-
|
| 1467 |
-
|
| 1468 |
-
|
| 1469 |
-
|
| 1470 |
-
|
| 1471 |
-
|
| 1472 |
-
|
| 1473 |
-
|
| 1474 |
-
|
| 1475 |
-
|
| 1476 |
-
|
| 1477 |
-
|
| 1478 |
-
|
| 1479 |
-
|
| 1480 |
-
|
| 1481 |
-
|
| 1482 |
-
|
| 1483 |
-
"
|
| 1484 |
-
|
| 1485 |
-
|
| 1486 |
-
|
| 1487 |
-
|
| 1488 |
-
|
| 1489 |
-
|
| 1490 |
-
|
| 1491 |
-
|
| 1492 |
-
"
|
| 1493 |
-
|
| 1494 |
-
|
| 1495 |
-
|
| 1496 |
-
|
| 1497 |
-
|
| 1498 |
-
|
| 1499 |
-
|
| 1500 |
-
|
| 1501 |
-
|
| 1502 |
-
|
| 1503 |
-
|
| 1504 |
-
|
| 1505 |
-
|
| 1506 |
-
0.0078125
|
| 1507 |
-
]
|
| 1508 |
-
},
|
| 1509 |
-
"v2-v1|sampled|held_out": {
|
| 1510 |
-
"n": 48,
|
| 1511 |
-
"a": 0.75,
|
| 1512 |
-
"b": 0.7708333333333334,
|
| 1513 |
-
"delta": -0.020833333333333332,
|
| 1514 |
-
"ci95": [
|
| 1515 |
-
-0.14583333333333334,
|
| 1516 |
-
0.10416666666666667
|
| 1517 |
-
]
|
| 1518 |
-
},
|
| 1519 |
-
"v2-v2_T1719|greedy|all": {
|
| 1520 |
-
"n": 176,
|
| 1521 |
-
"a": 0.9261363636363636,
|
| 1522 |
-
"b": 0.9261363636363636,
|
| 1523 |
-
"delta": 0.0,
|
| 1524 |
-
"ci95": [
|
| 1525 |
-
0.0,
|
| 1526 |
-
0.0
|
| 1527 |
-
]
|
| 1528 |
-
},
|
| 1529 |
-
"v2-v2_T1719|greedy|rewarded": {
|
| 1530 |
-
"n": 128,
|
| 1531 |
-
"a": 0.96875,
|
| 1532 |
-
"b": 0.96875,
|
| 1533 |
-
"delta": 0.0,
|
| 1534 |
-
"ci95": [
|
| 1535 |
-
0.0,
|
| 1536 |
-
0.0
|
| 1537 |
-
]
|
| 1538 |
-
},
|
| 1539 |
-
"v2-v2_T1719|greedy|held_out": {
|
| 1540 |
-
"n": 48,
|
| 1541 |
-
"a": 0.8125,
|
| 1542 |
-
"b": 0.8125,
|
| 1543 |
-
"delta": 0.0,
|
| 1544 |
-
"ci95": [
|
| 1545 |
-
0.0,
|
| 1546 |
-
0.0
|
| 1547 |
-
]
|
| 1548 |
-
},
|
| 1549 |
-
"v2-v2_T1719|sampled|all": {
|
| 1550 |
-
"n": 176,
|
| 1551 |
-
"a": 0.8806818181818182,
|
| 1552 |
-
"b": 0.8636363636363636,
|
| 1553 |
-
"delta": 0.017045454545454544,
|
| 1554 |
-
"ci95": [
|
| 1555 |
-
-0.022727272727272728,
|
| 1556 |
-
0.0625
|
| 1557 |
-
]
|
| 1558 |
-
},
|
| 1559 |
-
"v2-v2_T1719|sampled|rewarded": {
|
| 1560 |
-
"n": 128,
|
| 1561 |
-
"a": 0.9296875,
|
| 1562 |
-
"b": 0.8984375,
|
| 1563 |
-
"delta": 0.03125,
|
| 1564 |
-
"ci95": [
|
| 1565 |
-
-0.015625,
|
| 1566 |
-
0.078125
|
| 1567 |
-
]
|
| 1568 |
-
},
|
| 1569 |
-
"v2-v2_T1719|sampled|held_out": {
|
| 1570 |
-
"n": 48,
|
| 1571 |
-
"a": 0.75,
|
| 1572 |
-
"b": 0.7708333333333334,
|
| 1573 |
-
"delta": -0.020833333333333332,
|
| 1574 |
-
"ci95": [
|
| 1575 |
-
-0.10416666666666667,
|
| 1576 |
-
0.0625
|
| 1577 |
-
]
|
| 1578 |
-
},
|
| 1579 |
-
"v2-v10|greedy|all": {
|
| 1580 |
-
"n": 176,
|
| 1581 |
-
"a": 0.9261363636363636,
|
| 1582 |
-
"b": 0.9090909090909091,
|
| 1583 |
-
"delta": 0.017045454545454544,
|
| 1584 |
-
"ci95": [
|
| 1585 |
-
-0.03977272727272727,
|
| 1586 |
-
0.07386363636363637
|
| 1587 |
-
]
|
| 1588 |
-
},
|
| 1589 |
-
"v2-v10|greedy|rewarded": {
|
| 1590 |
-
"n": 128,
|
| 1591 |
-
"a": 0.96875,
|
| 1592 |
-
"b": 0.90625,
|
| 1593 |
-
"delta": 0.0625,
|
| 1594 |
-
"ci95": [
|
| 1595 |
-
0.0078125,
|
| 1596 |
-
0.125
|
| 1597 |
-
]
|
| 1598 |
-
},
|
| 1599 |
-
"v2-v10|greedy|held_out": {
|
| 1600 |
-
"n": 48,
|
| 1601 |
-
"a": 0.8125,
|
| 1602 |
-
"b": 0.9166666666666666,
|
| 1603 |
-
"delta": -0.10416666666666667,
|
| 1604 |
-
"ci95": [
|
| 1605 |
-
-0.22916666666666666,
|
| 1606 |
-
0.020833333333333332
|
| 1607 |
-
]
|
| 1608 |
-
},
|
| 1609 |
-
"v2-v10|sampled|all": {
|
| 1610 |
-
"n": 176,
|
| 1611 |
-
"a": 0.8806818181818182,
|
| 1612 |
-
"b": 0.9318181818181818,
|
| 1613 |
-
"delta": -0.05113636363636364,
|
| 1614 |
-
"ci95": [
|
| 1615 |
-
-0.10795454545454546,
|
| 1616 |
-
0.005681818181818182
|
| 1617 |
-
]
|
| 1618 |
-
},
|
| 1619 |
-
"v2-v10|sampled|rewarded": {
|
| 1620 |
-
"n": 128,
|
| 1621 |
-
"a": 0.9296875,
|
| 1622 |
-
"b": 0.9375,
|
| 1623 |
-
"delta": -0.0078125,
|
| 1624 |
-
"ci95": [
|
| 1625 |
-
-0.0703125,
|
| 1626 |
-
0.0546875
|
| 1627 |
-
]
|
| 1628 |
-
},
|
| 1629 |
-
"v2-v10|sampled|held_out": {
|
| 1630 |
-
"n": 48,
|
| 1631 |
-
"a": 0.75,
|
| 1632 |
-
"b": 0.9166666666666666,
|
| 1633 |
-
"delta": -0.16666666666666666,
|
| 1634 |
-
"ci95": [
|
| 1635 |
-
-0.3125,
|
| 1636 |
-
-0.041666666666666664
|
| 1637 |
-
]
|
| 1638 |
-
},
|
| 1639 |
-
"v2-35b|greedy|all": {
|
| 1640 |
-
"n": 176,
|
| 1641 |
-
"a": 0.9261363636363636,
|
| 1642 |
-
"b": 0.9715909090909091,
|
| 1643 |
-
"delta": -0.045454545454545456,
|
| 1644 |
-
"ci95": [
|
| 1645 |
-
-0.07954545454545454,
|
| 1646 |
-
-0.017045454545454544
|
| 1647 |
-
]
|
| 1648 |
-
},
|
| 1649 |
-
"v2-35b|greedy|rewarded": {
|
| 1650 |
-
"n": 128,
|
| 1651 |
-
"a": 0.96875,
|
| 1652 |
-
"b": 0.96875,
|
| 1653 |
-
"delta": 0.0,
|
| 1654 |
-
"ci95": [
|
| 1655 |
-
0.0,
|
| 1656 |
-
0.0
|
| 1657 |
-
]
|
| 1658 |
-
},
|
| 1659 |
-
"v2-35b|greedy|held_out": {
|
| 1660 |
-
"n": 48,
|
| 1661 |
-
"a": 0.8125,
|
| 1662 |
-
"b": 0.9791666666666666,
|
| 1663 |
-
"delta": -0.16666666666666666,
|
| 1664 |
-
"ci95": [
|
| 1665 |
-
-0.2708333333333333,
|
| 1666 |
-
-0.0625
|
| 1667 |
-
]
|
| 1668 |
-
},
|
| 1669 |
-
"v2-35b|sampled|all": {
|
| 1670 |
-
"n": 176,
|
| 1671 |
-
"a": 0.8806818181818182,
|
| 1672 |
-
"b": 0.8636363636363636,
|
| 1673 |
-
"delta": 0.017045454545454544,
|
| 1674 |
-
"ci95": [
|
| 1675 |
-
-0.03409090909090909,
|
| 1676 |
-
0.06832386363636415
|
| 1677 |
-
]
|
| 1678 |
-
},
|
| 1679 |
-
"v2-35b|sampled|rewarded": {
|
| 1680 |
-
"n": 128,
|
| 1681 |
-
"a": 0.9296875,
|
| 1682 |
-
"b": 0.890625,
|
| 1683 |
-
"delta": 0.0390625,
|
| 1684 |
-
"ci95": [
|
| 1685 |
-
-0.015625,
|
| 1686 |
-
0.09375
|
| 1687 |
-
]
|
| 1688 |
-
},
|
| 1689 |
-
"v2-35b|sampled|held_out": {
|
| 1690 |
-
"n": 48,
|
| 1691 |
-
"a": 0.75,
|
| 1692 |
-
"b": 0.7916666666666666,
|
| 1693 |
-
"delta": -0.041666666666666664,
|
| 1694 |
-
"ci95": [
|
| 1695 |
-
-0.16666666666666666,
|
| 1696 |
-
0.08333333333333333
|
| 1697 |
-
]
|
| 1698 |
-
},
|
| 1699 |
-
"per_task|greedy": {
|
| 1700 |
-
"click-button": 1.0,
|
| 1701 |
-
"click-button-sequence": 1.0,
|
| 1702 |
-
"click-checkboxes": 1.0,
|
| 1703 |
-
"click-checkboxes-large": 0.625,
|
| 1704 |
-
"click-checkboxes-soft": 1.0,
|
| 1705 |
-
"click-checkboxes-transfer": 1.0,
|
| 1706 |
-
"click-collapsible": 1.0,
|
| 1707 |
-
"click-collapsible-2": 0.5,
|
| 1708 |
-
"click-color": 1.0,
|
| 1709 |
-
"click-dialog": 1.0,
|
| 1710 |
-
"click-dialog-2": 1.0,
|
| 1711 |
-
"click-link": 1.0,
|
| 1712 |
-
"click-option": 1.0,
|
| 1713 |
-
"click-tab": 1.0,
|
| 1714 |
-
"click-tab-2": 0.875,
|
| 1715 |
-
"click-tab-2-hard": 1.0,
|
| 1716 |
-
"click-test": 1.0,
|
| 1717 |
-
"click-test-2": 1.0,
|
| 1718 |
-
"click-widget": 1.0,
|
| 1719 |
-
"focus-text": 1.0,
|
| 1720 |
-
"focus-text-2": 0.375,
|
| 1721 |
-
"navigate-tree": 1.0
|
| 1722 |
-
},
|
| 1723 |
-
"per_task|sampled": {
|
| 1724 |
-
"click-button": 1.0,
|
| 1725 |
-
"click-button-sequence": 1.0,
|
| 1726 |
-
"click-checkboxes": 1.0,
|
| 1727 |
-
"click-checkboxes-large": 0.75,
|
| 1728 |
-
"click-checkboxes-soft": 0.875,
|
| 1729 |
-
"click-checkboxes-transfer": 0.875,
|
| 1730 |
-
"click-collapsible": 1.0,
|
| 1731 |
-
"click-collapsible-2": 0.75,
|
| 1732 |
-
"click-color": 1.0,
|
| 1733 |
-
"click-dialog": 1.0,
|
| 1734 |
-
"click-dialog-2": 0.5,
|
| 1735 |
-
"click-link": 1.0,
|
| 1736 |
-
"click-option": 0.875,
|
| 1737 |
-
"click-tab": 0.625,
|
| 1738 |
-
"click-tab-2": 0.75,
|
| 1739 |
-
"click-tab-2-hard": 0.625,
|
| 1740 |
-
"click-test": 1.0,
|
| 1741 |
-
"click-test-2": 1.0,
|
| 1742 |
-
"click-widget": 1.0,
|
| 1743 |
-
"focus-text": 1.0,
|
| 1744 |
-
"focus-text-2": 0.75,
|
| 1745 |
-
"navigate-tree": 1.0
|
| 1746 |
}
|
| 1747 |
}
|
| 1748 |
},
|
| 1749 |
-
"
|
| 1750 |
-
"
|
| 1751 |
-
"
|
| 1752 |
-
"
|
| 1753 |
-
"
|
| 1754 |
-
"
|
| 1755 |
-
"
|
| 1756 |
-
|
| 1757 |
-
|
| 1758 |
-
|
| 1759 |
-
"a_only": 0,
|
| 1760 |
-
"b_only": 0,
|
| 1761 |
-
"argmax_differs": 0
|
| 1762 |
},
|
| 1763 |
-
"
|
| 1764 |
-
"
|
| 1765 |
-
|
| 1766 |
-
|
| 1767 |
-
|
| 1768 |
-
|
| 1769 |
-
|
| 1770 |
-
|
| 1771 |
-
|
| 1772 |
-
|
| 1773 |
-
|
| 1774 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1775 |
},
|
| 1776 |
-
"
|
| 1777 |
-
"
|
| 1778 |
-
|
| 1779 |
-
|
| 1780 |
-
|
| 1781 |
-
|
| 1782 |
-
|
| 1783 |
-
|
| 1784 |
-
|
| 1785 |
-
|
| 1786 |
-
|
| 1787 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1788 |
}
|
| 1789 |
}
|
| 1790 |
},
|
| 1791 |
"references": {
|
| 1792 |
-
"
|
| 1793 |
-
|
| 1794 |
-
|
| 1795 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1796 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"model": "decider-4b-v2.1",
|
| 3 |
+
"temperature": 1.099,
|
| 4 |
+
"temperature_by_type": {
|
| 5 |
+
"choice": 1.11,
|
| 6 |
+
"noul": 1.56,
|
| 7 |
+
"score": 1.287
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
},
|
| 9 |
+
"temperature_fit": {
|
| 10 |
+
"global": "one scalar, pooled row NLL on the in-task half of the public regression set without banking77, clinc_oos, mmlu, arc, winogrande and hellaswag (61 tasks, 102,804 rows)",
|
| 11 |
+
"by_type": "decider.calibrate.fit_by_type (decider-ai 1.4.0), min_rows 50, on the same regression rows plus our own validation rows (val_jb, val_teacher, val_teacher2, cal_human without mmlu, arc and banking77, val_replay without six excluded families)",
|
| 12 |
+
"rows": {
|
| 13 |
+
"choice": 108927,
|
| 14 |
+
"noul": 1677,
|
| 15 |
+
"score": 535
|
| 16 |
+
},
|
| 17 |
+
"note": "no JevBench or Decision Index item was used for training, selection or temperature"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 18 |
},
|
| 19 |
+
"protocol": "Regression set: eval half of the mixture, 95 tasks (67 in-task / 28 held-out), plain state-first layout, large label sets sub-sampled to 10 options, ECE 15 bins, per-task mean. Own sets (heldout_jb, test_teacher2, cal_human, guard): stored T=1 logits read at the served temperatures, ECE 10 bins top label. Card lanes (fixtures, games, browser, probes, Bespoke, issue #9) through decider-ai 1.4.0 with the map; parents through 1.3.0 at their released temperatures, same rows and seeds; intervals are 95% paired bootstrap. JevBench public items read in process through 1.4.0 with the map; parents' files read earlier through 1.2.1.",
|
| 20 |
+
"regression": {
|
| 21 |
+
"map": {
|
| 22 |
+
"reg_in": 0.8308,
|
| 23 |
+
"reg_out": 0.7838,
|
| 24 |
+
"reg_in_nll": 0.4145,
|
| 25 |
+
"reg_out_nll": 0.5689,
|
| 26 |
+
"reg_in_ece": 0.0309,
|
| 27 |
+
"reg_out_ece": 0.0775
|
| 28 |
+
},
|
| 29 |
+
"global_T": {
|
| 30 |
+
"reg_in": 0.8308,
|
| 31 |
+
"reg_out": 0.7838,
|
| 32 |
+
"reg_in_nll": 0.4145,
|
| 33 |
+
"reg_out_nll": 0.5703,
|
| 34 |
+
"reg_in_ece": 0.0308,
|
| 35 |
+
"reg_out_ece": 0.0781
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
}
|
| 37 |
},
|
| 38 |
+
"own_sets": {
|
| 39 |
+
"map": {
|
| 40 |
+
"heldout_jb": {
|
| 41 |
+
"all": {
|
| 42 |
+
"n": 5000,
|
| 43 |
+
"accuracy": 0.556,
|
| 44 |
+
"nll": 1.2571,
|
| 45 |
+
"ece": 0.1466,
|
| 46 |
+
"mean_conf": 0.7026
|
| 47 |
+
},
|
| 48 |
+
"by_type": {
|
| 49 |
+
"choice": {
|
| 50 |
+
"n": 3293,
|
| 51 |
+
"accuracy": 0.5056,
|
| 52 |
+
"nll": 1.4995,
|
| 53 |
+
"ece": 0.1701,
|
| 54 |
+
"mean_conf": 0.6757
|
| 55 |
+
},
|
| 56 |
+
"noul": {
|
| 57 |
+
"n": 1289,
|
| 58 |
+
"accuracy": 0.7114,
|
| 59 |
+
"nll": 0.6074,
|
| 60 |
+
"ece": 0.1013,
|
| 61 |
+
"mean_conf": 0.8127
|
| 62 |
+
},
|
| 63 |
+
"score": {
|
| 64 |
+
"n": 418,
|
| 65 |
+
"accuracy": 0.4737,
|
| 66 |
+
"nll": 1.3504,
|
| 67 |
+
"ece": 0.1081,
|
| 68 |
+
"mean_conf": 0.5747
|
| 69 |
+
}
|
| 70 |
+
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 71 |
},
|
| 72 |
+
"test_teacher2": {
|
| 73 |
+
"all": {
|
| 74 |
+
"n": 449,
|
| 75 |
+
"accuracy": 0.8196,
|
| 76 |
+
"nll": 0.4304,
|
| 77 |
+
"ece": 0.043,
|
| 78 |
+
"mean_conf": 0.8559
|
| 79 |
+
},
|
| 80 |
+
"by_type": {
|
| 81 |
+
"choice": {
|
| 82 |
+
"n": 257,
|
| 83 |
+
"accuracy": 0.8016,
|
| 84 |
+
"nll": 0.5041,
|
| 85 |
+
"ece": 0.0569,
|
| 86 |
+
"mean_conf": 0.8585
|
| 87 |
+
},
|
| 88 |
+
"noul": {
|
| 89 |
+
"n": 130,
|
| 90 |
+
"accuracy": 0.8692,
|
| 91 |
+
"nll": 0.2489,
|
| 92 |
+
"ece": 0.0479,
|
| 93 |
+
"mean_conf": 0.905
|
| 94 |
+
},
|
| 95 |
+
"score": {
|
| 96 |
+
"n": 62,
|
| 97 |
+
"accuracy": 0.7903,
|
| 98 |
+
"nll": 0.5051,
|
| 99 |
+
"ece": 0.0863,
|
| 100 |
+
"mean_conf": 0.7423
|
| 101 |
+
}
|
| 102 |
+
}
|
| 103 |
},
|
| 104 |
+
"guard": {
|
| 105 |
+
"all": {
|
| 106 |
+
"n": 2994,
|
| 107 |
+
"accuracy": 0.8186,
|
| 108 |
+
"nll": 0.4866,
|
| 109 |
+
"ece": 0.0168,
|
| 110 |
+
"mean_conf": 0.8344
|
| 111 |
+
},
|
| 112 |
+
"by_type": {
|
| 113 |
+
"choice": {
|
| 114 |
+
"n": 2994,
|
| 115 |
+
"accuracy": 0.8186,
|
| 116 |
+
"nll": 0.4866,
|
| 117 |
+
"ece": 0.0168,
|
| 118 |
+
"mean_conf": 0.8344
|
| 119 |
+
},
|
| 120 |
+
"noul": null,
|
| 121 |
+
"score": null
|
| 122 |
+
}
|
| 123 |
},
|
| 124 |
+
"val_jb": {
|
| 125 |
+
"all": {
|
| 126 |
+
"n": 5000,
|
| 127 |
+
"accuracy": 0.796,
|
| 128 |
+
"nll": 0.5264,
|
| 129 |
+
"ece": 0.0157,
|
| 130 |
+
"mean_conf": 0.8117
|
| 131 |
+
},
|
| 132 |
+
"by_type": {
|
| 133 |
+
"choice": {
|
| 134 |
+
"n": 3374,
|
| 135 |
+
"accuracy": 0.7647,
|
| 136 |
+
"nll": 0.6258,
|
| 137 |
+
"ece": 0.0289,
|
| 138 |
+
"mean_conf": 0.7935
|
| 139 |
+
},
|
| 140 |
+
"noul": {
|
| 141 |
+
"n": 1273,
|
| 142 |
+
"accuracy": 0.901,
|
| 143 |
+
"nll": 0.2233,
|
| 144 |
+
"ece": 0.0161,
|
| 145 |
+
"mean_conf": 0.8849
|
| 146 |
+
},
|
| 147 |
+
"score": {
|
| 148 |
+
"n": 353,
|
| 149 |
+
"accuracy": 0.7167,
|
| 150 |
+
"nll": 0.6681,
|
| 151 |
+
"ece": 0.0317,
|
| 152 |
+
"mean_conf": 0.722
|
| 153 |
+
}
|
| 154 |
+
}
|
| 155 |
},
|
| 156 |
+
"val_teacher": {
|
| 157 |
+
"all": {
|
| 158 |
+
"n": 944,
|
| 159 |
+
"accuracy": 0.8379,
|
| 160 |
+
"nll": 0.4169,
|
| 161 |
+
"ece": 0.0369,
|
| 162 |
+
"mean_conf": 0.8619
|
| 163 |
+
},
|
| 164 |
+
"by_type": {
|
| 165 |
+
"choice": {
|
| 166 |
+
"n": 540,
|
| 167 |
+
"accuracy": 0.8315,
|
| 168 |
+
"nll": 0.4643,
|
| 169 |
+
"ece": 0.0464,
|
| 170 |
+
"mean_conf": 0.8547
|
| 171 |
+
},
|
| 172 |
+
"noul": {
|
| 173 |
+
"n": 284,
|
| 174 |
+
"accuracy": 0.8873,
|
| 175 |
+
"nll": 0.2931,
|
| 176 |
+
"ece": 0.0453,
|
| 177 |
+
"mean_conf": 0.9088
|
| 178 |
+
},
|
| 179 |
+
"score": {
|
| 180 |
+
"n": 120,
|
| 181 |
+
"accuracy": 0.75,
|
| 182 |
+
"nll": 0.4969,
|
| 183 |
+
"ece": 0.0576,
|
| 184 |
+
"mean_conf": 0.7831
|
| 185 |
+
}
|
| 186 |
+
}
|
| 187 |
},
|
| 188 |
+
"val_teacher2": {
|
| 189 |
+
"all": {
|
| 190 |
+
"n": 447,
|
| 191 |
+
"accuracy": 0.8322,
|
| 192 |
+
"nll": 0.4435,
|
| 193 |
+
"ece": 0.0627,
|
| 194 |
+
"mean_conf": 0.8747
|
| 195 |
+
},
|
| 196 |
+
"by_type": {
|
| 197 |
+
"choice": {
|
| 198 |
+
"n": 265,
|
| 199 |
+
"accuracy": 0.8377,
|
| 200 |
+
"nll": 0.4219,
|
| 201 |
+
"ece": 0.0503,
|
| 202 |
+
"mean_conf": 0.8858
|
| 203 |
+
},
|
| 204 |
+
"noul": {
|
| 205 |
+
"n": 120,
|
| 206 |
+
"accuracy": 0.85,
|
| 207 |
+
"nll": 0.3702,
|
| 208 |
+
"ece": 0.0827,
|
| 209 |
+
"mean_conf": 0.9034
|
| 210 |
+
},
|
| 211 |
+
"score": {
|
| 212 |
+
"n": 62,
|
| 213 |
+
"accuracy": 0.7742,
|
| 214 |
+
"nll": 0.6773,
|
| 215 |
+
"ece": 0.1441,
|
| 216 |
+
"mean_conf": 0.772
|
| 217 |
+
}
|
| 218 |
+
}
|
| 219 |
},
|
| 220 |
+
"cal_human": {
|
| 221 |
+
"all": {
|
| 222 |
+
"n": 1595,
|
| 223 |
+
"accuracy": 0.8564,
|
| 224 |
+
"nll": 0.3892,
|
| 225 |
+
"ece": 0.0205,
|
| 226 |
+
"mean_conf": 0.8717
|
| 227 |
+
},
|
| 228 |
+
"by_type": {
|
| 229 |
+
"choice": {
|
| 230 |
+
"n": 1595,
|
| 231 |
+
"accuracy": 0.8564,
|
| 232 |
+
"nll": 0.3892,
|
| 233 |
+
"ece": 0.0205,
|
| 234 |
+
"mean_conf": 0.8717
|
| 235 |
+
},
|
| 236 |
+
"noul": null,
|
| 237 |
+
"score": null
|
| 238 |
+
}
|
| 239 |
},
|
| 240 |
+
"val_replay4": {
|
| 241 |
+
"all": {
|
| 242 |
+
"n": 1000,
|
| 243 |
+
"accuracy": 0.889,
|
| 244 |
+
"nll": 0.2848,
|
| 245 |
+
"ece": 0.0297,
|
| 246 |
+
"mean_conf": 0.8702
|
| 247 |
+
},
|
| 248 |
+
"by_type": {
|
| 249 |
+
"choice": {
|
| 250 |
+
"n": 1000,
|
| 251 |
+
"accuracy": 0.889,
|
| 252 |
+
"nll": 0.2848,
|
| 253 |
+
"ece": 0.0297,
|
| 254 |
+
"mean_conf": 0.8702
|
| 255 |
+
},
|
| 256 |
+
"noul": null,
|
| 257 |
+
"score": null
|
| 258 |
+
}
|
| 259 |
}
|
| 260 |
},
|
| 261 |
+
"global_T": {
|
| 262 |
+
"heldout_jb": {
|
| 263 |
+
"all": {
|
| 264 |
+
"n": 5000,
|
| 265 |
+
"accuracy": 0.5562,
|
| 266 |
+
"nll": 1.2937,
|
| 267 |
+
"ece": 0.163,
|
| 268 |
+
"mean_conf": 0.7192
|
| 269 |
+
},
|
| 270 |
+
"by_type": {
|
| 271 |
+
"choice": {
|
| 272 |
+
"n": 3293,
|
| 273 |
+
"accuracy": 0.5056,
|
| 274 |
+
"nll": 1.5059,
|
| 275 |
+
"ece": 0.1724,
|
| 276 |
+
"mean_conf": 0.678
|
| 277 |
+
},
|
| 278 |
+
"noul": {
|
| 279 |
+
"n": 1289,
|
| 280 |
+
"accuracy": 0.7114,
|
| 281 |
+
"nll": 0.7112,
|
| 282 |
+
"ece": 0.1463,
|
| 283 |
+
"mean_conf": 0.8577
|
| 284 |
+
},
|
| 285 |
+
"score": {
|
| 286 |
+
"n": 418,
|
| 287 |
+
"accuracy": 0.4761,
|
| 288 |
+
"nll": 1.4183,
|
| 289 |
+
"ece": 0.1402,
|
| 290 |
+
"mean_conf": 0.6163
|
| 291 |
+
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 292 |
}
|
| 293 |
},
|
| 294 |
+
"test_teacher2": {
|
| 295 |
+
"all": {
|
| 296 |
+
"n": 449,
|
| 297 |
+
"accuracy": 0.8196,
|
| 298 |
+
"nll": 0.4381,
|
| 299 |
+
"ece": 0.0497,
|
| 300 |
+
"mean_conf": 0.8664
|
| 301 |
+
},
|
| 302 |
+
"by_type": {
|
| 303 |
+
"choice": {
|
| 304 |
+
"n": 257,
|
| 305 |
+
"accuracy": 0.8016,
|
| 306 |
+
"nll": 0.5058,
|
| 307 |
+
"ece": 0.058,
|
| 308 |
+
"mean_conf": 0.8596
|
| 309 |
+
},
|
| 310 |
+
"noul": {
|
| 311 |
+
"n": 130,
|
| 312 |
+
"accuracy": 0.8692,
|
| 313 |
+
"nll": 0.2755,
|
| 314 |
+
"ece": 0.0602,
|
| 315 |
+
"mean_conf": 0.9294
|
| 316 |
+
},
|
| 317 |
+
"score": {
|
| 318 |
+
"n": 62,
|
| 319 |
+
"accuracy": 0.7903,
|
| 320 |
+
"nll": 0.4984,
|
| 321 |
+
"ece": 0.0798,
|
| 322 |
+
"mean_conf": 0.7623
|
| 323 |
+
}
|
|
|
|
|
|
|
|
|
|
| 324 |
}
|
| 325 |
},
|
| 326 |
+
"guard": {
|
| 327 |
+
"all": {
|
| 328 |
+
"n": 2994,
|
| 329 |
+
"accuracy": 0.8186,
|
| 330 |
+
"nll": 0.4871,
|
| 331 |
+
"ece": 0.0194,
|
| 332 |
+
"mean_conf": 0.836
|
| 333 |
+
},
|
| 334 |
+
"by_type": {
|
| 335 |
+
"choice": {
|
| 336 |
+
"n": 2994,
|
| 337 |
+
"accuracy": 0.8186,
|
| 338 |
+
"nll": 0.4871,
|
| 339 |
+
"ece": 0.0194,
|
| 340 |
+
"mean_conf": 0.836
|
| 341 |
+
},
|
| 342 |
+
"noul": null,
|
| 343 |
+
"score": null
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 344 |
}
|
| 345 |
},
|
| 346 |
+
"val_jb": {
|
| 347 |
+
"all": {
|
| 348 |
+
"n": 5000,
|
| 349 |
+
"accuracy": 0.796,
|
| 350 |
+
"nll": 0.5269,
|
| 351 |
+
"ece": 0.0264,
|
| 352 |
+
"mean_conf": 0.8224
|
| 353 |
+
},
|
| 354 |
+
"by_type": {
|
| 355 |
+
"choice": {
|
| 356 |
+
"n": 3374,
|
| 357 |
+
"accuracy": 0.7647,
|
| 358 |
+
"nll": 0.6267,
|
| 359 |
+
"ece": 0.0303,
|
| 360 |
+
"mean_conf": 0.795
|
| 361 |
+
},
|
| 362 |
+
"noul": {
|
| 363 |
+
"n": 1273,
|
| 364 |
+
"accuracy": 0.901,
|
| 365 |
+
"nll": 0.223,
|
| 366 |
+
"ece": 0.0172,
|
| 367 |
+
"mean_conf": 0.9146
|
| 368 |
+
},
|
| 369 |
+
"score": {
|
| 370 |
+
"n": 353,
|
| 371 |
+
"accuracy": 0.7167,
|
| 372 |
+
"nll": 0.6687,
|
| 373 |
+
"ece": 0.0363,
|
| 374 |
+
"mean_conf": 0.7514
|
| 375 |
+
}
|
|
|
|
|
|
|
|
|
|
| 376 |
}
|
| 377 |
},
|
| 378 |
+
"val_teacher": {
|
| 379 |
+
"all": {
|
| 380 |
+
"n": 944,
|
| 381 |
+
"accuracy": 0.8379,
|
| 382 |
+
"nll": 0.433,
|
| 383 |
+
"ece": 0.046,
|
| 384 |
+
"mean_conf": 0.8722
|
| 385 |
+
},
|
| 386 |
+
"by_type": {
|
| 387 |
+
"choice": {
|
| 388 |
+
"n": 540,
|
| 389 |
+
"accuracy": 0.8315,
|
| 390 |
+
"nll": 0.4654,
|
| 391 |
+
"ece": 0.0435,
|
| 392 |
+
"mean_conf": 0.8558
|
| 393 |
+
},
|
| 394 |
+
"noul": {
|
| 395 |
+
"n": 284,
|
| 396 |
+
"accuracy": 0.8873,
|
| 397 |
+
"nll": 0.3405,
|
| 398 |
+
"ece": 0.0634,
|
| 399 |
+
"mean_conf": 0.9331
|
| 400 |
+
},
|
| 401 |
+
"score": {
|
| 402 |
+
"n": 120,
|
| 403 |
+
"accuracy": 0.75,
|
| 404 |
+
"nll": 0.5064,
|
| 405 |
+
"ece": 0.0603,
|
| 406 |
+
"mean_conf": 0.8016
|
| 407 |
+
}
|
|
|
|
|
|
|
|
|
|
| 408 |
}
|
| 409 |
},
|
| 410 |
+
"val_teacher2": {
|
| 411 |
+
"all": {
|
| 412 |
+
"n": 447,
|
| 413 |
+
"accuracy": 0.8277,
|
| 414 |
+
"nll": 0.4695,
|
| 415 |
+
"ece": 0.0806,
|
| 416 |
+
"mean_conf": 0.8852
|
| 417 |
+
},
|
| 418 |
+
"by_type": {
|
| 419 |
+
"choice": {
|
| 420 |
+
"n": 265,
|
| 421 |
+
"accuracy": 0.8377,
|
| 422 |
+
"nll": 0.4233,
|
| 423 |
+
"ece": 0.0546,
|
| 424 |
+
"mean_conf": 0.8868
|
| 425 |
+
},
|
| 426 |
+
"noul": {
|
| 427 |
+
"n": 120,
|
| 428 |
+
"accuracy": 0.85,
|
| 429 |
+
"nll": 0.4441,
|
| 430 |
+
"ece": 0.1059,
|
| 431 |
+
"mean_conf": 0.9311
|
| 432 |
+
},
|
| 433 |
+
"score": {
|
| 434 |
+
"n": 62,
|
| 435 |
+
"accuracy": 0.7419,
|
| 436 |
+
"nll": 0.7165,
|
| 437 |
+
"ece": 0.1556,
|
| 438 |
+
"mean_conf": 0.7893
|
| 439 |
+
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 440 |
}
|
| 441 |
},
|
| 442 |
+
"cal_human": {
|
| 443 |
+
"all": {
|
| 444 |
+
"n": 1595,
|
| 445 |
+
"accuracy": 0.8564,
|
| 446 |
+
"nll": 0.3893,
|
| 447 |
+
"ece": 0.0195,
|
| 448 |
+
"mean_conf": 0.8733
|
| 449 |
+
},
|
| 450 |
+
"by_type": {
|
| 451 |
+
"choice": {
|
| 452 |
+
"n": 1595,
|
| 453 |
+
"accuracy": 0.8564,
|
| 454 |
+
"nll": 0.3893,
|
| 455 |
+
"ece": 0.0195,
|
| 456 |
+
"mean_conf": 0.8733
|
| 457 |
+
},
|
| 458 |
+
"noul": null,
|
| 459 |
+
"score": null
|
| 460 |
}
|
| 461 |
},
|
| 462 |
+
"val_replay4": {
|
| 463 |
+
"all": {
|
| 464 |
+
"n": 1000,
|
| 465 |
+
"accuracy": 0.889,
|
| 466 |
+
"nll": 0.2843,
|
| 467 |
+
"ece": 0.0287,
|
| 468 |
+
"mean_conf": 0.8713
|
| 469 |
+
},
|
| 470 |
+
"by_type": {
|
| 471 |
+
"choice": {
|
| 472 |
+
"n": 1000,
|
| 473 |
+
"accuracy": 0.889,
|
| 474 |
+
"nll": 0.2843,
|
| 475 |
+
"ece": 0.0287,
|
| 476 |
+
"mean_conf": 0.8713
|
| 477 |
+
},
|
| 478 |
+
"noul": null,
|
| 479 |
+
"score": null
|
| 480 |
}
|
| 481 |
+
}
|
| 482 |
+
}
|
| 483 |
+
},
|
| 484 |
+
"fixtures": {
|
| 485 |
+
"general_validation": {
|
| 486 |
+
"this": {
|
| 487 |
+
"rows": 847,
|
| 488 |
+
"accuracy": 0.8500590318772137,
|
| 489 |
+
"macro_accuracy": 0.8552272255409907,
|
| 490 |
+
"groups": 65,
|
| 491 |
+
"nll": 0.43237405419192004,
|
| 492 |
+
"ece": 0.03505942904541031,
|
| 493 |
+
"brier": 0.22353081852050935,
|
| 494 |
+
"mean_confidence": 0.8244136475475226
|
| 495 |
},
|
| 496 |
+
"paired": {
|
| 497 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2": {
|
| 498 |
+
"accuracy": {
|
| 499 |
+
"delta": 0.0,
|
| 500 |
+
"ci95": [
|
| 501 |
+
0.0,
|
| 502 |
+
0.0
|
| 503 |
+
]
|
| 504 |
+
},
|
| 505 |
+
"nll": {
|
| 506 |
+
"delta": 0.00012970495542040443,
|
| 507 |
+
"ci95": [
|
| 508 |
+
-0.00039863639304211823,
|
| 509 |
+
0.0005942988193083918
|
| 510 |
+
]
|
| 511 |
+
}
|
| 512 |
+
},
|
| 513 |
+
"cand4b_B4_e2_tt-v1_4b_p130": {
|
| 514 |
+
"accuracy": {
|
| 515 |
+
"delta": -0.010625737898465172,
|
| 516 |
+
"ci95": [
|
| 517 |
+
-0.021251475796930343,
|
| 518 |
+
0.0
|
| 519 |
+
]
|
| 520 |
+
},
|
| 521 |
+
"nll": {
|
| 522 |
+
"delta": 0.015307099810435664,
|
| 523 |
+
"ci95": [
|
| 524 |
+
0.0030912070436396634,
|
| 525 |
+
0.028028207832146323
|
| 526 |
+
]
|
| 527 |
+
}
|
| 528 |
+
},
|
| 529 |
+
"cand4b_B4_e2_tt-v2_4b_p130": {
|
| 530 |
+
"accuracy": {
|
| 531 |
+
"delta": 0.0,
|
| 532 |
+
"ci95": [
|
| 533 |
+
-0.015348288075560802,
|
| 534 |
+
0.015348288075560802
|
| 535 |
+
]
|
| 536 |
+
},
|
| 537 |
+
"nll": {
|
| 538 |
+
"delta": 0.013765296214476044,
|
| 539 |
+
"ci95": [
|
| 540 |
+
-0.006225581100979938,
|
| 541 |
+
0.0345539752628619
|
| 542 |
+
]
|
| 543 |
+
}
|
| 544 |
}
|
| 545 |
}
|
| 546 |
},
|
| 547 |
+
"typesafe": {
|
| 548 |
+
"this": {
|
| 549 |
+
"rows": 102,
|
| 550 |
+
"accuracy": 0.8431372549019608,
|
| 551 |
+
"macro_accuracy": 0.8464696223316912,
|
| 552 |
+
"groups": 4,
|
| 553 |
+
"nll": 0.44128428424617644,
|
| 554 |
+
"ece": 0.06267209380280735,
|
| 555 |
+
"brier": 0.24132394423929263,
|
| 556 |
+
"mean_confidence": 0.8996653276331285,
|
| 557 |
+
"mean_total_variation_to_reference": 0.19246147475275616
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 558 |
},
|
| 559 |
+
"paired": {
|
| 560 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2": {
|
| 561 |
+
"accuracy": {
|
| 562 |
+
"delta": 0.0,
|
| 563 |
+
"ci95": [
|
| 564 |
+
0.0,
|
| 565 |
+
0.0
|
| 566 |
+
]
|
| 567 |
+
},
|
| 568 |
+
"nll": {
|
| 569 |
+
"delta": -0.0017070072924830098,
|
| 570 |
+
"ci95": [
|
| 571 |
+
-0.003544773928864442,
|
| 572 |
+
-8.629601238631066e-05
|
| 573 |
+
]
|
| 574 |
+
}
|
| 575 |
+
},
|
| 576 |
+
"cand4b_B4_e2_tt-v1_4b_p130": {
|
| 577 |
+
"accuracy": {
|
| 578 |
+
"delta": 0.029411764705882353,
|
| 579 |
+
"ci95": [
|
| 580 |
+
-0.029411764705882353,
|
| 581 |
+
0.08823529411764706
|
| 582 |
+
]
|
| 583 |
+
},
|
| 584 |
+
"nll": {
|
| 585 |
+
"delta": -0.16956755666314244,
|
| 586 |
+
"ci95": [
|
| 587 |
+
-0.3236270765220259,
|
| 588 |
+
-0.035752501984007506
|
| 589 |
+
]
|
| 590 |
+
}
|
| 591 |
+
},
|
| 592 |
+
"cand4b_B4_e2_tt-v2_4b_p130": {
|
| 593 |
+
"accuracy": {
|
| 594 |
+
"delta": -0.0196078431372549,
|
| 595 |
+
"ci95": [
|
| 596 |
+
-0.0784313725490196,
|
| 597 |
+
0.0392156862745098
|
| 598 |
+
]
|
| 599 |
+
},
|
| 600 |
+
"nll": {
|
| 601 |
+
"delta": 0.03116252024698791,
|
| 602 |
+
"ci95": [
|
| 603 |
+
-0.06909971743191749,
|
| 604 |
+
0.13434316002254298
|
| 605 |
+
]
|
| 606 |
+
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 607 |
}
|
| 608 |
}
|
| 609 |
},
|
| 610 |
+
"openjev_external": {
|
| 611 |
+
"this": {
|
| 612 |
+
"rows": 5252,
|
| 613 |
+
"accuracy": 0.6597486671744097,
|
| 614 |
+
"macro_accuracy": 0.8063444444444444,
|
| 615 |
+
"groups": 3,
|
| 616 |
+
"nll": 0.8457289344223444,
|
| 617 |
+
"ece": 0.1539857582574297,
|
| 618 |
+
"brier": 0.5022800176356607,
|
| 619 |
+
"mean_confidence": 0.8111537299704116
|
|
|
|
|
|
|
|
|
|
| 620 |
},
|
| 621 |
+
"paired": {
|
| 622 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2": {
|
| 623 |
+
"accuracy": {
|
| 624 |
+
"delta": 0.0,
|
| 625 |
+
"ci95": [
|
| 626 |
+
0.0,
|
| 627 |
+
0.0
|
| 628 |
+
]
|
| 629 |
+
},
|
| 630 |
+
"nll": {
|
| 631 |
+
"delta": -0.003617502816960803,
|
| 632 |
+
"ci95": [
|
| 633 |
+
-0.003919285763621306,
|
| 634 |
+
-0.00332506503930084
|
| 635 |
+
]
|
| 636 |
+
}
|
| 637 |
+
},
|
| 638 |
+
"cand4b_B4_e2_tt-v1_4b_p130": {
|
| 639 |
+
"accuracy": {
|
| 640 |
+
"delta": 0.020944402132520943,
|
| 641 |
+
"ci95": [
|
| 642 |
+
0.012376237623762377,
|
| 643 |
+
0.029322162985529324
|
| 644 |
+
]
|
| 645 |
+
},
|
| 646 |
+
"nll": {
|
| 647 |
+
"delta": -0.047333755296254304,
|
| 648 |
+
"ci95": [
|
| 649 |
+
-0.059559262901540976,
|
| 650 |
+
-0.035193633735333484
|
| 651 |
+
]
|
| 652 |
+
}
|
| 653 |
+
},
|
| 654 |
+
"cand4b_B4_e2_tt-v2_4b_p130": {
|
| 655 |
+
"accuracy": {
|
| 656 |
+
"delta": -0.007425742574257425,
|
| 657 |
+
"ci95": [
|
| 658 |
+
-0.015993907083015995,
|
| 659 |
+
0.0017136329017517135
|
| 660 |
+
]
|
| 661 |
+
},
|
| 662 |
+
"nll": {
|
| 663 |
+
"delta": 0.05671352997929895,
|
| 664 |
+
"ci95": [
|
| 665 |
+
0.04359541016972144,
|
| 666 |
+
0.07033520922345497
|
| 667 |
+
]
|
| 668 |
+
}
|
| 669 |
+
}
|
| 670 |
+
}
|
| 671 |
+
},
|
| 672 |
+
"mind2web": {
|
| 673 |
+
"this": {
|
| 674 |
+
"rows": 1770,
|
| 675 |
+
"accuracy": 0.8751412429378531,
|
| 676 |
+
"macro_accuracy": 0.8773196126201471,
|
| 677 |
+
"groups": 3,
|
| 678 |
+
"nll": 0.37523612083453944,
|
| 679 |
+
"ece": 0.016833950116135952,
|
| 680 |
+
"brier": 0.18402882486455008,
|
| 681 |
+
"mean_confidence": 0.8624463401776923
|
| 682 |
},
|
| 683 |
+
"paired": {
|
| 684 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2": {
|
| 685 |
+
"accuracy": {
|
| 686 |
+
"delta": 0.0,
|
| 687 |
+
"ci95": [
|
| 688 |
+
0.0,
|
| 689 |
+
0.0
|
| 690 |
+
]
|
| 691 |
+
},
|
| 692 |
+
"nll": {
|
| 693 |
+
"delta": 0.00026909733615400604,
|
| 694 |
+
"ci95": [
|
| 695 |
+
-4.997823592689937e-05,
|
| 696 |
+
0.0006061414472582659
|
| 697 |
+
]
|
| 698 |
+
}
|
| 699 |
+
},
|
| 700 |
+
"cand4b_B4_e2_tt-v1_4b_p130": {
|
| 701 |
+
"accuracy": {
|
| 702 |
+
"delta": -0.00847457627118644,
|
| 703 |
+
"ci95": [
|
| 704 |
+
-0.01694915254237288,
|
| 705 |
+
0.0
|
| 706 |
+
]
|
| 707 |
+
},
|
| 708 |
+
"nll": {
|
| 709 |
+
"delta": 0.00924747244030173,
|
| 710 |
+
"ci95": [
|
| 711 |
+
-0.004092843962589085,
|
| 712 |
+
0.022512164121517977
|
| 713 |
+
]
|
| 714 |
+
}
|
| 715 |
+
},
|
| 716 |
+
"cand4b_B4_e2_tt-v2_4b_p130": {
|
| 717 |
+
"accuracy": {
|
| 718 |
+
"delta": 0.0011299435028248588,
|
| 719 |
+
"ci95": [
|
| 720 |
+
-0.00847457627118644,
|
| 721 |
+
0.011299435028248588
|
| 722 |
+
]
|
| 723 |
+
},
|
| 724 |
+
"nll": {
|
| 725 |
+
"delta": -0.018244809747118285,
|
| 726 |
+
"ci95": [
|
| 727 |
+
-0.032040489157299956,
|
| 728 |
+
-0.005027723688309381
|
| 729 |
+
]
|
| 730 |
+
}
|
| 731 |
+
}
|
| 732 |
+
}
|
| 733 |
+
}
|
| 734 |
+
},
|
| 735 |
+
"games": {
|
| 736 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|sampled|all": {
|
| 737 |
+
"n": 234,
|
| 738 |
+
"a": 0.2692307692307692,
|
| 739 |
+
"b": 0.2692307692307692,
|
| 740 |
+
"delta": 0.0,
|
| 741 |
+
"ci95": [
|
| 742 |
+
-0.004273504273504274,
|
| 743 |
+
0.004273504273504274
|
| 744 |
+
]
|
| 745 |
+
},
|
| 746 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|sampled|bag_game": {
|
| 747 |
+
"n": 64,
|
| 748 |
+
"a": 0.51953125,
|
| 749 |
+
"b": 0.51953125,
|
| 750 |
+
"delta": 0.0,
|
| 751 |
+
"ci95": [
|
| 752 |
+
-0.01171875,
|
| 753 |
+
0.01171875
|
| 754 |
+
]
|
| 755 |
+
},
|
| 756 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|sampled|stochastic_grid": {
|
| 757 |
+
"n": 64,
|
| 758 |
+
"a": 0.16796875,
|
| 759 |
+
"b": 0.16796875,
|
| 760 |
+
"delta": 0.0,
|
| 761 |
+
"ci95": [
|
| 762 |
+
-0.01171875,
|
| 763 |
+
0.01171875
|
| 764 |
+
]
|
| 765 |
+
},
|
| 766 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|sampled|tic_tac_toe": {
|
| 767 |
+
"n": 74,
|
| 768 |
+
"a": 0.2533783783783784,
|
| 769 |
+
"b": 0.2533783783783784,
|
| 770 |
+
"delta": 0.0,
|
| 771 |
+
"ci95": [
|
| 772 |
+
0.0,
|
| 773 |
+
0.0
|
| 774 |
+
]
|
| 775 |
+
},
|
| 776 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|sampled|exact_minesweeper": {
|
| 777 |
+
"n": 32,
|
| 778 |
+
"a": 0.0078125,
|
| 779 |
+
"b": 0.0078125,
|
| 780 |
+
"delta": 0.0,
|
| 781 |
+
"ci95": [
|
| 782 |
+
0.0,
|
| 783 |
+
0.0
|
| 784 |
+
]
|
| 785 |
+
},
|
| 786 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|greedy|all": {
|
| 787 |
+
"n": 234,
|
| 788 |
+
"a": 0.25213675213675213,
|
| 789 |
+
"b": 0.25213675213675213,
|
| 790 |
+
"delta": 0.0,
|
| 791 |
+
"ci95": [
|
| 792 |
+
0.0,
|
| 793 |
+
0.0
|
| 794 |
+
]
|
| 795 |
+
},
|
| 796 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|greedy|bag_game": {
|
| 797 |
+
"n": 64,
|
| 798 |
+
"a": 0.53125,
|
| 799 |
+
"b": 0.53125,
|
| 800 |
+
"delta": 0.0,
|
| 801 |
+
"ci95": [
|
| 802 |
+
0.0,
|
| 803 |
+
0.0
|
| 804 |
+
]
|
| 805 |
+
},
|
| 806 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|greedy|stochastic_grid": {
|
| 807 |
+
"n": 64,
|
| 808 |
+
"a": 0.109375,
|
| 809 |
+
"b": 0.109375,
|
| 810 |
+
"delta": 0.0,
|
| 811 |
+
"ci95": [
|
| 812 |
+
0.0,
|
| 813 |
+
0.0
|
| 814 |
+
]
|
| 815 |
+
},
|
| 816 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|greedy|tic_tac_toe": {
|
| 817 |
+
"n": 74,
|
| 818 |
+
"a": 0.24324324324324326,
|
| 819 |
+
"b": 0.24324324324324326,
|
| 820 |
+
"delta": 0.0,
|
| 821 |
+
"ci95": [
|
| 822 |
+
0.0,
|
| 823 |
+
0.0
|
| 824 |
+
]
|
| 825 |
+
},
|
| 826 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|greedy|exact_minesweeper": {
|
| 827 |
+
"n": 32,
|
| 828 |
+
"a": 0.0,
|
| 829 |
+
"b": 0.0,
|
| 830 |
+
"delta": 0.0,
|
| 831 |
+
"ci95": [
|
| 832 |
+
0.0,
|
| 833 |
+
0.0
|
| 834 |
+
]
|
| 835 |
+
},
|
| 836 |
+
"cand4b_B4_e2_tt-v1_4b_p130|sampled|all": {
|
| 837 |
+
"n": 234,
|
| 838 |
+
"a": 0.2692307692307692,
|
| 839 |
+
"b": 0.2777777777777778,
|
| 840 |
+
"delta": -0.008547008547008548,
|
| 841 |
+
"ci95": [
|
| 842 |
+
-0.02564102564102564,
|
| 843 |
+
0.009615384615384616
|
| 844 |
+
]
|
| 845 |
+
},
|
| 846 |
+
"cand4b_B4_e2_tt-v1_4b_p130|sampled|bag_game": {
|
| 847 |
+
"n": 64,
|
| 848 |
+
"a": 0.51953125,
|
| 849 |
+
"b": 0.56640625,
|
| 850 |
+
"delta": -0.046875,
|
| 851 |
+
"ci95": [
|
| 852 |
+
-0.08984375,
|
| 853 |
+
-0.00390625
|
| 854 |
+
]
|
| 855 |
+
},
|
| 856 |
+
"cand4b_B4_e2_tt-v1_4b_p130|sampled|stochastic_grid": {
|
| 857 |
+
"n": 64,
|
| 858 |
+
"a": 0.16796875,
|
| 859 |
+
"b": 0.1640625,
|
| 860 |
+
"delta": 0.00390625,
|
| 861 |
+
"ci95": [
|
| 862 |
+
-0.02734375,
|
| 863 |
+
0.03515625
|
| 864 |
+
]
|
| 865 |
+
},
|
| 866 |
+
"cand4b_B4_e2_tt-v1_4b_p130|sampled|tic_tac_toe": {
|
| 867 |
+
"n": 74,
|
| 868 |
+
"a": 0.2533783783783784,
|
| 869 |
+
"b": 0.24324324324324326,
|
| 870 |
+
"delta": 0.010135135135135136,
|
| 871 |
+
"ci95": [
|
| 872 |
+
-0.013513513513513514,
|
| 873 |
+
0.037162162162162164
|
| 874 |
+
]
|
| 875 |
+
},
|
| 876 |
+
"cand4b_B4_e2_tt-v1_4b_p130|sampled|exact_minesweeper": {
|
| 877 |
+
"n": 32,
|
| 878 |
+
"a": 0.0078125,
|
| 879 |
+
"b": 0.0078125,
|
| 880 |
+
"delta": 0.0,
|
| 881 |
+
"ci95": [
|
| 882 |
+
-0.0234375,
|
| 883 |
+
0.0234375
|
| 884 |
+
]
|
| 885 |
+
},
|
| 886 |
+
"cand4b_B4_e2_tt-v1_4b_p130|greedy|all": {
|
| 887 |
+
"n": 234,
|
| 888 |
+
"a": 0.25213675213675213,
|
| 889 |
+
"b": 0.2905982905982906,
|
| 890 |
+
"delta": -0.038461538461538464,
|
| 891 |
+
"ci95": [
|
| 892 |
+
-0.07264957264957266,
|
| 893 |
+
-0.004273504273504274
|
| 894 |
+
]
|
| 895 |
+
},
|
| 896 |
+
"cand4b_B4_e2_tt-v1_4b_p130|greedy|bag_game": {
|
| 897 |
+
"n": 64,
|
| 898 |
+
"a": 0.53125,
|
| 899 |
+
"b": 0.625,
|
| 900 |
+
"delta": -0.09375,
|
| 901 |
+
"ci95": [
|
| 902 |
+
-0.171875,
|
| 903 |
+
-0.03125
|
| 904 |
+
]
|
| 905 |
+
},
|
| 906 |
+
"cand4b_B4_e2_tt-v1_4b_p130|greedy|stochastic_grid": {
|
| 907 |
+
"n": 64,
|
| 908 |
+
"a": 0.109375,
|
| 909 |
+
"b": 0.125,
|
| 910 |
+
"delta": -0.015625,
|
| 911 |
+
"ci95": [
|
| 912 |
+
-0.07851562499999987,
|
| 913 |
+
0.046875
|
| 914 |
+
]
|
| 915 |
+
},
|
| 916 |
+
"cand4b_B4_e2_tt-v1_4b_p130|greedy|tic_tac_toe": {
|
| 917 |
+
"n": 74,
|
| 918 |
+
"a": 0.24324324324324326,
|
| 919 |
+
"b": 0.2702702702702703,
|
| 920 |
+
"delta": -0.02702702702702703,
|
| 921 |
+
"ci95": [
|
| 922 |
+
-0.0945945945945946,
|
| 923 |
+
0.04054054054054054
|
| 924 |
+
]
|
| 925 |
+
},
|
| 926 |
+
"cand4b_B4_e2_tt-v1_4b_p130|greedy|exact_minesweeper": {
|
| 927 |
+
"n": 32,
|
| 928 |
+
"a": 0.0,
|
| 929 |
+
"b": 0.0,
|
| 930 |
+
"delta": 0.0,
|
| 931 |
+
"ci95": [
|
| 932 |
+
0.0,
|
| 933 |
+
0.0
|
| 934 |
+
]
|
| 935 |
+
},
|
| 936 |
+
"cand4b_B4_e2_tt-v2_4b_p130|sampled|all": {
|
| 937 |
+
"n": 234,
|
| 938 |
+
"a": 0.2692307692307692,
|
| 939 |
+
"b": 0.22435897435897437,
|
| 940 |
+
"delta": 0.04487179487179487,
|
| 941 |
+
"ci95": [
|
| 942 |
+
0.024572649572649572,
|
| 943 |
+
0.06623931623931624
|
| 944 |
+
]
|
| 945 |
+
},
|
| 946 |
+
"cand4b_B4_e2_tt-v2_4b_p130|sampled|bag_game": {
|
| 947 |
+
"n": 64,
|
| 948 |
+
"a": 0.51953125,
|
| 949 |
+
"b": 0.37890625,
|
| 950 |
+
"delta": 0.140625,
|
| 951 |
+
"ci95": [
|
| 952 |
+
0.08203125,
|
| 953 |
+
0.19921875
|
| 954 |
+
]
|
| 955 |
+
},
|
| 956 |
+
"cand4b_B4_e2_tt-v2_4b_p130|sampled|stochastic_grid": {
|
| 957 |
+
"n": 64,
|
| 958 |
+
"a": 0.16796875,
|
| 959 |
+
"b": 0.16015625,
|
| 960 |
+
"delta": 0.0078125,
|
| 961 |
+
"ci95": [
|
| 962 |
+
-0.0234375,
|
| 963 |
+
0.0390625
|
| 964 |
+
]
|
| 965 |
+
},
|
| 966 |
+
"cand4b_B4_e2_tt-v2_4b_p130|sampled|tic_tac_toe": {
|
| 967 |
+
"n": 74,
|
| 968 |
+
"a": 0.2533783783783784,
|
| 969 |
+
"b": 0.24324324324324326,
|
| 970 |
+
"delta": 0.010135135135135136,
|
| 971 |
+
"ci95": [
|
| 972 |
+
-0.010135135135135136,
|
| 973 |
+
0.033783783783783786
|
| 974 |
+
]
|
| 975 |
+
},
|
| 976 |
+
"cand4b_B4_e2_tt-v2_4b_p130|sampled|exact_minesweeper": {
|
| 977 |
+
"n": 32,
|
| 978 |
+
"a": 0.0078125,
|
| 979 |
+
"b": 0.0,
|
| 980 |
+
"delta": 0.0078125,
|
| 981 |
+
"ci95": [
|
| 982 |
+
0.0,
|
| 983 |
+
0.0234375
|
| 984 |
+
]
|
| 985 |
+
},
|
| 986 |
+
"cand4b_B4_e2_tt-v2_4b_p130|greedy|all": {
|
| 987 |
+
"n": 234,
|
| 988 |
+
"a": 0.25213675213675213,
|
| 989 |
+
"b": 0.2777777777777778,
|
| 990 |
+
"delta": -0.02564102564102564,
|
| 991 |
+
"ci95": [
|
| 992 |
+
-0.0641025641025641,
|
| 993 |
+
0.01282051282051282
|
| 994 |
+
]
|
| 995 |
+
},
|
| 996 |
+
"cand4b_B4_e2_tt-v2_4b_p130|greedy|bag_game": {
|
| 997 |
+
"n": 64,
|
| 998 |
+
"a": 0.53125,
|
| 999 |
+
"b": 0.484375,
|
| 1000 |
+
"delta": 0.046875,
|
| 1001 |
+
"ci95": [
|
| 1002 |
+
-0.03125,
|
| 1003 |
+
0.125
|
| 1004 |
+
]
|
| 1005 |
+
},
|
| 1006 |
+
"cand4b_B4_e2_tt-v2_4b_p130|greedy|stochastic_grid": {
|
| 1007 |
+
"n": 64,
|
| 1008 |
+
"a": 0.109375,
|
| 1009 |
+
"b": 0.1875,
|
| 1010 |
+
"delta": -0.078125,
|
| 1011 |
+
"ci95": [
|
| 1012 |
+
-0.15625,
|
| 1013 |
+
0.0
|
| 1014 |
+
]
|
| 1015 |
+
},
|
| 1016 |
+
"cand4b_B4_e2_tt-v2_4b_p130|greedy|tic_tac_toe": {
|
| 1017 |
+
"n": 74,
|
| 1018 |
+
"a": 0.24324324324324326,
|
| 1019 |
+
"b": 0.2972972972972973,
|
| 1020 |
+
"delta": -0.05405405405405406,
|
| 1021 |
+
"ci95": [
|
| 1022 |
+
-0.12162162162162163,
|
| 1023 |
+
0.02702702702702703
|
| 1024 |
+
]
|
| 1025 |
+
},
|
| 1026 |
+
"cand4b_B4_e2_tt-v2_4b_p130|greedy|exact_minesweeper": {
|
| 1027 |
+
"n": 32,
|
| 1028 |
+
"a": 0.0,
|
| 1029 |
+
"b": 0.0,
|
| 1030 |
+
"delta": 0.0,
|
| 1031 |
+
"ci95": [
|
| 1032 |
+
0.0,
|
| 1033 |
+
0.0
|
| 1034 |
+
]
|
| 1035 |
+
}
|
| 1036 |
+
},
|
| 1037 |
+
"browser": {
|
| 1038 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|sampled|all": {
|
| 1039 |
+
"n": 176,
|
| 1040 |
+
"a": 0.9318181818181818,
|
| 1041 |
+
"b": 0.9318181818181818,
|
| 1042 |
+
"delta": 0.0,
|
| 1043 |
+
"ci95": [
|
| 1044 |
+
0.0,
|
| 1045 |
+
0.0
|
| 1046 |
+
]
|
| 1047 |
+
},
|
| 1048 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|sampled|rewarded": {
|
| 1049 |
+
"n": 128,
|
| 1050 |
+
"a": 0.953125,
|
| 1051 |
+
"b": 0.953125,
|
| 1052 |
+
"delta": 0.0,
|
| 1053 |
+
"ci95": [
|
| 1054 |
+
0.0,
|
| 1055 |
+
0.0
|
| 1056 |
+
]
|
| 1057 |
+
},
|
| 1058 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|sampled|held_out": {
|
| 1059 |
+
"n": 48,
|
| 1060 |
+
"a": 0.875,
|
| 1061 |
+
"b": 0.875,
|
| 1062 |
+
"delta": 0.0,
|
| 1063 |
+
"ci95": [
|
| 1064 |
+
0.0,
|
| 1065 |
+
0.0
|
| 1066 |
+
]
|
| 1067 |
+
},
|
| 1068 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|greedy|all": {
|
| 1069 |
+
"n": 176,
|
| 1070 |
+
"a": 0.9375,
|
| 1071 |
+
"b": 0.9375,
|
| 1072 |
+
"delta": 0.0,
|
| 1073 |
+
"ci95": [
|
| 1074 |
+
0.0,
|
| 1075 |
+
0.0
|
| 1076 |
+
]
|
| 1077 |
+
},
|
| 1078 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|greedy|rewarded": {
|
| 1079 |
+
"n": 128,
|
| 1080 |
+
"a": 0.9609375,
|
| 1081 |
+
"b": 0.9609375,
|
| 1082 |
+
"delta": 0.0,
|
| 1083 |
+
"ci95": [
|
| 1084 |
+
0.0,
|
| 1085 |
+
0.0
|
| 1086 |
+
]
|
| 1087 |
+
},
|
| 1088 |
+
"cand4b_B4_e2_tt-cand4b_B4_e2|greedy|held_out": {
|
| 1089 |
+
"n": 48,
|
| 1090 |
+
"a": 0.875,
|
| 1091 |
+
"b": 0.875,
|
| 1092 |
+
"delta": 0.0,
|
| 1093 |
+
"ci95": [
|
| 1094 |
+
0.0,
|
| 1095 |
+
0.0
|
| 1096 |
+
]
|
| 1097 |
+
},
|
| 1098 |
+
"cand4b_B4_e2_tt-v1_4b_p130|sampled|all": {
|
| 1099 |
+
"n": 176,
|
| 1100 |
+
"a": 0.9318181818181818,
|
| 1101 |
+
"b": 0.9090909090909091,
|
| 1102 |
+
"delta": 0.022727272727272728,
|
| 1103 |
+
"ci95": [
|
| 1104 |
+
-0.017045454545454544,
|
| 1105 |
+
0.0625
|
| 1106 |
+
]
|
| 1107 |
+
},
|
| 1108 |
+
"cand4b_B4_e2_tt-v1_4b_p130|sampled|rewarded": {
|
| 1109 |
+
"n": 128,
|
| 1110 |
+
"a": 0.953125,
|
| 1111 |
+
"b": 0.9609375,
|
| 1112 |
+
"delta": -0.0078125,
|
| 1113 |
+
"ci95": [
|
| 1114 |
+
-0.0390625,
|
| 1115 |
+
0.0234375
|
| 1116 |
+
]
|
| 1117 |
+
},
|
| 1118 |
+
"cand4b_B4_e2_tt-v1_4b_p130|sampled|held_out": {
|
| 1119 |
+
"n": 48,
|
| 1120 |
+
"a": 0.875,
|
| 1121 |
+
"b": 0.7708333333333334,
|
| 1122 |
+
"delta": 0.10416666666666667,
|
| 1123 |
+
"ci95": [
|
| 1124 |
+
0.0,
|
| 1125 |
+
0.20833333333333334
|
| 1126 |
+
]
|
| 1127 |
+
},
|
| 1128 |
+
"cand4b_B4_e2_tt-v1_4b_p130|greedy|all": {
|
| 1129 |
+
"n": 176,
|
| 1130 |
+
"a": 0.9375,
|
| 1131 |
+
"b": 0.9147727272727273,
|
| 1132 |
+
"delta": 0.022727272727272728,
|
| 1133 |
+
"ci95": [
|
| 1134 |
+
-0.005681818181818182,
|
| 1135 |
+
0.056818181818181816
|
| 1136 |
+
]
|
| 1137 |
+
},
|
| 1138 |
+
"cand4b_B4_e2_tt-v1_4b_p130|greedy|rewarded": {
|
| 1139 |
+
"n": 128,
|
| 1140 |
+
"a": 0.9609375,
|
| 1141 |
+
"b": 0.9765625,
|
| 1142 |
+
"delta": -0.015625,
|
| 1143 |
+
"ci95": [
|
| 1144 |
+
-0.0390625,
|
| 1145 |
+
0.0
|
| 1146 |
+
]
|
| 1147 |
+
},
|
| 1148 |
+
"cand4b_B4_e2_tt-v1_4b_p130|greedy|held_out": {
|
| 1149 |
+
"n": 48,
|
| 1150 |
+
"a": 0.875,
|
| 1151 |
+
"b": 0.75,
|
| 1152 |
+
"delta": 0.125,
|
| 1153 |
+
"ci95": [
|
| 1154 |
+
0.041666666666666664,
|
| 1155 |
+
0.22916666666666666
|
| 1156 |
+
]
|
| 1157 |
+
},
|
| 1158 |
+
"cand4b_B4_e2_tt-v2_4b_p130|sampled|all": {
|
| 1159 |
+
"n": 176,
|
| 1160 |
+
"a": 0.9318181818181818,
|
| 1161 |
+
"b": 0.8806818181818182,
|
| 1162 |
+
"delta": 0.05113636363636364,
|
| 1163 |
+
"ci95": [
|
| 1164 |
+
0.005681818181818182,
|
| 1165 |
+
0.09659090909090909
|
| 1166 |
+
]
|
| 1167 |
+
},
|
| 1168 |
+
"cand4b_B4_e2_tt-v2_4b_p130|sampled|rewarded": {
|
| 1169 |
+
"n": 128,
|
| 1170 |
+
"a": 0.953125,
|
| 1171 |
+
"b": 0.9296875,
|
| 1172 |
+
"delta": 0.0234375,
|
| 1173 |
+
"ci95": [
|
| 1174 |
+
-0.0234375,
|
| 1175 |
+
0.0703125
|
| 1176 |
+
]
|
| 1177 |
+
},
|
| 1178 |
+
"cand4b_B4_e2_tt-v2_4b_p130|sampled|held_out": {
|
| 1179 |
+
"n": 48,
|
| 1180 |
+
"a": 0.875,
|
| 1181 |
+
"b": 0.75,
|
| 1182 |
+
"delta": 0.125,
|
| 1183 |
+
"ci95": [
|
| 1184 |
+
0.0,
|
| 1185 |
+
0.25
|
| 1186 |
+
]
|
| 1187 |
+
},
|
| 1188 |
+
"cand4b_B4_e2_tt-v2_4b_p130|greedy|all": {
|
| 1189 |
+
"n": 176,
|
| 1190 |
+
"a": 0.9375,
|
| 1191 |
+
"b": 0.9261363636363636,
|
| 1192 |
+
"delta": 0.011363636363636364,
|
| 1193 |
+
"ci95": [
|
| 1194 |
+
-0.017045454545454544,
|
| 1195 |
+
0.045454545454545456
|
| 1196 |
+
]
|
| 1197 |
+
},
|
| 1198 |
+
"cand4b_B4_e2_tt-v2_4b_p130|greedy|rewarded": {
|
| 1199 |
+
"n": 128,
|
| 1200 |
+
"a": 0.9609375,
|
| 1201 |
+
"b": 0.96875,
|
| 1202 |
+
"delta": -0.0078125,
|
| 1203 |
+
"ci95": [
|
| 1204 |
+
-0.0234375,
|
| 1205 |
+
0.0
|
| 1206 |
+
]
|
| 1207 |
+
},
|
| 1208 |
+
"cand4b_B4_e2_tt-v2_4b_p130|greedy|held_out": {
|
| 1209 |
+
"n": 48,
|
| 1210 |
+
"a": 0.875,
|
| 1211 |
+
"b": 0.8125,
|
| 1212 |
+
"delta": 0.0625,
|
| 1213 |
+
"ci95": [
|
| 1214 |
+
-0.041666666666666664,
|
| 1215 |
+
0.16666666666666666
|
| 1216 |
+
]
|
| 1217 |
+
}
|
| 1218 |
+
},
|
| 1219 |
+
"text_games_greedy": {
|
| 1220 |
+
"pong": -5.0,
|
| 1221 |
+
"freeway": 0.0,
|
| 1222 |
+
"breakout": 35.0,
|
| 1223 |
+
"frozenlake": 0.0,
|
| 1224 |
+
"cliffwalking": -13.0,
|
| 1225 |
+
"blackjack": -0.6,
|
| 1226 |
+
"minigrid_empty": 0.0,
|
| 1227 |
+
"minigrid_lavagap": 0.0,
|
| 1228 |
+
"minigrid_doorkey": 0.0,
|
| 1229 |
+
"babyai_goto": 0.1915625
|
| 1230 |
+
},
|
| 1231 |
+
"probes": {
|
| 1232 |
+
"router_n": 31,
|
| 1233 |
+
"router_tier": 0.968,
|
| 1234 |
+
"needs_live_data": 0.839,
|
| 1235 |
+
"commands_n": 45,
|
| 1236 |
+
"command_risk": 0.889,
|
| 1237 |
+
"destructive_called_safe": 0,
|
| 1238 |
+
"outside_project": 0.933,
|
| 1239 |
+
"browser_element": 0.938,
|
| 1240 |
+
"browser_action": 0.75,
|
| 1241 |
+
"batteries": {
|
| 1242 |
+
"generic": 1.0,
|
| 1243 |
+
"specific": 1.0,
|
| 1244 |
+
"catchall": 0.95,
|
| 1245 |
+
"noul_acc": 1.0,
|
| 1246 |
+
"noul_brier": 0.01,
|
| 1247 |
+
"n_choice": 60,
|
| 1248 |
+
"n_noul": 49,
|
| 1249 |
+
"abstain_battery": "8/8"
|
| 1250 |
+
}
|
| 1251 |
+
},
|
| 1252 |
+
"issue9": {
|
| 1253 |
+
"c_1": {
|
| 1254 |
+
"pred": 26,
|
| 1255 |
+
"gold": 6,
|
| 1256 |
+
"ok": false,
|
| 1257 |
+
"p_gold": 0.198,
|
| 1258 |
+
"p_pred": 0.764,
|
| 1259 |
+
"T": 1.099
|
| 1260 |
+
},
|
| 1261 |
+
"c_2": {
|
| 1262 |
+
"pred": 20,
|
| 1263 |
+
"gold": 20,
|
| 1264 |
+
"ok": true,
|
| 1265 |
+
"p_gold": 0.409,
|
| 1266 |
+
"p_pred": 0.409,
|
| 1267 |
+
"T": 1.099
|
| 1268 |
+
},
|
| 1269 |
+
"b_10": {
|
| 1270 |
+
"pred": 0,
|
| 1271 |
+
"gold": 0,
|
| 1272 |
+
"ok": true,
|
| 1273 |
+
"p_gold": 0.732,
|
| 1274 |
+
"p_pred": 0.732,
|
| 1275 |
+
"T": 1.099
|
| 1276 |
+
},
|
| 1277 |
+
"a2_4": {
|
| 1278 |
+
"pred": 4,
|
| 1279 |
+
"gold": 4,
|
| 1280 |
+
"ok": true,
|
| 1281 |
+
"p_gold": 0.577,
|
| 1282 |
+
"p_pred": 0.577,
|
| 1283 |
+
"T": 1.099
|
| 1284 |
+
}
|
| 1285 |
+
},
|
| 1286 |
+
"bespoke": {
|
| 1287 |
+
"vitaminc-dev": {
|
| 1288 |
+
"n": 599,
|
| 1289 |
+
"type": "choice",
|
| 1290 |
+
"acc": 0.7429,
|
| 1291 |
+
"ece": 0.1456,
|
| 1292 |
+
"brier": 0.3945,
|
| 1293 |
+
"nll": 0.7789,
|
| 1294 |
+
"trained": false,
|
| 1295 |
+
"sec": 11.8
|
| 1296 |
+
},
|
| 1297 |
+
"massive-en-US": {
|
| 1298 |
+
"n": 350,
|
| 1299 |
+
"type": "choice",
|
| 1300 |
+
"acc": 0.8629,
|
| 1301 |
+
"ece": 0.0411,
|
| 1302 |
+
"brier": 0.1936,
|
| 1303 |
+
"nll": 0.4173,
|
| 1304 |
+
"trained": true,
|
| 1305 |
+
"sec": 6.2
|
| 1306 |
+
},
|
| 1307 |
+
"massive-de-DE": {
|
| 1308 |
+
"n": 350,
|
| 1309 |
+
"type": "choice",
|
| 1310 |
+
"acc": 0.84,
|
| 1311 |
+
"ece": 0.0297,
|
| 1312 |
+
"brier": 0.2269,
|
| 1313 |
+
"nll": 0.4947,
|
| 1314 |
+
"trained": false,
|
| 1315 |
+
"sec": 4.5
|
| 1316 |
+
},
|
| 1317 |
+
"boolq": {
|
| 1318 |
+
"n": 300,
|
| 1319 |
+
"type": "noul",
|
| 1320 |
+
"acc": 0.8667,
|
| 1321 |
+
"ece": 0.0319,
|
| 1322 |
+
"brier": 0.1862,
|
| 1323 |
+
"nll": 0.3007,
|
| 1324 |
+
"trained": true,
|
| 1325 |
+
"sec": 6.2
|
| 1326 |
+
},
|
| 1327 |
+
"squad2": {
|
| 1328 |
+
"n": 299,
|
| 1329 |
+
"type": "noul",
|
| 1330 |
+
"acc": 0.7893,
|
| 1331 |
+
"ece": 0.0441,
|
| 1332 |
+
"brier": 0.2955,
|
| 1333 |
+
"nll": 0.452,
|
| 1334 |
+
"trained": false,
|
| 1335 |
+
"sec": 7.2
|
| 1336 |
+
},
|
| 1337 |
+
"paws": {
|
| 1338 |
+
"n": 250,
|
| 1339 |
+
"type": "noul",
|
| 1340 |
+
"acc": 0.764,
|
| 1341 |
+
"ece": 0.1238,
|
| 1342 |
+
"brier": 0.3398,
|
| 1343 |
+
"nll": 0.5538,
|
| 1344 |
+
"trained": true,
|
| 1345 |
+
"sec": 3.0
|
| 1346 |
+
},
|
| 1347 |
+
"multinli": {
|
| 1348 |
+
"n": 299,
|
| 1349 |
+
"type": "choice",
|
| 1350 |
+
"acc": 0.9331,
|
| 1351 |
+
"ece": 0.0103,
|
| 1352 |
+
"brier": 0.1089,
|
| 1353 |
+
"nll": 0.2099,
|
| 1354 |
+
"trained": true,
|
| 1355 |
+
"sec": 3.1
|
| 1356 |
+
},
|
| 1357 |
+
"civil_comments": {
|
| 1358 |
+
"n": 300,
|
| 1359 |
+
"type": "noul",
|
| 1360 |
+
"acc": 0.8733,
|
| 1361 |
+
"ece": 0.0566,
|
| 1362 |
+
"brier": 0.1917,
|
| 1363 |
+
"nll": 0.3199,
|
| 1364 |
+
"trained": true,
|
| 1365 |
+
"sec": 3.2
|
| 1366 |
+
},
|
| 1367 |
+
"aegis2": {
|
| 1368 |
+
"n": 250,
|
| 1369 |
+
"type": "noul",
|
| 1370 |
+
"acc": 0.804,
|
| 1371 |
+
"ece": 0.0467,
|
| 1372 |
+
"brier": 0.2797,
|
| 1373 |
+
"nll": 0.4527,
|
| 1374 |
+
"trained": false,
|
| 1375 |
+
"sec": 6.1
|
| 1376 |
+
},
|
| 1377 |
+
"helpsteer2": {
|
| 1378 |
+
"n": 249,
|
| 1379 |
+
"type": "score",
|
| 1380 |
+
"acc": 0.4498,
|
| 1381 |
+
"ece": 0.142,
|
| 1382 |
+
"brier": 0.6713,
|
| 1383 |
+
"nll": 1.2502,
|
| 1384 |
+
"trained": true,
|
| 1385 |
+
"sec": 25.8,
|
| 1386 |
+
"score_mae": 0.797
|
| 1387 |
+
},
|
| 1388 |
+
"summeval-relevance": {
|
| 1389 |
+
"n": 240,
|
| 1390 |
+
"type": "score",
|
| 1391 |
+
"acc": 0.3958,
|
| 1392 |
+
"ece": 0.0696,
|
| 1393 |
+
"brier": 0.6844,
|
| 1394 |
+
"nll": 1.2541,
|
| 1395 |
+
"trained": false,
|
| 1396 |
+
"sec": 20.1,
|
| 1397 |
+
"score_mae": 0.6172
|
| 1398 |
+
},
|
| 1399 |
+
"summeval-consistency": {
|
| 1400 |
+
"n": 144,
|
| 1401 |
+
"type": "score",
|
| 1402 |
+
"acc": 0.7986,
|
| 1403 |
+
"ece": 0.3095,
|
| 1404 |
+
"brier": 0.4741,
|
| 1405 |
+
"nll": 0.9007,
|
| 1406 |
+
"trained": false,
|
| 1407 |
+
"sec": 12.2,
|
| 1408 |
+
"score_mae": 0.7974
|
| 1409 |
+
},
|
| 1410 |
+
"pubmedqa": {
|
| 1411 |
+
"n": 250,
|
| 1412 |
+
"type": "choice",
|
| 1413 |
+
"acc": 0.708,
|
| 1414 |
+
"ece": 0.0863,
|
| 1415 |
+
"brier": 0.4012,
|
| 1416 |
+
"nll": 0.7205,
|
| 1417 |
+
"trained": true,
|
| 1418 |
+
"sec": 3.8
|
| 1419 |
+
},
|
| 1420 |
+
"_macro": 0.756,
|
| 1421 |
+
"_micro": 0.7652,
|
| 1422 |
+
"_macro_untrained": 0.7284
|
| 1423 |
+
},
|
| 1424 |
+
"jevbench_public": {
|
| 1425 |
+
"easy": {
|
| 1426 |
+
"n": 48,
|
| 1427 |
+
"accuracy": 1.0,
|
| 1428 |
+
"invalid": 0,
|
| 1429 |
+
"ece": 0.0027979166666670663,
|
| 1430 |
+
"by_family": {
|
| 1431 |
+
"extraction": {
|
| 1432 |
+
"n": 12,
|
| 1433 |
+
"accuracy": 1.0
|
| 1434 |
+
},
|
| 1435 |
+
"fact": {
|
| 1436 |
+
"n": 12,
|
| 1437 |
+
"accuracy": 1.0
|
| 1438 |
+
},
|
| 1439 |
+
"intent": {
|
| 1440 |
+
"n": 12,
|
| 1441 |
+
"accuracy": 1.0
|
| 1442 |
+
},
|
| 1443 |
+
"tool_selection": {
|
| 1444 |
+
"n": 12,
|
| 1445 |
+
"accuracy": 1.0
|
| 1446 |
+
}
|
| 1447 |
+
}
|
| 1448 |
+
},
|
| 1449 |
+
"standard": {
|
| 1450 |
+
"n": 72,
|
| 1451 |
+
"accuracy": 0.9861111111111112,
|
| 1452 |
+
"invalid": 0,
|
| 1453 |
+
"ece": 0.0753152777777778,
|
| 1454 |
+
"by_family": {
|
| 1455 |
+
"adequacy": {
|
| 1456 |
+
"n": 12,
|
| 1457 |
+
"accuracy": 0.9166666666666666
|
| 1458 |
+
},
|
| 1459 |
+
"extraction": {
|
| 1460 |
+
"n": 12,
|
| 1461 |
+
"accuracy": 1.0
|
| 1462 |
+
},
|
| 1463 |
+
"intent": {
|
| 1464 |
+
"n": 12,
|
| 1465 |
+
"accuracy": 1.0
|
| 1466 |
+
},
|
| 1467 |
+
"ordinal": {
|
| 1468 |
+
"n": 12,
|
| 1469 |
+
"accuracy": 1.0
|
| 1470 |
+
},
|
| 1471 |
+
"policy": {
|
| 1472 |
+
"n": 12,
|
| 1473 |
+
"accuracy": 1.0
|
| 1474 |
+
},
|
| 1475 |
+
"routing": {
|
| 1476 |
+
"n": 12,
|
| 1477 |
+
"accuracy": 1.0
|
| 1478 |
+
}
|
| 1479 |
+
}
|
| 1480 |
+
},
|
| 1481 |
+
"hard": {
|
| 1482 |
+
"n": 111,
|
| 1483 |
+
"accuracy": 0.6486486486486487,
|
| 1484 |
+
"invalid": 0,
|
| 1485 |
+
"ece": 0.18370720720720732,
|
| 1486 |
+
"by_family": {
|
| 1487 |
+
"adversarial": {
|
| 1488 |
+
"n": 6,
|
| 1489 |
+
"accuracy": 1.0
|
| 1490 |
+
},
|
| 1491 |
+
"ambiguous": {
|
| 1492 |
+
"n": 7,
|
| 1493 |
+
"accuracy": 0.7142857142857143
|
| 1494 |
+
},
|
| 1495 |
+
"judge_hard": {
|
| 1496 |
+
"n": 17,
|
| 1497 |
+
"accuracy": 0.5294117647058824
|
| 1498 |
+
},
|
| 1499 |
+
"long_policy": {
|
| 1500 |
+
"n": 19,
|
| 1501 |
+
"accuracy": 0.47368421052631576
|
| 1502 |
+
},
|
| 1503 |
+
"multi_hop": {
|
| 1504 |
+
"n": 18,
|
| 1505 |
+
"accuracy": 0.8333333333333334
|
| 1506 |
+
},
|
| 1507 |
+
"probability": {
|
| 1508 |
+
"n": 10,
|
| 1509 |
+
"accuracy": 0.8
|
| 1510 |
+
},
|
| 1511 |
+
"routing_hard": {
|
| 1512 |
+
"n": 5,
|
| 1513 |
+
"accuracy": 1.0
|
| 1514 |
+
},
|
| 1515 |
+
"temporal_numeric": {
|
| 1516 |
+
"n": 15,
|
| 1517 |
+
"accuracy": 0.3333333333333333
|
| 1518 |
+
},
|
| 1519 |
+
"tradeoff": {
|
| 1520 |
+
"n": 6,
|
| 1521 |
+
"accuracy": 0.3333333333333333
|
| 1522 |
+
},
|
| 1523 |
+
"trap": {
|
| 1524 |
+
"n": 8,
|
| 1525 |
+
"accuracy": 1.0
|
| 1526 |
+
}
|
| 1527 |
}
|
| 1528 |
}
|
| 1529 |
},
|
| 1530 |
"references": {
|
| 1531 |
+
"v1_4b": {
|
| 1532 |
+
"reg_in": 0.8338,
|
| 1533 |
+
"reg_out": 0.7876,
|
| 1534 |
+
"reg_in_nll": 0.4044,
|
| 1535 |
+
"reg_out_nll": 0.5583,
|
| 1536 |
+
"reg_in_ece": 0.0269,
|
| 1537 |
+
"reg_out_ece": 0.0708
|
| 1538 |
+
},
|
| 1539 |
+
"v2_4b": {
|
| 1540 |
+
"reg_in": 0.8241,
|
| 1541 |
+
"reg_out": 0.7788,
|
| 1542 |
+
"reg_in_nll": 0.4411,
|
| 1543 |
+
"reg_out_nll": 0.5663,
|
| 1544 |
+
"reg_in_ece": 0.0413,
|
| 1545 |
+
"reg_out_ece": 0.0803
|
| 1546 |
+
}
|
| 1547 |
+
},
|
| 1548 |
+
"speed": "not measured again: the architecture and size equal the parent's, and the temperature map does not change the computation"
|
| 1549 |
}
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 8411558400
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ee8ce585b3cedd93206dd149b09b4bdd683174874211f85c77a36090b90c9fdd
|
| 3 |
size 8411558400
|
tokenizer_config.json
CHANGED
|
@@ -10,7 +10,7 @@
|
|
| 10 |
"errors": "replace",
|
| 11 |
"image_token": "<|image_pad|>",
|
| 12 |
"is_local": true,
|
| 13 |
-
"local_files_only":
|
| 14 |
"model_max_length": 262144,
|
| 15 |
"model_specific_special_tokens": {
|
| 16 |
"audio_bos_token": "<|audio_start|>",
|
|
|
|
| 10 |
"errors": "replace",
|
| 11 |
"image_token": "<|image_pad|>",
|
| 12 |
"is_local": true,
|
| 13 |
+
"local_files_only": false,
|
| 14 |
"model_max_length": 262144,
|
| 15 |
"model_specific_special_tokens": {
|
| 16 |
"audio_bos_token": "<|audio_start|>",
|