Independent run: our MNLI 0.7667 against your published 0.926
Following up on our independent run of decider-4b - a large gap on MNLI we cannot yet explain.
Setup. Revision 533964dae8be, one sealed manifest, seed 42, 1,190 cases / 1,240 scored evaluation decisions across 9 typed-decision suites, bf16 on one T4. Same harness and same pinned revision as our decider-2b post.
What we measured (overall 0.8411):
| suite | ours | your card |
|---|---|---|
| MNLI | 0.7667 | 0.926 |
| our moderation suite | 0.8790 | - |
| our guardrails suite | 0.8333 | - |
The gap. Both are choice questions over MNLI. Ours is 0.1593 below yours. That is far outside anything our label grouping or option-count could explain, so we do not think we can attribute it to task shape.
What we suspect on our side, most likely first.
- Prompt and option verbalisation. We declare the label set per case and score with
per_option_conditional_logprob. If your harness instead renders options as an enumerated list and reads a single next-token distribution, the two are not measuring the same function, and MNLI is where that diverges most for us. - Scoring path. Our
choicescoring for decider-4b went through the vendor-default readout with no calibration applied. - Split. Our manifest pins case IDs and order indices, so we can diff row-by-row instead of arguing in aggregate.
What would help. The exact MNLI prompt, label set and answer-extraction rule behind your 0.926 figure. We will re-run and publish whichever number holds, including if yours is right and our harness is the bug. We would rather correct our record than leave a 0.17 gap unexplained.
Full data, per-suite matrices and per-row predictions: https://sysone.sdad.pro/
Correction to the setup line above. Two fields in my post were wrong, both mine, neither affecting the MNLI comparison:
| field | stated above | correct |
|---|---|---|
| revision | 533964dae8be |
eb5fbdfc9448 |
| scoring path | no calibration applied |
vendor-default readout, dtype bf16 |
| our moderation suite | 0.8790 |
0.9444 |
| our guardrails suite | 0.8333 |
0.8333 |
The revision is the one to care about: it was transcribed from the wrong row of our results, so the
checkpoint I named was not the checkpoint I ran. eb5fbdfc9448 is the pinned revision fromdecider-4b-t4-d0. MNLI remains 0.7667 overall 0.8411, and that is unchanged.
Apologies for the noise. Source of truth for every field is the run directory named above.