Independent run: our MNLI 0.7417 against your published 0.856 - which is right?
Independent run of decider-2b, and a gap on MNLI we would like to resolve.
Setup. Revision 533964dae8be, one sealed manifest, seed 42, 1,190 cases / 1,240 scored evaluation decisions across 9 typed-decision suites, bf16 on one T4.
What we measured (overall 0.7895):
| suite | ours |
|---|---|
| MNLI | 0.7417 |
| our moderation suite | 0.8819 |
| our guardrails suite | 0.6354 |
| emotion | 0.6719 |
The gap. Your card reports multinli (choice; trained) at 0.856 / 0.853. Ours is 0.7417 on MNLI - about 0.11 lower. Since both are choice questions over MNLI, we could not explain it from task shape alone.
Three candidates on our side, in the order we suspect them:
- Prompt shape. We present the full label set as declared options and score with
per_option_conditional_logprob. A card that reports a trained/choice row may use a different verbalisation. - Label set. We use the 3-way MNLI test split. A different label grouping changes the ceiling.
- Split membership. Our manifest pins case IDs and order indices, so we can diff row-by-row rather than argue in aggregate.
What would help. The exact MNLI prompt and label set behind the 0.856 / 0.853 row. We will re-run and report whichever number is correct, including if yours holds and ours is the bug.
Note our guardrails (0.6354) and moderation (0.8819) spread far more than your aegis2 (noul) 0.728 / 0.716 would suggest, but our two suites are different datasets from aegis2, so we are not treating that as a comparison.
Full data, per-suite matrices, per-row predictions: https://sysone.sdad.pro/