# Experiment 1 results — OpenJEV E4B 1.0 The public evaluation has two protocols. GPQA Diamond, Chess legal, GSM8K-4, and GSM8K-10 use frozen choice-panel rows and cyclic option rotations. Five standard public tasks retain their native option order. Gemma and OpenJEV 1.0 completed the choice-panel evaluation on the Windows RTX 3060. The completed untouched-Gemma results below come from [`../eval/choice_panel/summary.json`](../eval/choice_panel/summary.json). ## Choice-panel evaluation Probabilities from each cyclic rotation are mapped to original option identities and averaged per row. The main accuracy chooses the largest mean. Wilson 95% intervals treat rows as observations. Mean per-order accuracy averages individual presentations; same-answer rate is the share of rows whose chosen option never changes across rotations. | Task | Rows | Rotations | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] | Difference [95% CI] | |---|---:|---:|---:|---:|---:|---:| | GPQA Diamond, shuffled | 198 | 4 | 25.0% | 32.8% [26.7, 39.6] | 32.8% [26.7, 39.6] | +0.00 [-7.07, +7.07] pp | | Chess legal | 500 | 4 | 25.0% | 43.8% [39.5, 48.2] | 42.6% [38.3, 47.0] | -1.20 [-6.40, +3.80] pp | | GSM8K, 4 choices | 1,319 | 4 | 25.0% | 43.7% [41.1, 46.4] | 53.8% [51.1, 56.5] | +10.08 [+6.44, +13.72] pp | | GSM8K, 10 choices | 1,319 | 10 | 10.0% | 20.5% [18.5, 22.8] | 30.6% [28.2, 33.2] | +10.08 [+7.05, +13.04] pp | | Task | Gemma per order / same answer | OpenJEV 1.0 per order / same answer | |---|---:|---:| | GPQA Diamond, shuffled | 31.3% / 10.1% | 33.7% / 23.2% | | Chess legal | 31.5% / 1.4% | 40.1% / 46.6% | | GSM8K, 4 choices | 36.7% / 6.9% | 51.5% / 43.3% | | GSM8K, 10 choices | 16.5% / 0.6% | 28.6% / 14.9% | Paired intervals use 2,000 row bootstrap resamples. Exact McNemar p-values and flip counts are in [`../eval/choice_panel/summary.json`](../eval/choice_panel/summary.json). ## Options-only control State and question are replaced by neutral text. Ten-choice tasks use five evenly spaced rotations in this condition. | Task | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] | |---|---:|---:|---:| | GPQA Diamond, shuffled | 25.0% | 25.8% [20.2, 32.3] | 32.3% [26.2, 39.1] | | Chess legal | 25.0% | 29.0% [25.2, 33.1] | 30.8% [26.9, 35.0] | | GSM8K, 4 choices | 25.0% | 24.7% [22.5, 27.1] | 24.3% [22.1, 26.7] | | GSM8K, 10 choices | 10.0% | 10.5% [9.0, 12.3] | 10.8% [9.3, 12.6] | The frozen-row construction audit passes its chance gate on all four tasks; see [`../eval/choice_panel/leak_audit.json`](../eval/choice_panel/leak_audit.json). OpenJEV 1.0 nevertheless scores above chance with options only on GPQA (32.3%) and Chess legal (30.8%). Its GPQA full-condition score is 32.8%, so that result alone does not establish use of the question. The GSM8K options-only scores are near chance. Chess legal measures legality rather than move quality; the GSM8K tasks do not measure open-ended solution generation. ## Standard public tasks The following are single-order results on identical rows. They compare the complete native-readout Gemma system with OpenJEV 1.0. They do not isolate the training effect of the adapter or backbone delta. ARC, HellaSwag, and MMLU have training-family exposure; see [`BENCHMARK_EXPOSURE.md`](./BENCHMARK_EXPOSURE.md). | Task | Gemma base NF4 | OpenJEV 1.0 | Difference [95% CI] | |---|---:|---:|---:| | ARC-Easy | 95.58% | 95.62% | +0.04 [-0.80, +0.84] pp | | ARC-Challenge | 87.37% | 87.88% | +0.51 [-1.37, +2.39] pp | | WinoGrande | 60.14% | 66.46% | +6.31 [+3.79, +9.00] pp | | HellaSwag | 76.51% | 89.77% | +13.26 [+12.40, +14.06] pp | | MMLU | 66.19% | 64.41% | -1.77 [-2.58, -1.01] pp | The standard-task and choice-panel scores use different protocols. Cross-task averages require explicit weights and a statement of the protocol difference.