openjev-e4b / docs /EXPERIMENT_1_RESULTS.md
bambamdevs's picture
Publish OpenJEV E4B 1.0
03223d7
|
Raw History Blame Contribute Delete
3.87 kB

Experiment 1 results — OpenJEV E4B 1.0

The public evaluation has two protocols. GPQA Diamond, Chess legal, GSM8K-4, and GSM8K-10 use frozen choice-panel rows and cyclic option rotations. Five standard public tasks retain their native option order. Gemma and OpenJEV 1.0 completed the choice-panel evaluation on the Windows RTX 3060. The completed untouched-Gemma results below come from ../eval/choice_panel/summary.json.

Choice-panel evaluation

Probabilities from each cyclic rotation are mapped to original option identities and averaged per row. The main accuracy chooses the largest mean. Wilson 95% intervals treat rows as observations. Mean per-order accuracy averages individual presentations; same-answer rate is the share of rows whose chosen option never changes across rotations.

Task Rows Rotations Chance Gemma base [95% CI] OpenJEV 1.0 [95% CI] Difference [95% CI]
GPQA Diamond, shuffled 198 4 25.0% 32.8% [26.7, 39.6] 32.8% [26.7, 39.6] +0.00 [-7.07, +7.07] pp
Chess legal 500 4 25.0% 43.8% [39.5, 48.2] 42.6% [38.3, 47.0] -1.20 [-6.40, +3.80] pp
GSM8K, 4 choices 1,319 4 25.0% 43.7% [41.1, 46.4] 53.8% [51.1, 56.5] +10.08 [+6.44, +13.72] pp
GSM8K, 10 choices 1,319 10 10.0% 20.5% [18.5, 22.8] 30.6% [28.2, 33.2] +10.08 [+7.05, +13.04] pp
Task Gemma per order / same answer OpenJEV 1.0 per order / same answer
GPQA Diamond, shuffled 31.3% / 10.1% 33.7% / 23.2%
Chess legal 31.5% / 1.4% 40.1% / 46.6%
GSM8K, 4 choices 36.7% / 6.9% 51.5% / 43.3%
GSM8K, 10 choices 16.5% / 0.6% 28.6% / 14.9%

Paired intervals use 2,000 row bootstrap resamples. Exact McNemar p-values and flip counts are in ../eval/choice_panel/summary.json.

Options-only control

State and question are replaced by neutral text. Ten-choice tasks use five evenly spaced rotations in this condition.

Task Chance Gemma base [95% CI] OpenJEV 1.0 [95% CI]
GPQA Diamond, shuffled 25.0% 25.8% [20.2, 32.3] 32.3% [26.2, 39.1]
Chess legal 25.0% 29.0% [25.2, 33.1] 30.8% [26.9, 35.0]
GSM8K, 4 choices 25.0% 24.7% [22.5, 27.1] 24.3% [22.1, 26.7]
GSM8K, 10 choices 10.0% 10.5% [9.0, 12.3] 10.8% [9.3, 12.6]

The frozen-row construction audit passes its chance gate on all four tasks; see ../eval/choice_panel/leak_audit.json. OpenJEV 1.0 nevertheless scores above chance with options only on GPQA (32.3%) and Chess legal (30.8%). Its GPQA full-condition score is 32.8%, so that result alone does not establish use of the question. The GSM8K options-only scores are near chance. Chess legal measures legality rather than move quality; the GSM8K tasks do not measure open-ended solution generation.

Standard public tasks

The following are single-order results on identical rows. They compare the complete native-readout Gemma system with OpenJEV 1.0. They do not isolate the training effect of the adapter or backbone delta. ARC, HellaSwag, and MMLU have training-family exposure; see BENCHMARK_EXPOSURE.md.

Task Gemma base NF4 OpenJEV 1.0 Difference [95% CI]
ARC-Easy 95.58% 95.62% +0.04 [-0.80, +0.84] pp
ARC-Challenge 87.37% 87.88% +0.51 [-1.37, +2.39] pp
WinoGrande 60.14% 66.46% +6.31 [+3.79, +9.00] pp
HellaSwag 76.51% 89.77% +13.26 [+12.40, +14.06] pp
MMLU 66.19% 64.41% -1.77 [-2.58, -1.01] pp

The standard-task and choice-panel scores use different protocols. Cross-task averages require explicit weights and a statement of the protocol difference.