Instructions to use bambamdevs/openjev-e4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use bambamdevs/openjev-e4b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Download docs/EXPERIMENT_1_RESULTS.md from bambamdevs/openjev-e4b: direct link, hf CLI and curl.
- Browser
- Download file 3.87 kB
-
https://huggingface.co/bambamdevs/openjev-e4b/resolve/main/docs/EXPERIMENT_1_RESULTS.md
- Command line
-
hf download hf://bambamdevs/openjev-e4b/docs/EXPERIMENT_1_RESULTS.md
-
curl -L -o EXPERIMENT_1_RESULTS.md https://huggingface.co/bambamdevs/openjev-e4b/resolve/main/docs/EXPERIMENT_1_RESULTS.md
Experiment 1 results — OpenJEV E4B 1.0
The public evaluation has two protocols. GPQA Diamond, Chess legal, GSM8K-4, and GSM8K-10 use frozen choice-panel rows and cyclic option rotations. Five standard public tasks retain their native option order. Gemma and OpenJEV 1.0 completed the choice-panel evaluation on the Windows RTX 3060. The completed untouched-Gemma results below come from ../eval/choice_panel/summary.json.
Choice-panel evaluation
Probabilities from each cyclic rotation are mapped to original option identities and averaged per row. The main accuracy chooses the largest mean. Wilson 95% intervals treat rows as observations. Mean per-order accuracy averages individual presentations; same-answer rate is the share of rows whose chosen option never changes across rotations.
| Task | Rows | Rotations | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] | Difference [95% CI] |
|---|---|---|---|---|---|---|
| GPQA Diamond, shuffled | 198 | 4 | 25.0% | 32.8% [26.7, 39.6] | 32.8% [26.7, 39.6] | +0.00 [-7.07, +7.07] pp |
| Chess legal | 500 | 4 | 25.0% | 43.8% [39.5, 48.2] | 42.6% [38.3, 47.0] | -1.20 [-6.40, +3.80] pp |
| GSM8K, 4 choices | 1,319 | 4 | 25.0% | 43.7% [41.1, 46.4] | 53.8% [51.1, 56.5] | +10.08 [+6.44, +13.72] pp |
| GSM8K, 10 choices | 1,319 | 10 | 10.0% | 20.5% [18.5, 22.8] | 30.6% [28.2, 33.2] | +10.08 [+7.05, +13.04] pp |
| Task | Gemma per order / same answer | OpenJEV 1.0 per order / same answer |
|---|---|---|
| GPQA Diamond, shuffled | 31.3% / 10.1% | 33.7% / 23.2% |
| Chess legal | 31.5% / 1.4% | 40.1% / 46.6% |
| GSM8K, 4 choices | 36.7% / 6.9% | 51.5% / 43.3% |
| GSM8K, 10 choices | 16.5% / 0.6% | 28.6% / 14.9% |
Paired intervals use 2,000 row bootstrap resamples. Exact McNemar p-values and flip counts are in ../eval/choice_panel/summary.json.
Options-only control
State and question are replaced by neutral text. Ten-choice tasks use five evenly spaced rotations in this condition.
| Task | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] |
|---|---|---|---|
| GPQA Diamond, shuffled | 25.0% | 25.8% [20.2, 32.3] | 32.3% [26.2, 39.1] |
| Chess legal | 25.0% | 29.0% [25.2, 33.1] | 30.8% [26.9, 35.0] |
| GSM8K, 4 choices | 25.0% | 24.7% [22.5, 27.1] | 24.3% [22.1, 26.7] |
| GSM8K, 10 choices | 10.0% | 10.5% [9.0, 12.3] | 10.8% [9.3, 12.6] |
The frozen-row construction audit passes its chance gate on all four tasks; see ../eval/choice_panel/leak_audit.json. OpenJEV 1.0 nevertheless scores above chance with options only on GPQA (32.3%) and Chess legal (30.8%). Its GPQA full-condition score is 32.8%, so that result alone does not establish use of the question. The GSM8K options-only scores are near chance. Chess legal measures legality rather than move quality; the GSM8K tasks do not measure open-ended solution generation.
Standard public tasks
The following are single-order results on identical rows. They compare the complete native-readout Gemma system with OpenJEV 1.0. They do not isolate the training effect of the adapter or backbone delta. ARC, HellaSwag, and MMLU have training-family exposure; see BENCHMARK_EXPOSURE.md.
| Task | Gemma base NF4 | OpenJEV 1.0 | Difference [95% CI] |
|---|---|---|---|
| ARC-Easy | 95.58% | 95.62% | +0.04 [-0.80, +0.84] pp |
| ARC-Challenge | 87.37% | 87.88% | +0.51 [-1.37, +2.39] pp |
| WinoGrande | 60.14% | 66.46% | +6.31 [+3.79, +9.00] pp |
| HellaSwag | 76.51% | 89.77% | +13.26 [+12.40, +14.06] pp |
| MMLU | 66.19% | 64.41% | -1.77 [-2.58, -1.01] pp |
The standard-task and choice-panel scores use different protocols. Cross-task averages require explicit weights and a statement of the protocol difference.