File size: 9,902 Bytes
03223d7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
---
license: apache-2.0
library_name: openjev
pipeline_tag: zero-shot-classification
base_model: google/gemma-4-E4B-it
language:
  - en
tags:
  - openjev
  - classification
  - zero-shot-classification
  - decision-model
  - listwise
  - peft
  - gemma4
  - research
---

# OpenJEV E4B 1.0

OpenJEV turns Gemma 4 E4B into a direct-decision classifier. A caller supplies state, a question, and a closed set of options; the model returns a probability for each option without generating text. This is the public 1.0 release of Experiment 1.

OpenJEV was trained and evaluated on text. An optional, experimental path (`multimodal=True`) keeps Gemma's vision and audio encoders so a call can also take an image or an audio clip; see [Experimental: image and audio input](#experimental-image-and-audio-input).

## Download and use

The release is hosted at [bambamdevs/openjev-e4b](https://huggingface.co/bambamdevs/openjev-e4b). Install CUDA-enabled PyTorch 2.6 or newer first on Windows ([PyTorch instructions](https://pytorch.org/get-started/locally/)). The release contains the adapter, cumulative backbone delta, decision head, calibration, and loader. The loader downloads the pinned `google/gemma-4-E4B-it` base revision on first use. Hugging Face login is optional for public downloads.

```bash
python -m pip install --upgrade "huggingface_hub<2"
hf auth login
hf download bambamdevs/openjev-e4b --local-dir openjev-e4b
cd openjev-e4b
python -m pip install -r requirements.txt
python -m examples.basic_choice
```

To score your own options:

```python
from openjev import OpenJEV

model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4")
result = model.choice(
    state="Checkout fails after a deployment; customers cannot place orders.",
    instruction="Which team should own this incident?",
    options=["Billing", "Technical outages", "Sales"],
)
print(result["probabilities"], result["selected_option"])
```

The NF4 path keeps frozen base layers quantized and restores delta-touched layers 37–41 in dense FP16. The decision head runs in FP32. `load_mode="bf16"` is available with enough memory; its public benchmarks have not been run.

### Use as a classifier

Put the text to classify in `state`, the question in `instruction`, and your labels in `options`. Any label set of two or more works on each call; no retraining is needed.

```python
r = model.choice(
    state="I was charged twice for my subscription this month.",
    instruction="Which support queue should handle this message?",
    options=["billing", "technical support", "account access", "sales"],
)
r["selected_option"]  # "billing"
r["probabilities"]    # one probability per option, in the same order as options
```

The same pattern covers sentiment, topic, intent, and entailment labels. Probabilities use the release calibration in `model/calibration.json`.

### Experimental: image and audio input

Load with `multimodal=True` to pass `image=` (path or PIL image) or `audio=` (`.wav` path) to `choice()`:

```python
model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4", multimodal=True)
r = model.choice(
    state="",
    audio="call.wav",
    instruction="Which support queue should handle the spoken message?",
    options=["billing", "technical support", "account access", "sales"],
)
```

The adapter, backbone delta, and decision head were trained on text only, so this path is experimental. In a small smoke test on clean synthetic inputs (NF4, RTX 3060), it scored 12/12 on image shape, 12/12 on image colour, 8/8 on printed messages, 16/16 on spoken messages, and 12/12 on spoken reviews. With the media removed, the same questions scored at chance. Image calls took about 5 s and audio calls about 2 s. This is not a benchmark. Details, limits, and the rerun script are in [`docs/MULTIMODAL_EXPERIMENTAL.md`](./docs/MULTIMODAL_EXPERIMENTAL.md).

## Evaluation

The four choice-panel tasks use frozen rows and cyclic option rotations. For each row, probabilities are mapped back to the original option identities and averaged across rotations; accuracy is calculated from that mean. The table also reports mean accuracy for individual presentations, which exposes option-order sensitivity. A separate options-only condition removes the state and question. The five standard public tasks below were evaluated in their native option order. All model evaluations use NF4 on the Windows RTX 3060.

<!-- CHOICE_RESULTS_START -->
### Choice-panel evaluation

The main score selects the largest probability after averaging each row's cyclic option rotations. Wilson 95% intervals use rows as observations. Both systems use the same frozen rows and Windows NF4 evaluation path.

| Task | Rows | Rotations | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] | Difference [95% CI] |
|---|---:|---:|---:|---:|---:|---:|
| GPQA Diamond, shuffled | 198 | 4 | 25.0% | 32.8% [26.7, 39.6] | 32.8% [26.7, 39.6] | +0.00 [-7.07, +7.07] pp |
| Chess legal | 500 | 4 | 25.0% | 43.8% [39.5, 48.2] | 42.6% [38.3, 47.0] | -1.20 [-6.40, +3.80] pp |
| GSM8K, 4 choices | 1,319 | 4 | 25.0% | 43.7% [41.1, 46.4] | 53.8% [51.1, 56.5] | +10.08 [+6.44, +13.72] pp |
| GSM8K, 10 choices | 1,319 | 10 | 10.0% | 20.5% [18.5, 22.8] | 30.6% [28.2, 33.2] | +10.08 [+7.05, +13.04] pp |

Mean accuracy per presentation and share of rows with the same predicted option in every rotation:

| Task | Gemma per order / same answer | OpenJEV 1.0 per order / same answer |
|---|---:|---:|
| GPQA Diamond, shuffled | 31.3% / 10.1% | 33.7% / 23.2% |
| Chess legal | 31.5% / 1.4% | 40.1% / 46.6% |
| GSM8K, 4 choices | 36.7% / 6.9% | 51.5% / 43.3% |
| GSM8K, 10 choices | 16.5% / 0.6% | 28.6% / 14.9% |

### Options-only control

State and question are replaced by neutral text. Four-choice tasks use four rotations; GSM8K-10 uses five evenly spaced rotations.

| Task | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] |
|---|---:|---:|---:|
| GPQA Diamond, shuffled | 25.0% | 25.8% [20.2, 32.3] | 32.3% [26.2, 39.1] |
| Chess legal | 25.0% | 29.0% [25.2, 33.1] | 30.8% [26.9, 35.0] |
| GSM8K, 4 choices | 25.0% | 24.7% [22.5, 27.1] | 24.3% [22.1, 26.7] |
| GSM8K, 10 choices | 10.0% | 10.5% [9.0, 12.3] | 10.8% [9.3, 12.6] |

OpenJEV 1.0 scores 32.3% on GPQA and 30.8% on Chess legal with only the options present, both above the 25% chance rate. Its GPQA full-condition score is 32.8%; this result alone does not establish use of the question. The GSM8K options-only scores are near chance. The row-construction audit passes its chance gate, but the model control reveals additional choice-only signal. Chess legal tests move legality; the GSM8K tasks test closed-choice arithmetic. Gemma uses a native answer-letter readout and OpenJEV uses a learned decision head, so comparisons measure complete systems.

Evidence: [`eval/choice_panel/`](./eval/choice_panel/). The five standard public tasks below use native option order and are reported separately.
<!-- CHOICE_RESULTS_END -->

### Standard public tasks

**Benchmark status: COMPLETE.** All five standard public tasks have both Gemma base NF4 control and OpenJEV 1.0 results.

| Benchmark | Gemma base NF4 | OpenJEV 1.0 | OpenJEV - Gemma |
|---|---:|---:|---:|
| ARC-Challenge | 87.37% | 87.88% | +0.51 pp |
| WinoGrande | 60.14% | 66.46% | +6.32 pp |
| ARC-Easy | 95.58% | 95.62% | +0.04 pp |
| HellaSwag | 76.51% | 89.77% | +13.26 pp |
| MMLU | 66.19% | 64.41% | -1.78 pp |

The untouched base uses native LM-head answer-label scoring; OpenJEV 1.0 uses a learned decision head. Their comparison measures the complete decision systems. ARC, HellaSwag, and MMLU have training-family exposure described in [`docs/BENCHMARK_EXPOSURE.md`](./docs/BENCHMARK_EXPOSURE.md).

## Paper and evidence

[Experiment 1 whitepaper](./docs/OpenJEV_E4B_Experiment_1_Whitepaper.pdf) describes the model, training lineage, evaluation protocol, results, limitations, and reproduction steps. It accompanies this release.

```text
model/                      adapter, head, cumulative backbone delta, calibration
openjev/                    loader and scoring implementation
eval/choice_panel/          frozen rows, audit, scoring code, row-level results
eval/harness/               standard public-task runner
docs/                       whitepaper, source, and result notes
provenance/                 checkpoint metadata
```

OpenJEV 1.0 is a delta release; Google base weights are not republished. `backbone_delta.safetensors` is required to restore the model.

## Training and limitations

The model evolved through readout and adapter training, then an Arena curriculum with exact-oracle examples and replay across 20 decision families. The public checkpoint also updates selected decoder layers and includes 65 cumulative backbone-delta tensors. The development sequence, promotion rule, and release hashes are in the whitepaper and `RELEASE_METADATA.json`.

The experiment has one training seed per stage. Internal populations changed during development. The choice-panel results measure accuracy on their stated tasks; they do not isolate the effect of the head, LoRA, backbone updates, or inference precision. GPQA and GSM8K are multiple-choice tasks, and Chess legal tests move legality rather than move quality.

## Reproduce and license

See [`REPRODUCIBILITY.md`](./REPRODUCIBILITY.md), [`RELEASE_METADATA.json`](./RELEASE_METADATA.json), and [`MANIFEST.sha256`](./MANIFEST.sha256). OpenJEV-authored material is Apache-2.0. The Gemma base remains under Google's terms. Training code is not part of this release.

OpenJEV is independent research. It is not a TypeSafe AI product, is not affiliated with or endorsed by TypeSafe AI, and is unrelated to `AlexWortega/openjev`. Its design was inspired by System 1-style decision making: a fast, direct choice over a closed set of options, without step-by-step text generation.

Generated: 2026-09-27T17:07:53.606842+00:00