--- license: apache-2.0 library_name: openjev pipeline_tag: zero-shot-classification base_model: google/gemma-4-E4B-it language: - en tags: - openjev - classification - zero-shot-classification - decision-model - listwise - peft - gemma4 - research --- # OpenJEV E4B 1.0 OpenJEV turns Gemma 4 E4B into a direct-decision classifier. A caller supplies state, a question, and a closed set of options; the model returns a probability for each option without generating text. This is the public 1.0 release of Experiment 1. OpenJEV was trained and evaluated on text. An optional, experimental path (`multimodal=True`) keeps Gemma's vision and audio encoders so a call can also take an image or an audio clip; see [Experimental: image and audio input](#experimental-image-and-audio-input). ## Download and use The release is hosted at [bambamdevs/openjev-e4b](https://huggingface.co/bambamdevs/openjev-e4b). Install CUDA-enabled PyTorch 2.6 or newer first on Windows ([PyTorch instructions](https://pytorch.org/get-started/locally/)). The release contains the adapter, cumulative backbone delta, decision head, calibration, and loader. The loader downloads the pinned `google/gemma-4-E4B-it` base revision on first use. Hugging Face login is optional for public downloads. ```bash python -m pip install --upgrade "huggingface_hub<2" hf auth login hf download bambamdevs/openjev-e4b --local-dir openjev-e4b cd openjev-e4b python -m pip install -r requirements.txt python -m examples.basic_choice ``` To score your own options: ```python from openjev import OpenJEV model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4") result = model.choice( state="Checkout fails after a deployment; customers cannot place orders.", instruction="Which team should own this incident?", options=["Billing", "Technical outages", "Sales"], ) print(result["probabilities"], result["selected_option"]) ``` The NF4 path keeps frozen base layers quantized and restores delta-touched layers 37–41 in dense FP16. The decision head runs in FP32. `load_mode="bf16"` is available with enough memory; its public benchmarks have not been run. ### Use as a classifier Put the text to classify in `state`, the question in `instruction`, and your labels in `options`. Any label set of two or more works on each call; no retraining is needed. ```python r = model.choice( state="I was charged twice for my subscription this month.", instruction="Which support queue should handle this message?", options=["billing", "technical support", "account access", "sales"], ) r["selected_option"] # "billing" r["probabilities"] # one probability per option, in the same order as options ``` The same pattern covers sentiment, topic, intent, and entailment labels. Probabilities use the release calibration in `model/calibration.json`. ### Experimental: image and audio input Load with `multimodal=True` to pass `image=` (path or PIL image) or `audio=` (`.wav` path) to `choice()`: ```python model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4", multimodal=True) r = model.choice( state="", audio="call.wav", instruction="Which support queue should handle the spoken message?", options=["billing", "technical support", "account access", "sales"], ) ``` The adapter, backbone delta, and decision head were trained on text only, so this path is experimental. In a small smoke test on clean synthetic inputs (NF4, RTX 3060), it scored 12/12 on image shape, 12/12 on image colour, 8/8 on printed messages, 16/16 on spoken messages, and 12/12 on spoken reviews. With the media removed, the same questions scored at chance. Image calls took about 5 s and audio calls about 2 s. This is not a benchmark. Details, limits, and the rerun script are in [`docs/MULTIMODAL_EXPERIMENTAL.md`](./docs/MULTIMODAL_EXPERIMENTAL.md). ## Evaluation The four choice-panel tasks use frozen rows and cyclic option rotations. For each row, probabilities are mapped back to the original option identities and averaged across rotations; accuracy is calculated from that mean. The table also reports mean accuracy for individual presentations, which exposes option-order sensitivity. A separate options-only condition removes the state and question. The five standard public tasks below were evaluated in their native option order. All model evaluations use NF4 on the Windows RTX 3060. ### Choice-panel evaluation The main score selects the largest probability after averaging each row's cyclic option rotations. Wilson 95% intervals use rows as observations. Both systems use the same frozen rows and Windows NF4 evaluation path. | Task | Rows | Rotations | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] | Difference [95% CI] | |---|---:|---:|---:|---:|---:|---:| | GPQA Diamond, shuffled | 198 | 4 | 25.0% | 32.8% [26.7, 39.6] | 32.8% [26.7, 39.6] | +0.00 [-7.07, +7.07] pp | | Chess legal | 500 | 4 | 25.0% | 43.8% [39.5, 48.2] | 42.6% [38.3, 47.0] | -1.20 [-6.40, +3.80] pp | | GSM8K, 4 choices | 1,319 | 4 | 25.0% | 43.7% [41.1, 46.4] | 53.8% [51.1, 56.5] | +10.08 [+6.44, +13.72] pp | | GSM8K, 10 choices | 1,319 | 10 | 10.0% | 20.5% [18.5, 22.8] | 30.6% [28.2, 33.2] | +10.08 [+7.05, +13.04] pp | Mean accuracy per presentation and share of rows with the same predicted option in every rotation: | Task | Gemma per order / same answer | OpenJEV 1.0 per order / same answer | |---|---:|---:| | GPQA Diamond, shuffled | 31.3% / 10.1% | 33.7% / 23.2% | | Chess legal | 31.5% / 1.4% | 40.1% / 46.6% | | GSM8K, 4 choices | 36.7% / 6.9% | 51.5% / 43.3% | | GSM8K, 10 choices | 16.5% / 0.6% | 28.6% / 14.9% | ### Options-only control State and question are replaced by neutral text. Four-choice tasks use four rotations; GSM8K-10 uses five evenly spaced rotations. | Task | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] | |---|---:|---:|---:| | GPQA Diamond, shuffled | 25.0% | 25.8% [20.2, 32.3] | 32.3% [26.2, 39.1] | | Chess legal | 25.0% | 29.0% [25.2, 33.1] | 30.8% [26.9, 35.0] | | GSM8K, 4 choices | 25.0% | 24.7% [22.5, 27.1] | 24.3% [22.1, 26.7] | | GSM8K, 10 choices | 10.0% | 10.5% [9.0, 12.3] | 10.8% [9.3, 12.6] | OpenJEV 1.0 scores 32.3% on GPQA and 30.8% on Chess legal with only the options present, both above the 25% chance rate. Its GPQA full-condition score is 32.8%; this result alone does not establish use of the question. The GSM8K options-only scores are near chance. The row-construction audit passes its chance gate, but the model control reveals additional choice-only signal. Chess legal tests move legality; the GSM8K tasks test closed-choice arithmetic. Gemma uses a native answer-letter readout and OpenJEV uses a learned decision head, so comparisons measure complete systems. Evidence: [`eval/choice_panel/`](./eval/choice_panel/). The five standard public tasks below use native option order and are reported separately. ### Standard public tasks **Benchmark status: COMPLETE.** All five standard public tasks have both Gemma base NF4 control and OpenJEV 1.0 results. | Benchmark | Gemma base NF4 | OpenJEV 1.0 | OpenJEV - Gemma | |---|---:|---:|---:| | ARC-Challenge | 87.37% | 87.88% | +0.51 pp | | WinoGrande | 60.14% | 66.46% | +6.32 pp | | ARC-Easy | 95.58% | 95.62% | +0.04 pp | | HellaSwag | 76.51% | 89.77% | +13.26 pp | | MMLU | 66.19% | 64.41% | -1.78 pp | The untouched base uses native LM-head answer-label scoring; OpenJEV 1.0 uses a learned decision head. Their comparison measures the complete decision systems. ARC, HellaSwag, and MMLU have training-family exposure described in [`docs/BENCHMARK_EXPOSURE.md`](./docs/BENCHMARK_EXPOSURE.md). ## Paper and evidence [Experiment 1 whitepaper](./docs/OpenJEV_E4B_Experiment_1_Whitepaper.pdf) describes the model, training lineage, evaluation protocol, results, limitations, and reproduction steps. It accompanies this release. ```text model/ adapter, head, cumulative backbone delta, calibration openjev/ loader and scoring implementation eval/choice_panel/ frozen rows, audit, scoring code, row-level results eval/harness/ standard public-task runner docs/ whitepaper, source, and result notes provenance/ checkpoint metadata ``` OpenJEV 1.0 is a delta release; Google base weights are not republished. `backbone_delta.safetensors` is required to restore the model. ## Training and limitations The model evolved through readout and adapter training, then an Arena curriculum with exact-oracle examples and replay across 20 decision families. The public checkpoint also updates selected decoder layers and includes 65 cumulative backbone-delta tensors. The development sequence, promotion rule, and release hashes are in the whitepaper and `RELEASE_METADATA.json`. The experiment has one training seed per stage. Internal populations changed during development. The choice-panel results measure accuracy on their stated tasks; they do not isolate the effect of the head, LoRA, backbone updates, or inference precision. GPQA and GSM8K are multiple-choice tasks, and Chess legal tests move legality rather than move quality. ## Reproduce and license See [`REPRODUCIBILITY.md`](./REPRODUCIBILITY.md), [`RELEASE_METADATA.json`](./RELEASE_METADATA.json), and [`MANIFEST.sha256`](./MANIFEST.sha256). OpenJEV-authored material is Apache-2.0. The Gemma base remains under Google's terms. Training code is not part of this release. OpenJEV is independent research. It is not a TypeSafe AI product, is not affiliated with or endorsed by TypeSafe AI, and is unrelated to `AlexWortega/openjev`. Its design was inspired by System 1-style decision making: a fast, direct choice over a closed set of options, without step-by-step text generation. Generated: 2026-09-27T17:07:53.606842+00:00