OpenJEV E4B 1.0

OpenJEV turns Gemma 4 E4B into a direct-decision classifier. A caller supplies state, a question, and a closed set of options; the model returns a probability for each option without generating text. This is the public 1.0 release of Experiment 1.

OpenJEV was trained and evaluated on text. An optional, experimental path (multimodal=True) keeps Gemma's vision and audio encoders so a call can also take an image or an audio clip; see Experimental: image and audio input.

Download and use

The release is hosted at bambamdevs/openjev-e4b. Install CUDA-enabled PyTorch 2.6 or newer first on Windows (PyTorch instructions). The release contains the adapter, cumulative backbone delta, decision head, calibration, and loader. The loader downloads the pinned google/gemma-4-E4B-it base revision on first use. Hugging Face login is optional for public downloads.

python -m pip install --upgrade "huggingface_hub<2"
hf auth login
hf download bambamdevs/openjev-e4b --local-dir openjev-e4b
cd openjev-e4b
python -m pip install -r requirements.txt
python -m examples.basic_choice

To score your own options:

from openjev import OpenJEV

model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4")
result = model.choice(
    state="Checkout fails after a deployment; customers cannot place orders.",
    instruction="Which team should own this incident?",
    options=["Billing", "Technical outages", "Sales"],
)
print(result["probabilities"], result["selected_option"])

The NF4 path keeps frozen base layers quantized and restores delta-touched layers 37–41 in dense FP16. The decision head runs in FP32. load_mode="bf16" is available with enough memory; its public benchmarks have not been run.

Use as a classifier

Put the text to classify in state, the question in instruction, and your labels in options. Any label set of two or more works on each call; no retraining is needed.

r = model.choice(
    state="I was charged twice for my subscription this month.",
    instruction="Which support queue should handle this message?",
    options=["billing", "technical support", "account access", "sales"],
)
r["selected_option"]  # "billing"
r["probabilities"]    # one probability per option, in the same order as options

The same pattern covers sentiment, topic, intent, and entailment labels. Probabilities use the release calibration in model/calibration.json.

Experimental: image and audio input

Load with multimodal=True to pass image= (path or PIL image) or audio= (.wav path) to choice():

model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4", multimodal=True)
r = model.choice(
    state="",
    audio="call.wav",
    instruction="Which support queue should handle the spoken message?",
    options=["billing", "technical support", "account access", "sales"],
)

The adapter, backbone delta, and decision head were trained on text only, so this path is experimental. In a small smoke test on clean synthetic inputs (NF4, RTX 3060), it scored 12/12 on image shape, 12/12 on image colour, 8/8 on printed messages, 16/16 on spoken messages, and 12/12 on spoken reviews. With the media removed, the same questions scored at chance. Image calls took about 5 s and audio calls about 2 s. This is not a benchmark. Details, limits, and the rerun script are in docs/MULTIMODAL_EXPERIMENTAL.md.

Evaluation

The four choice-panel tasks use frozen rows and cyclic option rotations. For each row, probabilities are mapped back to the original option identities and averaged across rotations; accuracy is calculated from that mean. The table also reports mean accuracy for individual presentations, which exposes option-order sensitivity. A separate options-only condition removes the state and question. The five standard public tasks below were evaluated in their native option order. All model evaluations use NF4 on the Windows RTX 3060.

Choice-panel evaluation

The main score selects the largest probability after averaging each row's cyclic option rotations. Wilson 95% intervals use rows as observations. Both systems use the same frozen rows and Windows NF4 evaluation path.

Task Rows Rotations Chance Gemma base [95% CI] OpenJEV 1.0 [95% CI] Difference [95% CI]
GPQA Diamond, shuffled 198 4 25.0% 32.8% [26.7, 39.6] 32.8% [26.7, 39.6] +0.00 [-7.07, +7.07] pp
Chess legal 500 4 25.0% 43.8% [39.5, 48.2] 42.6% [38.3, 47.0] -1.20 [-6.40, +3.80] pp
GSM8K, 4 choices 1,319 4 25.0% 43.7% [41.1, 46.4] 53.8% [51.1, 56.5] +10.08 [+6.44, +13.72] pp
GSM8K, 10 choices 1,319 10 10.0% 20.5% [18.5, 22.8] 30.6% [28.2, 33.2] +10.08 [+7.05, +13.04] pp

Mean accuracy per presentation and share of rows with the same predicted option in every rotation:

Task Gemma per order / same answer OpenJEV 1.0 per order / same answer
GPQA Diamond, shuffled 31.3% / 10.1% 33.7% / 23.2%
Chess legal 31.5% / 1.4% 40.1% / 46.6%
GSM8K, 4 choices 36.7% / 6.9% 51.5% / 43.3%
GSM8K, 10 choices 16.5% / 0.6% 28.6% / 14.9%

Options-only control

State and question are replaced by neutral text. Four-choice tasks use four rotations; GSM8K-10 uses five evenly spaced rotations.

Task Chance Gemma base [95% CI] OpenJEV 1.0 [95% CI]
GPQA Diamond, shuffled 25.0% 25.8% [20.2, 32.3] 32.3% [26.2, 39.1]
Chess legal 25.0% 29.0% [25.2, 33.1] 30.8% [26.9, 35.0]
GSM8K, 4 choices 25.0% 24.7% [22.5, 27.1] 24.3% [22.1, 26.7]
GSM8K, 10 choices 10.0% 10.5% [9.0, 12.3] 10.8% [9.3, 12.6]

OpenJEV 1.0 scores 32.3% on GPQA and 30.8% on Chess legal with only the options present, both above the 25% chance rate. Its GPQA full-condition score is 32.8%; this result alone does not establish use of the question. The GSM8K options-only scores are near chance. The row-construction audit passes its chance gate, but the model control reveals additional choice-only signal. Chess legal tests move legality; the GSM8K tasks test closed-choice arithmetic. Gemma uses a native answer-letter readout and OpenJEV uses a learned decision head, so comparisons measure complete systems.

Evidence: eval/choice_panel/. The five standard public tasks below use native option order and are reported separately.

Standard public tasks

Benchmark status: COMPLETE. All five standard public tasks have both Gemma base NF4 control and OpenJEV 1.0 results.

Benchmark Gemma base NF4 OpenJEV 1.0 OpenJEV - Gemma
ARC-Challenge 87.37% 87.88% +0.51 pp
WinoGrande 60.14% 66.46% +6.32 pp
ARC-Easy 95.58% 95.62% +0.04 pp
HellaSwag 76.51% 89.77% +13.26 pp
MMLU 66.19% 64.41% -1.78 pp

The untouched base uses native LM-head answer-label scoring; OpenJEV 1.0 uses a learned decision head. Their comparison measures the complete decision systems. ARC, HellaSwag, and MMLU have training-family exposure described in docs/BENCHMARK_EXPOSURE.md.

Paper and evidence

Experiment 1 whitepaper describes the model, training lineage, evaluation protocol, results, limitations, and reproduction steps. It accompanies this release.

model/                      adapter, head, cumulative backbone delta, calibration
openjev/                    loader and scoring implementation
eval/choice_panel/          frozen rows, audit, scoring code, row-level results
eval/harness/               standard public-task runner
docs/                       whitepaper, source, and result notes
provenance/                 checkpoint metadata

OpenJEV 1.0 is a delta release; Google base weights are not republished. backbone_delta.safetensors is required to restore the model.

Training and limitations

The model evolved through readout and adapter training, then an Arena curriculum with exact-oracle examples and replay across 20 decision families. The public checkpoint also updates selected decoder layers and includes 65 cumulative backbone-delta tensors. The development sequence, promotion rule, and release hashes are in the whitepaper and RELEASE_METADATA.json.

The experiment has one training seed per stage. Internal populations changed during development. The choice-panel results measure accuracy on their stated tasks; they do not isolate the effect of the head, LoRA, backbone updates, or inference precision. GPQA and GSM8K are multiple-choice tasks, and Chess legal tests move legality rather than move quality.

Reproduce and license

See REPRODUCIBILITY.md, RELEASE_METADATA.json, and MANIFEST.sha256. OpenJEV-authored material is Apache-2.0. The Gemma base remains under Google's terms. Training code is not part of this release.

OpenJEV is independent research. It is not a TypeSafe AI product, is not affiliated with or endorsed by TypeSafe AI, and is unrelated to AlexWortega/openjev. Its design was inspired by System 1-style decision making: a fast, direct choice over a closed set of options, without step-by-step text generation.

Generated: 2026-09-27T17:07:53.606842+00:00

Downloads last month
45
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bambamdevs/openjev-e4b

Finetuned
(400)
this model