Instructions to use bambamdevs/openjev-e4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use bambamdevs/openjev-e4b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
OpenJEV E4B 1.0
OpenJEV turns Gemma 4 E4B into a direct-decision classifier. A caller supplies state, a question, and a closed set of options; the model returns a probability for each option without generating text. This is the public 1.0 release of Experiment 1.
OpenJEV was trained and evaluated on text. An optional, experimental path (multimodal=True) keeps Gemma's vision and audio encoders so a call can also take an image or an audio clip; see Experimental: image and audio input.
Download and use
The release is hosted at bambamdevs/openjev-e4b. Install CUDA-enabled PyTorch 2.6 or newer first on Windows (PyTorch instructions). The release contains the adapter, cumulative backbone delta, decision head, calibration, and loader. The loader downloads the pinned google/gemma-4-E4B-it base revision on first use. Hugging Face login is optional for public downloads.
python -m pip install --upgrade "huggingface_hub<2"
hf auth login
hf download bambamdevs/openjev-e4b --local-dir openjev-e4b
cd openjev-e4b
python -m pip install -r requirements.txt
python -m examples.basic_choice
To score your own options:
from openjev import OpenJEV
model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4")
result = model.choice(
state="Checkout fails after a deployment; customers cannot place orders.",
instruction="Which team should own this incident?",
options=["Billing", "Technical outages", "Sales"],
)
print(result["probabilities"], result["selected_option"])
The NF4 path keeps frozen base layers quantized and restores delta-touched layers 37–41 in dense FP16. The decision head runs in FP32. load_mode="bf16" is available with enough memory; its public benchmarks have not been run.
Use as a classifier
Put the text to classify in state, the question in instruction, and your labels in options. Any label set of two or more works on each call; no retraining is needed.
r = model.choice(
state="I was charged twice for my subscription this month.",
instruction="Which support queue should handle this message?",
options=["billing", "technical support", "account access", "sales"],
)
r["selected_option"] # "billing"
r["probabilities"] # one probability per option, in the same order as options
The same pattern covers sentiment, topic, intent, and entailment labels. Probabilities use the release calibration in model/calibration.json.
Experimental: image and audio input
Load with multimodal=True to pass image= (path or PIL image) or audio= (.wav path) to choice():
model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4", multimodal=True)
r = model.choice(
state="",
audio="call.wav",
instruction="Which support queue should handle the spoken message?",
options=["billing", "technical support", "account access", "sales"],
)
The adapter, backbone delta, and decision head were trained on text only, so this path is experimental. In a small smoke test on clean synthetic inputs (NF4, RTX 3060), it scored 12/12 on image shape, 12/12 on image colour, 8/8 on printed messages, 16/16 on spoken messages, and 12/12 on spoken reviews. With the media removed, the same questions scored at chance. Image calls took about 5 s and audio calls about 2 s. This is not a benchmark. Details, limits, and the rerun script are in docs/MULTIMODAL_EXPERIMENTAL.md.
Evaluation
The four choice-panel tasks use frozen rows and cyclic option rotations. For each row, probabilities are mapped back to the original option identities and averaged across rotations; accuracy is calculated from that mean. The table also reports mean accuracy for individual presentations, which exposes option-order sensitivity. A separate options-only condition removes the state and question. The five standard public tasks below were evaluated in their native option order. All model evaluations use NF4 on the Windows RTX 3060.
Choice-panel evaluation
The main score selects the largest probability after averaging each row's cyclic option rotations. Wilson 95% intervals use rows as observations. Both systems use the same frozen rows and Windows NF4 evaluation path.
| Task | Rows | Rotations | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] | Difference [95% CI] |
|---|---|---|---|---|---|---|
| GPQA Diamond, shuffled | 198 | 4 | 25.0% | 32.8% [26.7, 39.6] | 32.8% [26.7, 39.6] | +0.00 [-7.07, +7.07] pp |
| Chess legal | 500 | 4 | 25.0% | 43.8% [39.5, 48.2] | 42.6% [38.3, 47.0] | -1.20 [-6.40, +3.80] pp |
| GSM8K, 4 choices | 1,319 | 4 | 25.0% | 43.7% [41.1, 46.4] | 53.8% [51.1, 56.5] | +10.08 [+6.44, +13.72] pp |
| GSM8K, 10 choices | 1,319 | 10 | 10.0% | 20.5% [18.5, 22.8] | 30.6% [28.2, 33.2] | +10.08 [+7.05, +13.04] pp |
Mean accuracy per presentation and share of rows with the same predicted option in every rotation:
| Task | Gemma per order / same answer | OpenJEV 1.0 per order / same answer |
|---|---|---|
| GPQA Diamond, shuffled | 31.3% / 10.1% | 33.7% / 23.2% |
| Chess legal | 31.5% / 1.4% | 40.1% / 46.6% |
| GSM8K, 4 choices | 36.7% / 6.9% | 51.5% / 43.3% |
| GSM8K, 10 choices | 16.5% / 0.6% | 28.6% / 14.9% |
Options-only control
State and question are replaced by neutral text. Four-choice tasks use four rotations; GSM8K-10 uses five evenly spaced rotations.
| Task | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] |
|---|---|---|---|
| GPQA Diamond, shuffled | 25.0% | 25.8% [20.2, 32.3] | 32.3% [26.2, 39.1] |
| Chess legal | 25.0% | 29.0% [25.2, 33.1] | 30.8% [26.9, 35.0] |
| GSM8K, 4 choices | 25.0% | 24.7% [22.5, 27.1] | 24.3% [22.1, 26.7] |
| GSM8K, 10 choices | 10.0% | 10.5% [9.0, 12.3] | 10.8% [9.3, 12.6] |
OpenJEV 1.0 scores 32.3% on GPQA and 30.8% on Chess legal with only the options present, both above the 25% chance rate. Its GPQA full-condition score is 32.8%; this result alone does not establish use of the question. The GSM8K options-only scores are near chance. The row-construction audit passes its chance gate, but the model control reveals additional choice-only signal. Chess legal tests move legality; the GSM8K tasks test closed-choice arithmetic. Gemma uses a native answer-letter readout and OpenJEV uses a learned decision head, so comparisons measure complete systems.
Evidence: eval/choice_panel/. The five standard public tasks below use native option order and are reported separately.
Standard public tasks
Benchmark status: COMPLETE. All five standard public tasks have both Gemma base NF4 control and OpenJEV 1.0 results.
| Benchmark | Gemma base NF4 | OpenJEV 1.0 | OpenJEV - Gemma |
|---|---|---|---|
| ARC-Challenge | 87.37% | 87.88% | +0.51 pp |
| WinoGrande | 60.14% | 66.46% | +6.32 pp |
| ARC-Easy | 95.58% | 95.62% | +0.04 pp |
| HellaSwag | 76.51% | 89.77% | +13.26 pp |
| MMLU | 66.19% | 64.41% | -1.78 pp |
The untouched base uses native LM-head answer-label scoring; OpenJEV 1.0 uses a learned decision head. Their comparison measures the complete decision systems. ARC, HellaSwag, and MMLU have training-family exposure described in docs/BENCHMARK_EXPOSURE.md.
Paper and evidence
Experiment 1 whitepaper describes the model, training lineage, evaluation protocol, results, limitations, and reproduction steps. It accompanies this release.
model/ adapter, head, cumulative backbone delta, calibration
openjev/ loader and scoring implementation
eval/choice_panel/ frozen rows, audit, scoring code, row-level results
eval/harness/ standard public-task runner
docs/ whitepaper, source, and result notes
provenance/ checkpoint metadata
OpenJEV 1.0 is a delta release; Google base weights are not republished. backbone_delta.safetensors is required to restore the model.
Training and limitations
The model evolved through readout and adapter training, then an Arena curriculum with exact-oracle examples and replay across 20 decision families. The public checkpoint also updates selected decoder layers and includes 65 cumulative backbone-delta tensors. The development sequence, promotion rule, and release hashes are in the whitepaper and RELEASE_METADATA.json.
The experiment has one training seed per stage. Internal populations changed during development. The choice-panel results measure accuracy on their stated tasks; they do not isolate the effect of the head, LoRA, backbone updates, or inference precision. GPQA and GSM8K are multiple-choice tasks, and Chess legal tests move legality rather than move quality.
Reproduce and license
See REPRODUCIBILITY.md, RELEASE_METADATA.json, and MANIFEST.sha256. OpenJEV-authored material is Apache-2.0. The Gemma base remains under Google's terms. Training code is not part of this release.
OpenJEV is independent research. It is not a TypeSafe AI product, is not affiliated with or endorsed by TypeSafe AI, and is unrelated to AlexWortega/openjev. Its design was inspired by System 1-style decision making: a fast, direct choice over a closed set of options, without step-by-step text generation.
Generated: 2026-09-27T17:07:53.606842+00:00
- Downloads last month
- 45