openjev-e4b / README.md
bambamdevs's picture
Publish OpenJEV E4B 1.0
03223d7
|
Raw History Blame Contribute Delete
9.9 kB
---
license: apache-2.0
library_name: openjev
pipeline_tag: zero-shot-classification
base_model: google/gemma-4-E4B-it
language:
- en
tags:
- openjev
- classification
- zero-shot-classification
- decision-model
- listwise
- peft
- gemma4
- research
---
# OpenJEV E4B 1.0
OpenJEV turns Gemma 4 E4B into a direct-decision classifier. A caller supplies state, a question, and a closed set of options; the model returns a probability for each option without generating text. This is the public 1.0 release of Experiment 1.
OpenJEV was trained and evaluated on text. An optional, experimental path (`multimodal=True`) keeps Gemma's vision and audio encoders so a call can also take an image or an audio clip; see [Experimental: image and audio input](#experimental-image-and-audio-input).
## Download and use
The release is hosted at [bambamdevs/openjev-e4b](https://huggingface.co/bambamdevs/openjev-e4b). Install CUDA-enabled PyTorch 2.6 or newer first on Windows ([PyTorch instructions](https://pytorch.org/get-started/locally/)). The release contains the adapter, cumulative backbone delta, decision head, calibration, and loader. The loader downloads the pinned `google/gemma-4-E4B-it` base revision on first use. Hugging Face login is optional for public downloads.
```bash
python -m pip install --upgrade "huggingface_hub<2"
hf auth login
hf download bambamdevs/openjev-e4b --local-dir openjev-e4b
cd openjev-e4b
python -m pip install -r requirements.txt
python -m examples.basic_choice
```
To score your own options:
```python
from openjev import OpenJEV
model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4")
result = model.choice(
state="Checkout fails after a deployment; customers cannot place orders.",
instruction="Which team should own this incident?",
options=["Billing", "Technical outages", "Sales"],
)
print(result["probabilities"], result["selected_option"])
```
The NF4 path keeps frozen base layers quantized and restores delta-touched layers 37–41 in dense FP16. The decision head runs in FP32. `load_mode="bf16"` is available with enough memory; its public benchmarks have not been run.
### Use as a classifier
Put the text to classify in `state`, the question in `instruction`, and your labels in `options`. Any label set of two or more works on each call; no retraining is needed.
```python
r = model.choice(
state="I was charged twice for my subscription this month.",
instruction="Which support queue should handle this message?",
options=["billing", "technical support", "account access", "sales"],
)
r["selected_option"] # "billing"
r["probabilities"] # one probability per option, in the same order as options
```
The same pattern covers sentiment, topic, intent, and entailment labels. Probabilities use the release calibration in `model/calibration.json`.
### Experimental: image and audio input
Load with `multimodal=True` to pass `image=` (path or PIL image) or `audio=` (`.wav` path) to `choice()`:
```python
model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4", multimodal=True)
r = model.choice(
state="",
audio="call.wav",
instruction="Which support queue should handle the spoken message?",
options=["billing", "technical support", "account access", "sales"],
)
```
The adapter, backbone delta, and decision head were trained on text only, so this path is experimental. In a small smoke test on clean synthetic inputs (NF4, RTX 3060), it scored 12/12 on image shape, 12/12 on image colour, 8/8 on printed messages, 16/16 on spoken messages, and 12/12 on spoken reviews. With the media removed, the same questions scored at chance. Image calls took about 5 s and audio calls about 2 s. This is not a benchmark. Details, limits, and the rerun script are in [`docs/MULTIMODAL_EXPERIMENTAL.md`](./docs/MULTIMODAL_EXPERIMENTAL.md).
## Evaluation
The four choice-panel tasks use frozen rows and cyclic option rotations. For each row, probabilities are mapped back to the original option identities and averaged across rotations; accuracy is calculated from that mean. The table also reports mean accuracy for individual presentations, which exposes option-order sensitivity. A separate options-only condition removes the state and question. The five standard public tasks below were evaluated in their native option order. All model evaluations use NF4 on the Windows RTX 3060.
<!-- CHOICE_RESULTS_START -->
### Choice-panel evaluation
The main score selects the largest probability after averaging each row's cyclic option rotations. Wilson 95% intervals use rows as observations. Both systems use the same frozen rows and Windows NF4 evaluation path.
| Task | Rows | Rotations | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] | Difference [95% CI] |
|---|---:|---:|---:|---:|---:|---:|
| GPQA Diamond, shuffled | 198 | 4 | 25.0% | 32.8% [26.7, 39.6] | 32.8% [26.7, 39.6] | +0.00 [-7.07, +7.07] pp |
| Chess legal | 500 | 4 | 25.0% | 43.8% [39.5, 48.2] | 42.6% [38.3, 47.0] | -1.20 [-6.40, +3.80] pp |
| GSM8K, 4 choices | 1,319 | 4 | 25.0% | 43.7% [41.1, 46.4] | 53.8% [51.1, 56.5] | +10.08 [+6.44, +13.72] pp |
| GSM8K, 10 choices | 1,319 | 10 | 10.0% | 20.5% [18.5, 22.8] | 30.6% [28.2, 33.2] | +10.08 [+7.05, +13.04] pp |
Mean accuracy per presentation and share of rows with the same predicted option in every rotation:
| Task | Gemma per order / same answer | OpenJEV 1.0 per order / same answer |
|---|---:|---:|
| GPQA Diamond, shuffled | 31.3% / 10.1% | 33.7% / 23.2% |
| Chess legal | 31.5% / 1.4% | 40.1% / 46.6% |
| GSM8K, 4 choices | 36.7% / 6.9% | 51.5% / 43.3% |
| GSM8K, 10 choices | 16.5% / 0.6% | 28.6% / 14.9% |
### Options-only control
State and question are replaced by neutral text. Four-choice tasks use four rotations; GSM8K-10 uses five evenly spaced rotations.
| Task | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] |
|---|---:|---:|---:|
| GPQA Diamond, shuffled | 25.0% | 25.8% [20.2, 32.3] | 32.3% [26.2, 39.1] |
| Chess legal | 25.0% | 29.0% [25.2, 33.1] | 30.8% [26.9, 35.0] |
| GSM8K, 4 choices | 25.0% | 24.7% [22.5, 27.1] | 24.3% [22.1, 26.7] |
| GSM8K, 10 choices | 10.0% | 10.5% [9.0, 12.3] | 10.8% [9.3, 12.6] |
OpenJEV 1.0 scores 32.3% on GPQA and 30.8% on Chess legal with only the options present, both above the 25% chance rate. Its GPQA full-condition score is 32.8%; this result alone does not establish use of the question. The GSM8K options-only scores are near chance. The row-construction audit passes its chance gate, but the model control reveals additional choice-only signal. Chess legal tests move legality; the GSM8K tasks test closed-choice arithmetic. Gemma uses a native answer-letter readout and OpenJEV uses a learned decision head, so comparisons measure complete systems.
Evidence: [`eval/choice_panel/`](./eval/choice_panel/). The five standard public tasks below use native option order and are reported separately.
<!-- CHOICE_RESULTS_END -->
### Standard public tasks
**Benchmark status: COMPLETE.** All five standard public tasks have both Gemma base NF4 control and OpenJEV 1.0 results.
| Benchmark | Gemma base NF4 | OpenJEV 1.0 | OpenJEV - Gemma |
|---|---:|---:|---:|
| ARC-Challenge | 87.37% | 87.88% | +0.51 pp |
| WinoGrande | 60.14% | 66.46% | +6.32 pp |
| ARC-Easy | 95.58% | 95.62% | +0.04 pp |
| HellaSwag | 76.51% | 89.77% | +13.26 pp |
| MMLU | 66.19% | 64.41% | -1.78 pp |
The untouched base uses native LM-head answer-label scoring; OpenJEV 1.0 uses a learned decision head. Their comparison measures the complete decision systems. ARC, HellaSwag, and MMLU have training-family exposure described in [`docs/BENCHMARK_EXPOSURE.md`](./docs/BENCHMARK_EXPOSURE.md).
## Paper and evidence
[Experiment 1 whitepaper](./docs/OpenJEV_E4B_Experiment_1_Whitepaper.pdf) describes the model, training lineage, evaluation protocol, results, limitations, and reproduction steps. It accompanies this release.
```text
model/ adapter, head, cumulative backbone delta, calibration
openjev/ loader and scoring implementation
eval/choice_panel/ frozen rows, audit, scoring code, row-level results
eval/harness/ standard public-task runner
docs/ whitepaper, source, and result notes
provenance/ checkpoint metadata
```
OpenJEV 1.0 is a delta release; Google base weights are not republished. `backbone_delta.safetensors` is required to restore the model.
## Training and limitations
The model evolved through readout and adapter training, then an Arena curriculum with exact-oracle examples and replay across 20 decision families. The public checkpoint also updates selected decoder layers and includes 65 cumulative backbone-delta tensors. The development sequence, promotion rule, and release hashes are in the whitepaper and `RELEASE_METADATA.json`.
The experiment has one training seed per stage. Internal populations changed during development. The choice-panel results measure accuracy on their stated tasks; they do not isolate the effect of the head, LoRA, backbone updates, or inference precision. GPQA and GSM8K are multiple-choice tasks, and Chess legal tests move legality rather than move quality.
## Reproduce and license
See [`REPRODUCIBILITY.md`](./REPRODUCIBILITY.md), [`RELEASE_METADATA.json`](./RELEASE_METADATA.json), and [`MANIFEST.sha256`](./MANIFEST.sha256). OpenJEV-authored material is Apache-2.0. The Gemma base remains under Google's terms. Training code is not part of this release.
OpenJEV is independent research. It is not a TypeSafe AI product, is not affiliated with or endorsed by TypeSafe AI, and is unrelated to `AlexWortega/openjev`. Its design was inspired by System 1-style decision making: a fast, direct choice over a closed set of options, without step-by-step text generation.
Generated: 2026-09-27T17:07:53.606842+00:00