Zero-Shot Classification
Safetensors
PEFT
English
openjev
classification
decision-model
listwise
gemma4
research
Instructions to use bambamdevs/openjev-e4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use bambamdevs/openjev-e4b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download README.md from bambamdevs/openjev-e4b: direct link, hf CLI and curl.
- Browser
- Download file 9.9 kB
-
https://huggingface.co/bambamdevs/openjev-e4b/resolve/main/README.md
- Command line
-
hf download hf://bambamdevs/openjev-e4b/README.md
-
curl -L -o README.md https://huggingface.co/bambamdevs/openjev-e4b/resolve/main/README.md
9.9 kB
| license: apache-2.0 | |
| library_name: openjev | |
| pipeline_tag: zero-shot-classification | |
| base_model: google/gemma-4-E4B-it | |
| language: | |
| - en | |
| tags: | |
| - openjev | |
| - classification | |
| - zero-shot-classification | |
| - decision-model | |
| - listwise | |
| - peft | |
| - gemma4 | |
| - research | |
| # OpenJEV E4B 1.0 | |
| OpenJEV turns Gemma 4 E4B into a direct-decision classifier. A caller supplies state, a question, and a closed set of options; the model returns a probability for each option without generating text. This is the public 1.0 release of Experiment 1. | |
| OpenJEV was trained and evaluated on text. An optional, experimental path (`multimodal=True`) keeps Gemma's vision and audio encoders so a call can also take an image or an audio clip; see [Experimental: image and audio input](#experimental-image-and-audio-input). | |
| ## Download and use | |
| The release is hosted at [bambamdevs/openjev-e4b](https://huggingface.co/bambamdevs/openjev-e4b). Install CUDA-enabled PyTorch 2.6 or newer first on Windows ([PyTorch instructions](https://pytorch.org/get-started/locally/)). The release contains the adapter, cumulative backbone delta, decision head, calibration, and loader. The loader downloads the pinned `google/gemma-4-E4B-it` base revision on first use. Hugging Face login is optional for public downloads. | |
| ```bash | |
| python -m pip install --upgrade "huggingface_hub<2" | |
| hf auth login | |
| hf download bambamdevs/openjev-e4b --local-dir openjev-e4b | |
| cd openjev-e4b | |
| python -m pip install -r requirements.txt | |
| python -m examples.basic_choice | |
| ``` | |
| To score your own options: | |
| ```python | |
| from openjev import OpenJEV | |
| model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4") | |
| result = model.choice( | |
| state="Checkout fails after a deployment; customers cannot place orders.", | |
| instruction="Which team should own this incident?", | |
| options=["Billing", "Technical outages", "Sales"], | |
| ) | |
| print(result["probabilities"], result["selected_option"]) | |
| ``` | |
| The NF4 path keeps frozen base layers quantized and restores delta-touched layers 37–41 in dense FP16. The decision head runs in FP32. `load_mode="bf16"` is available with enough memory; its public benchmarks have not been run. | |
| ### Use as a classifier | |
| Put the text to classify in `state`, the question in `instruction`, and your labels in `options`. Any label set of two or more works on each call; no retraining is needed. | |
| ```python | |
| r = model.choice( | |
| state="I was charged twice for my subscription this month.", | |
| instruction="Which support queue should handle this message?", | |
| options=["billing", "technical support", "account access", "sales"], | |
| ) | |
| r["selected_option"] # "billing" | |
| r["probabilities"] # one probability per option, in the same order as options | |
| ``` | |
| The same pattern covers sentiment, topic, intent, and entailment labels. Probabilities use the release calibration in `model/calibration.json`. | |
| ### Experimental: image and audio input | |
| Load with `multimodal=True` to pass `image=` (path or PIL image) or `audio=` (`.wav` path) to `choice()`: | |
| ```python | |
| model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4", multimodal=True) | |
| r = model.choice( | |
| state="", | |
| audio="call.wav", | |
| instruction="Which support queue should handle the spoken message?", | |
| options=["billing", "technical support", "account access", "sales"], | |
| ) | |
| ``` | |
| The adapter, backbone delta, and decision head were trained on text only, so this path is experimental. In a small smoke test on clean synthetic inputs (NF4, RTX 3060), it scored 12/12 on image shape, 12/12 on image colour, 8/8 on printed messages, 16/16 on spoken messages, and 12/12 on spoken reviews. With the media removed, the same questions scored at chance. Image calls took about 5 s and audio calls about 2 s. This is not a benchmark. Details, limits, and the rerun script are in [`docs/MULTIMODAL_EXPERIMENTAL.md`](./docs/MULTIMODAL_EXPERIMENTAL.md). | |
| ## Evaluation | |
| The four choice-panel tasks use frozen rows and cyclic option rotations. For each row, probabilities are mapped back to the original option identities and averaged across rotations; accuracy is calculated from that mean. The table also reports mean accuracy for individual presentations, which exposes option-order sensitivity. A separate options-only condition removes the state and question. The five standard public tasks below were evaluated in their native option order. All model evaluations use NF4 on the Windows RTX 3060. | |
| <!-- CHOICE_RESULTS_START --> | |
| ### Choice-panel evaluation | |
| The main score selects the largest probability after averaging each row's cyclic option rotations. Wilson 95% intervals use rows as observations. Both systems use the same frozen rows and Windows NF4 evaluation path. | |
| | Task | Rows | Rotations | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] | Difference [95% CI] | | |
| |---|---:|---:|---:|---:|---:|---:| | |
| | GPQA Diamond, shuffled | 198 | 4 | 25.0% | 32.8% [26.7, 39.6] | 32.8% [26.7, 39.6] | +0.00 [-7.07, +7.07] pp | | |
| | Chess legal | 500 | 4 | 25.0% | 43.8% [39.5, 48.2] | 42.6% [38.3, 47.0] | -1.20 [-6.40, +3.80] pp | | |
| | GSM8K, 4 choices | 1,319 | 4 | 25.0% | 43.7% [41.1, 46.4] | 53.8% [51.1, 56.5] | +10.08 [+6.44, +13.72] pp | | |
| | GSM8K, 10 choices | 1,319 | 10 | 10.0% | 20.5% [18.5, 22.8] | 30.6% [28.2, 33.2] | +10.08 [+7.05, +13.04] pp | | |
| Mean accuracy per presentation and share of rows with the same predicted option in every rotation: | |
| | Task | Gemma per order / same answer | OpenJEV 1.0 per order / same answer | | |
| |---|---:|---:| | |
| | GPQA Diamond, shuffled | 31.3% / 10.1% | 33.7% / 23.2% | | |
| | Chess legal | 31.5% / 1.4% | 40.1% / 46.6% | | |
| | GSM8K, 4 choices | 36.7% / 6.9% | 51.5% / 43.3% | | |
| | GSM8K, 10 choices | 16.5% / 0.6% | 28.6% / 14.9% | | |
| ### Options-only control | |
| State and question are replaced by neutral text. Four-choice tasks use four rotations; GSM8K-10 uses five evenly spaced rotations. | |
| | Task | Chance | Gemma base [95% CI] | OpenJEV 1.0 [95% CI] | | |
| |---|---:|---:|---:| | |
| | GPQA Diamond, shuffled | 25.0% | 25.8% [20.2, 32.3] | 32.3% [26.2, 39.1] | | |
| | Chess legal | 25.0% | 29.0% [25.2, 33.1] | 30.8% [26.9, 35.0] | | |
| | GSM8K, 4 choices | 25.0% | 24.7% [22.5, 27.1] | 24.3% [22.1, 26.7] | | |
| | GSM8K, 10 choices | 10.0% | 10.5% [9.0, 12.3] | 10.8% [9.3, 12.6] | | |
| OpenJEV 1.0 scores 32.3% on GPQA and 30.8% on Chess legal with only the options present, both above the 25% chance rate. Its GPQA full-condition score is 32.8%; this result alone does not establish use of the question. The GSM8K options-only scores are near chance. The row-construction audit passes its chance gate, but the model control reveals additional choice-only signal. Chess legal tests move legality; the GSM8K tasks test closed-choice arithmetic. Gemma uses a native answer-letter readout and OpenJEV uses a learned decision head, so comparisons measure complete systems. | |
| Evidence: [`eval/choice_panel/`](./eval/choice_panel/). The five standard public tasks below use native option order and are reported separately. | |
| <!-- CHOICE_RESULTS_END --> | |
| ### Standard public tasks | |
| **Benchmark status: COMPLETE.** All five standard public tasks have both Gemma base NF4 control and OpenJEV 1.0 results. | |
| | Benchmark | Gemma base NF4 | OpenJEV 1.0 | OpenJEV - Gemma | | |
| |---|---:|---:|---:| | |
| | ARC-Challenge | 87.37% | 87.88% | +0.51 pp | | |
| | WinoGrande | 60.14% | 66.46% | +6.32 pp | | |
| | ARC-Easy | 95.58% | 95.62% | +0.04 pp | | |
| | HellaSwag | 76.51% | 89.77% | +13.26 pp | | |
| | MMLU | 66.19% | 64.41% | -1.78 pp | | |
| The untouched base uses native LM-head answer-label scoring; OpenJEV 1.0 uses a learned decision head. Their comparison measures the complete decision systems. ARC, HellaSwag, and MMLU have training-family exposure described in [`docs/BENCHMARK_EXPOSURE.md`](./docs/BENCHMARK_EXPOSURE.md). | |
| ## Paper and evidence | |
| [Experiment 1 whitepaper](./docs/OpenJEV_E4B_Experiment_1_Whitepaper.pdf) describes the model, training lineage, evaluation protocol, results, limitations, and reproduction steps. It accompanies this release. | |
| ```text | |
| model/ adapter, head, cumulative backbone delta, calibration | |
| openjev/ loader and scoring implementation | |
| eval/choice_panel/ frozen rows, audit, scoring code, row-level results | |
| eval/harness/ standard public-task runner | |
| docs/ whitepaper, source, and result notes | |
| provenance/ checkpoint metadata | |
| ``` | |
| OpenJEV 1.0 is a delta release; Google base weights are not republished. `backbone_delta.safetensors` is required to restore the model. | |
| ## Training and limitations | |
| The model evolved through readout and adapter training, then an Arena curriculum with exact-oracle examples and replay across 20 decision families. The public checkpoint also updates selected decoder layers and includes 65 cumulative backbone-delta tensors. The development sequence, promotion rule, and release hashes are in the whitepaper and `RELEASE_METADATA.json`. | |
| The experiment has one training seed per stage. Internal populations changed during development. The choice-panel results measure accuracy on their stated tasks; they do not isolate the effect of the head, LoRA, backbone updates, or inference precision. GPQA and GSM8K are multiple-choice tasks, and Chess legal tests move legality rather than move quality. | |
| ## Reproduce and license | |
| See [`REPRODUCIBILITY.md`](./REPRODUCIBILITY.md), [`RELEASE_METADATA.json`](./RELEASE_METADATA.json), and [`MANIFEST.sha256`](./MANIFEST.sha256). OpenJEV-authored material is Apache-2.0. The Gemma base remains under Google's terms. Training code is not part of this release. | |
| OpenJEV is independent research. It is not a TypeSafe AI product, is not affiliated with or endorsed by TypeSafe AI, and is unrelated to `AlexWortega/openjev`. Its design was inspired by System 1-style decision making: a fast, direct choice over a closed set of options, without step-by-step text generation. | |
| Generated: 2026-09-27T17:07:53.606842+00:00 | |