Zero-Shot Classification
Safetensors
PEFT
English
openjev
classification
decision-model
listwise
gemma4
research
Instructions to use bambamdevs/openjev-e4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use bambamdevs/openjev-e4b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download docs/MULTIMODAL_EXPERIMENTAL.md from bambamdevs/openjev-e4b: direct link, hf CLI and curl.
- Browser
- Download file 3.45 kB
-
https://huggingface.co/bambamdevs/openjev-e4b/resolve/main/docs/MULTIMODAL_EXPERIMENTAL.md
- Command line
-
hf download hf://bambamdevs/openjev-e4b/docs/MULTIMODAL_EXPERIMENTAL.md
-
curl -L -o MULTIMODAL_EXPERIMENTAL.md https://huggingface.co/bambamdevs/openjev-e4b/resolve/main/docs/MULTIMODAL_EXPERIMENTAL.md
3.45 kB
| # Experimental image and audio input | |
| OpenJEV E4B 1.0 was trained and evaluated on text. The loader can optionally keep Gemma 4's own vision and audio encoders, so `choice()` also accepts an image or a short audio clip. This path is **experimental**: none of the adapter, backbone delta, or decision head was trained on image or audio input, and it has only passed the small smoke test below. | |
| ## Use | |
| ```python | |
| from openjev import OpenJEV | |
| model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4", multimodal=True) | |
| model.choice(state="", image="photo.png", | |
| instruction="What shape is shown in the image?", | |
| options=["square", "circle", "triangle"]) | |
| model.choice(state="", audio="call.wav", | |
| instruction="Which support queue should handle the spoken message?", | |
| options=["billing", "technical support", "account access", "sales"]) | |
| ``` | |
| - `image` is a file path or a PIL image. `audio` is a PCM `.wav` path, or an array plus `sampling_rate=`; it is converted to 16 kHz mono. | |
| - `state` can hold extra text alongside the media. The media tokens are placed at the start of the state. | |
| - Without `multimodal=True`, passing `image=` or `audio=` raises an error. Text calls behave the same in both modes. | |
| ## How it works | |
| The text-only loader discards Gemma's vision and audio encoders. With `multimodal=True`, it keeps them, wraps Gemma's text model in place with the OpenJEV adapter and backbone delta, and inserts Gemma's image or audio soft tokens after `State:` in the usual decision prompt. The decision head reads the same option and decision positions as for text. | |
| In NF4 mode the encoders stay in FP16, because Gemma's audio encoder cannot run with 4-bit weights. The text layers are quantized exactly as in the text-only loader. | |
| ## Smoke test | |
| Run on 2026-09-27 with NF4 on an RTX 3060 12 GB. Images are generated shapes and printed messages. Audio is the same messages and short reviews spoken by the two built-in Windows voices. Every question was also asked with the media removed. | |
| | Test | With media | Media removed | | |
| |---|---:|---:| | |
| | Image: shape, 3 options | 12/12 | 4/12 | | |
| | Image: colour, 4 options | 12/12 | 3/12 | | |
| | Image: printed support message → queue, 4 options | 8/8 | 2/8 | | |
| | Audio: spoken support message → queue, 4 options | 16/16 | 2/8 | | |
| | Audio: spoken review → sentiment, 2 options | 12/12 | 3/6 | | |
| - For comparison, the same support messages given as text scored 7/8, and the reviews as text scored 6/6. | |
| - For text-only input, the multimodal forward path gives the same probabilities as the text path (maximum difference 0). | |
| - Median latency per call: about 4.9 s with an image, 2.3 s with audio, and 0.2 s for text. Peak GPU memory was 10.7 GB. | |
| Full rows are in [`eval/multimodal_smoke/results.json`](../eval/multimodal_smoke/results.json). To rerun: run `eval/multimodal_smoke/make_audio_windows.ps1`, then `python -m eval.multimodal_smoke.run_smoke` from the release root. | |
| ## Limits | |
| - These inputs are clean, synthetic, and easy. The test shows that the path works end to end; it is not a benchmark, and there are no results on real photos, noisy audio, or public image or audio datasets. | |
| - Probabilities use the text calibration, which has not been checked for image or audio input. | |
| - Only NF4 on one GPU was tested; `load_mode="bf16"` with `multimodal=True` has not been run. | |
| - One image and one audio clip per call. Video is not supported. | |