# Experimental image and audio input OpenJEV E4B 1.0 was trained and evaluated on text. The loader can optionally keep Gemma 4's own vision and audio encoders, so `choice()` also accepts an image or a short audio clip. This path is **experimental**: none of the adapter, backbone delta, or decision head was trained on image or audio input, and it has only passed the small smoke test below. ## Use ```python from openjev import OpenJEV model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4", multimodal=True) model.choice(state="", image="photo.png", instruction="What shape is shown in the image?", options=["square", "circle", "triangle"]) model.choice(state="", audio="call.wav", instruction="Which support queue should handle the spoken message?", options=["billing", "technical support", "account access", "sales"]) ``` - `image` is a file path or a PIL image. `audio` is a PCM `.wav` path, or an array plus `sampling_rate=`; it is converted to 16 kHz mono. - `state` can hold extra text alongside the media. The media tokens are placed at the start of the state. - Without `multimodal=True`, passing `image=` or `audio=` raises an error. Text calls behave the same in both modes. ## How it works The text-only loader discards Gemma's vision and audio encoders. With `multimodal=True`, it keeps them, wraps Gemma's text model in place with the OpenJEV adapter and backbone delta, and inserts Gemma's image or audio soft tokens after `State:` in the usual decision prompt. The decision head reads the same option and decision positions as for text. In NF4 mode the encoders stay in FP16, because Gemma's audio encoder cannot run with 4-bit weights. The text layers are quantized exactly as in the text-only loader. ## Smoke test Run on 2026-09-27 with NF4 on an RTX 3060 12 GB. Images are generated shapes and printed messages. Audio is the same messages and short reviews spoken by the two built-in Windows voices. Every question was also asked with the media removed. | Test | With media | Media removed | |---|---:|---:| | Image: shape, 3 options | 12/12 | 4/12 | | Image: colour, 4 options | 12/12 | 3/12 | | Image: printed support message → queue, 4 options | 8/8 | 2/8 | | Audio: spoken support message → queue, 4 options | 16/16 | 2/8 | | Audio: spoken review → sentiment, 2 options | 12/12 | 3/6 | - For comparison, the same support messages given as text scored 7/8, and the reviews as text scored 6/6. - For text-only input, the multimodal forward path gives the same probabilities as the text path (maximum difference 0). - Median latency per call: about 4.9 s with an image, 2.3 s with audio, and 0.2 s for text. Peak GPU memory was 10.7 GB. Full rows are in [`eval/multimodal_smoke/results.json`](../eval/multimodal_smoke/results.json). To rerun: run `eval/multimodal_smoke/make_audio_windows.ps1`, then `python -m eval.multimodal_smoke.run_smoke` from the release root. ## Limits - These inputs are clean, synthetic, and easy. The test shows that the path works end to end; it is not a benchmark, and there are no results on real photos, noisy audio, or public image or audio datasets. - Probabilities use the text calibration, which has not been checked for image or audio input. - Only NF4 on one GPU was tested; `load_mode="bf16"` with `multimodal=True` has not been run. - One image and one audio clip per call. Video is not supported.