openjev-e4b / docs /MULTIMODAL_EXPERIMENTAL.md
bambamdevs's picture
Publish OpenJEV E4B 1.0
03223d7
|
Raw History Blame Contribute Delete
3.45 kB

Experimental image and audio input

OpenJEV E4B 1.0 was trained and evaluated on text. The loader can optionally keep Gemma 4's own vision and audio encoders, so choice() also accepts an image or a short audio clip. This path is experimental: none of the adapter, backbone delta, or decision head was trained on image or audio input, and it has only passed the small smoke test below.

Use

from openjev import OpenJEV

model = OpenJEV.from_pretrained(".", device="cuda", load_mode="nf4", multimodal=True)

model.choice(state="", image="photo.png",
             instruction="What shape is shown in the image?",
             options=["square", "circle", "triangle"])

model.choice(state="", audio="call.wav",
             instruction="Which support queue should handle the spoken message?",
             options=["billing", "technical support", "account access", "sales"])
  • image is a file path or a PIL image. audio is a PCM .wav path, or an array plus sampling_rate=; it is converted to 16 kHz mono.
  • state can hold extra text alongside the media. The media tokens are placed at the start of the state.
  • Without multimodal=True, passing image= or audio= raises an error. Text calls behave the same in both modes.

How it works

The text-only loader discards Gemma's vision and audio encoders. With multimodal=True, it keeps them, wraps Gemma's text model in place with the OpenJEV adapter and backbone delta, and inserts Gemma's image or audio soft tokens after State: in the usual decision prompt. The decision head reads the same option and decision positions as for text.

In NF4 mode the encoders stay in FP16, because Gemma's audio encoder cannot run with 4-bit weights. The text layers are quantized exactly as in the text-only loader.

Smoke test

Run on 2026-09-27 with NF4 on an RTX 3060 12 GB. Images are generated shapes and printed messages. Audio is the same messages and short reviews spoken by the two built-in Windows voices. Every question was also asked with the media removed.

Test With media Media removed
Image: shape, 3 options 12/12 4/12
Image: colour, 4 options 12/12 3/12
Image: printed support message → queue, 4 options 8/8 2/8
Audio: spoken support message → queue, 4 options 16/16 2/8
Audio: spoken review → sentiment, 2 options 12/12 3/6
  • For comparison, the same support messages given as text scored 7/8, and the reviews as text scored 6/6.
  • For text-only input, the multimodal forward path gives the same probabilities as the text path (maximum difference 0).
  • Median latency per call: about 4.9 s with an image, 2.3 s with audio, and 0.2 s for text. Peak GPU memory was 10.7 GB.

Full rows are in eval/multimodal_smoke/results.json. To rerun: run eval/multimodal_smoke/make_audio_windows.ps1, then python -m eval.multimodal_smoke.run_smoke from the release root.

Limits

  • These inputs are clean, synthetic, and easy. The test shows that the path works end to end; it is not a benchmark, and there are no results on real photos, noisy audio, or public image or audio datasets.
  • Probabilities use the text calibration, which has not been checked for image or audio input.
  • Only NF4 on one GPU was tested; load_mode="bf16" with multimodal=True has not been run.
  • One image and one audio clip per call. Video is not supported.