Instructions to use mohdyaser/jev-multimodal with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mohdyaser/jev-multimodal with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("mohdyaser/jev-multimodal") config = load_config("mohdyaser/jev-multimodal") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
jev-multimodal
An experimental Gemma 4 E4B instruction model LoRA adapter for typed decisions from text and up to four ordered images. It is inspired by Jev's noul (yes/no), choice, and score interaction pattern. This is an independent prototype; it is not the Jev model and has no affiliation with its creators.
This repository contains only the native MLX-VLM adapter files (adapters.safetensors and adapter_config.json). It does not contain the approximately 5 GB base checkpoint, a merged model, or a PEFT/Transformers adapter. The adapter needs the matching 4-bit base and the decision prompt/scoring code in the project source.
Base and adapter
| Item | Value |
|---|---|
| Base | mlx-community/gemma-4-e4b-it-4bit |
| Base revision | 475b9088d29754a3379866cf5aeb6b41acd313c2 |
| Original model | google/gemma-4-E4B-it |
| Adapter format | MLX-VLM LoRA, rank 8, scale 2 (alpha 16), dropout 0 |
| Tuned modules | Query and output projections in language layers 34โ41 |
| Weight SHA-256 | 87ef1029610d23688512883700c35aefa60f751850e8483225a64354701c3e45 |
| Experiment | decision-v1, selected 2,000-step checkpoint |
The base model's own model card and terms apply when obtaining and using its weights. This adapter was trained with Python 3.11, MLX-VLM 0.7.3, and MLX 0.32.2 on Apple Silicon/macOS.
Use
Download the pinned base and adapter into separate folders:
hf download mlx-community/gemma-4-e4b-it-4bit \
--revision 475b9088d29754a3379866cf5aeb6b41acd313c2 \
--local-dir models/gemma-4-e4b-it-4bit
hf download mohdyaser/jev-multimodal --local-dir models/jev-multimodal
From the project source, with its dependencies installed, a request JSON can be scored with:
.venv/bin/python src/jev_cli.py \
--model models/gemma-4-e4b-it-4bit \
--adapter models/jev-multimodal \
--temperature 1.4914477134397612 \
--input request.json
request.json contains a question, optional context, an ordered list of 0โ4 image paths, an output_type (noul, choice, or score), and 2โ16 { "id", "description" } options. For noul, option IDs are false and true; score may also specify increasing level_values. See the project README for the local UI and examples. The code builds the decision prompt, obtains the next-token logits for one-token option markers AโP in a single forward pass, then normalizes across the supplied options. The reported probabilities are conditional on those options; they are not confidence guarantees. The temperature above was fitted on the experiment's 100-example synthetic calibration split.
Training and evaluation
The frozen 4-bit base was adapted for one epoch on 2,000 locally generated synthetic decisions (0โ4 images), with 250 optimizer updates. A 100-example development split selected this checkpoint. An independent 100-example split fitted temperature, and a 200-example synthetic test split measured the following:
| Metric | Frozen base | This adapter |
|---|---|---|
| Decision accuracy | 158/200 (79%) | 158/200 (79%) |
| Calibrated negative log likelihood | 0.551 | 0.494 |
| Calibrated Brier score | 0.289 | 0.272 |
| Score mean absolute error (63 cases) | 0.549 | 0.487 |
| Paired image probes | 6/9 | 6/9 |
The adapter improved probability quality on this small synthetic test but did not improve answer accuracy or paired image probe performance. Three- and four-image cases remained weak. All evaluated images were synthetic; no rights-checked natural-image test was included. These results do not establish performance on real photos, documents, dense OCR, unfamiliar prompts, or Jev-equivalent behavior. See the experiment report for the protocol and per-category results.
Quantized
Model tree for mohdyaser/jev-multimodal
Base model
google/gemma-4-E4B