PEFT
Safetensors

JEMM

Like Jev, but multimodal and open-weight.

JEMM makes the same kind of fast, structured decision as Jev: given a state and a question, it picks one candidate and returns a probability for every candidate. It can also look at a screenshot, and it runs on your own GPU.

JEMM Jev
Input Text, or text + screenshot Text only
Weights Open, self-hosted Hosted API
Question types choice, noul, score choice, noul, score

JEMM stands for Judgment Engine for MultiModal decisions. It is a LoRA adapter for Qwen/Qwen3.8-27B. Not affiliated with TypeSafe AI.

Results

JEMM vs open Jev-like models

Accuracy of JEMM vs Jev 1.13

Latency of JEMM vs Jev 1.13

Usage

import json, torch
from huggingface_hub import hf_hub_download
from peft import PeftModel
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

base, adapter = "Qwen/Qwen3.8-27B", "MaestroYan/JEMM"
processor = AutoProcessor.from_pretrained(base)
model = Qwen3_5ForConditionalGeneration.from_pretrained(base, dtype=torch.bfloat16, device_map="cuda")
model = PeftModel.from_pretrained(model, adapter).eval()
config = json.load(open(hf_hub_download(adapter, "decision_config.json")))
labels = [chr(65 + i) for i in range(26)] + list("012345")
system = "Choose the best available candidate for the question using only the supplied state. Return exactly one candidate label."


def decide(state, question, candidates, image=None):
    flat = lambda s: " ".join(s.split())
    options = "\n".join(f"{labels[i]}) {flat(c)}" for i, c in enumerate(candidates))
    text = f"State:\n{state.strip()}\n\nQuestion: {flat(question)}\n\nCandidates:\n{options}\n\nAnswer with exactly one candidate label."
    content = ([{"type": "image"}] if image else []) + [{"type": "text", "text": text}]
    messages = [{"role": "system", "content": system}, {"role": "user", "content": content}]
    prompt = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
    inputs = processor(text=[prompt], images=[image] if image else None, return_tensors="pt").to(model.device, dtype=torch.bfloat16)
    with torch.inference_mode():
        logits = model(**inputs, logits_to_keep=1).logits[0, -1]
    ids = [processor.tokenizer.encode(x, add_special_tokens=False)[0] for x in labels[:len(candidates)]]
    probs = torch.softmax(logits[ids].float() / config["mm_temperature" if image else "temperature"], -1)
    return dict(zip(candidates, probs.tolist()))


decide("User: Will it rain in Paris tomorrow?", "Which tool should handle the request?",
       ["get_weather: weather forecast for a city", "search_flights: flight search", "No tool applies"])

2 to 32 candidates per question. Screenshots are PIL images; training used 1280x800. Treat a top probability below threshold in decision_config.json as undecided. HTTP server: JEMM on GitHub.

Training data

Mind2Web, xLAM function-calling-60k (APIGen), xlam-irrelevance, When2Call, Banking77, MASSIVE, Aegis 2.0 (CC-BY-4.0); CLINC150 (CC-BY-3.0); BFCL, glaive-function-calling-v2, ToolACE, hermes-function-calling-v1 (Apache-2.0); Multimodal-Mind2Web (OpenRAIL). No Jev outputs were used.

Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MaestroYan/JEMM

Base model

Qwen/Qwen3.8-27B
Adapter
(131)
this model

Space using MaestroYan/JEMM 1