Instructions to use MaestroYan/JEMM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use MaestroYan/JEMM with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.8-27B") model = PeftModel.from_pretrained(base_model, "MaestroYan/JEMM") - Notebooks
- Google Colab
- Kaggle
JEMM
Like Jev, but multimodal and open-weight.
JEMM makes the same kind of fast, structured decision as Jev: given a state and a question, it picks one candidate and returns a probability for every candidate. It can also look at a screenshot, and it runs on your own GPU.
| JEMM | Jev | |
|---|---|---|
| Input | Text, or text + screenshot | Text only |
| Weights | Open, self-hosted | Hosted API |
| Question types | choice, noul, score | choice, noul, score |
JEMM stands for Judgment Engine for MultiModal decisions. It is a LoRA adapter for Qwen/Qwen3.8-27B. Not affiliated with TypeSafe AI.
Results
Usage
import json, torch
from huggingface_hub import hf_hub_download
from peft import PeftModel
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
base, adapter = "Qwen/Qwen3.8-27B", "MaestroYan/JEMM"
processor = AutoProcessor.from_pretrained(base)
model = Qwen3_5ForConditionalGeneration.from_pretrained(base, dtype=torch.bfloat16, device_map="cuda")
model = PeftModel.from_pretrained(model, adapter).eval()
config = json.load(open(hf_hub_download(adapter, "decision_config.json")))
labels = [chr(65 + i) for i in range(26)] + list("012345")
system = "Choose the best available candidate for the question using only the supplied state. Return exactly one candidate label."
def decide(state, question, candidates, image=None):
flat = lambda s: " ".join(s.split())
options = "\n".join(f"{labels[i]}) {flat(c)}" for i, c in enumerate(candidates))
text = f"State:\n{state.strip()}\n\nQuestion: {flat(question)}\n\nCandidates:\n{options}\n\nAnswer with exactly one candidate label."
content = ([{"type": "image"}] if image else []) + [{"type": "text", "text": text}]
messages = [{"role": "system", "content": system}, {"role": "user", "content": content}]
prompt = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = processor(text=[prompt], images=[image] if image else None, return_tensors="pt").to(model.device, dtype=torch.bfloat16)
with torch.inference_mode():
logits = model(**inputs, logits_to_keep=1).logits[0, -1]
ids = [processor.tokenizer.encode(x, add_special_tokens=False)[0] for x in labels[:len(candidates)]]
probs = torch.softmax(logits[ids].float() / config["mm_temperature" if image else "temperature"], -1)
return dict(zip(candidates, probs.tolist()))
decide("User: Will it rain in Paris tomorrow?", "Which tool should handle the request?",
["get_weather: weather forecast for a city", "search_flights: flight search", "No tool applies"])
2 to 32 candidates per question. Screenshots are PIL images; training used 1280x800. Treat a top probability below threshold in decision_config.json as undecided. HTTP server: JEMM on GitHub.
Training data
Mind2Web, xLAM function-calling-60k (APIGen), xlam-irrelevance, When2Call, Banking77, MASSIVE, Aegis 2.0 (CC-BY-4.0); CLINC150 (CC-BY-3.0); BFCL, glaive-function-calling-v2, ToolACE, hermes-function-calling-v1 (Apache-2.0); Multimodal-Mind2Web (OpenRAIL). No Jev outputs were used.
- Downloads last month
- 13
Model tree for MaestroYan/JEMM
Base model
Qwen/Qwen3.8-27B

