autotrust/JEV-27B-VL

JEV-27B that can see: the same System 1 decisions and System 2, now on images

autotrust/JEV-27B-VL is autotrust/JEV-27B with vision. It pairs the multimodal Qwen3.8-27B (unmodified, including its vision encoder) with JEV-27B's System 1 adapter and decision head. One vLLM engine serves both systems, for text and for images:

what it does output
System 1 (jev-decision) typed decisions: yes/no · pick one of 2-16 options · rate 0-5, over text and images a calibrated probability for every option, in one forward pass
System 2 (autotrust/JEV-27B-VL) the unmodified Qwen3.8-27B, optionally thinking step by step, with image input text / reasoning

JEV-27B's language weights are bit-identical to Qwen3.8-27B's, so on text this model gives the same System 1 decisions as JEV-27B. On 24 reference decisions the largest probability difference was 0.024 (bf16 noise), with the same top option 24/24.

Highlight: #1 on the VL-RewardBench leaderboard (multimodal judge)

Given an image, a question and two answers, System 1 says which answer is better, zero-shot and in one forward pass per order. On VL-RewardBench (CVPR 2025, 1,247 human-verified pairs) it reaches 78.3% overall accuracy, the highest on the official leaderboard (26 models, updated May 2025):

judge general hallucination reasoning overall macro
JEV-27B-VL System 1 58.0 83.2 78.2 78.3 73.1
Skywork-VL-Reward-7B 65.6 80.2 61.3 73.3 69.0
Gemini 2.0 Flash 50.8 72.6 70.1 68.8 64.5
Gemini 1.5 Pro 50.8 72.5 64.2 67.2 62.5
GPT-4o 49.1 67.6 70.5 65.8 62.4
Claude 3.5 Sonnet 43.4 55.0 62.3 55.3 53.6

2,494 decisions in 163 seconds on one GPU. Code: JEV-27B-DEMO / 09-multimodal-judge.

multimodal judge

Highlight: zero-shot image recommendation

System 1's decision head was trained on text only. Given images, it still works zero-shot. On MicroLens-100k (real users of a short-video app, raw cover images), System 1 looked at the covers of the last 5 videos a user watched and ranked 20 candidate covers: the video the user actually watched next, hidden among 19 videos other users watched within ±3 days. 200 users.

zero-shot image recommendation

method uses interaction logs? AUC HR@5 NDCG@10
random order no 0.489 0.205 0.219
title similarity (TF-IDF) no 0.602 0.385 0.367
JEV-27B-VL System 1, titles only no, zero-shot 0.649 0.455 0.399
JEV-27B-VL System 1, covers only no, zero-shot 0.727 0.590 0.498
item-based collaborative filtering yes, 59,045 other users 0.728 0.490 0.503
  • Looking beats reading: covers alone +0.078 AUC over titles (95% interval +0.031 to +0.126).
  • Zero-shot matches collaborative filtering learned from 59,045 users' watch histories (AUC difference 0.000, interval −0.041 to +0.040), with a higher top-5 hit rate. Brand-new videos can be ranked from their cover alone.

Code, data pipeline and a live web demo: JEV-27B-DEMO / 07-image-recommendation.

Benchmarks

JEV-27B-VL gives the same text decisions as autotrust/JEV-27B: on 1,000 answer-checking decisions the mean probability difference is 0.010 and 99.8% fall on the same side of 0.5. The text results below were measured with JEV-27B; the image result was measured with JEV-27B-VL.

Six public decision benchmarks

Scores in %, higher is better. JEV-27B and TypeSafe Jev 1.13 (hosted API) were run in full by AutoTrust (27 September 2026); the other rows are as reported in the NeoHorse-Jev-4B evaluation.

Model JevBench Kev OpenJev text Nimble VitaminC MASSIVE-en Six-group mean
JEV-27B / JEV-27B-VL 88.70 83.75 73.89 92.91 77.46 87.71 84.07
TypeSafe Jev 1.13, hosted API 87.18 85.52 72.96 91.84 78.46 87.14 83.85
NeoHorse-Jev-4B 75.73 81.92 58.74 87.23 77.13 85.43 77.70
Open-Jev-9B 77.13 77.87 65.39 80.50 68.28 84.86 75.67
Kev-4B 73.71 81.47 54.75 73.40 76.46 85.71 74.25
Laya English 55.82 61.30 40.07 45.04 78.63 68.57 58.24

System 1 fidelity and calibration

Held-out test_set_30k of jev-distill-corpus-v3.

metric JEV-27B / JEV-27B-VL
KL divergence from TypeSafe Jev 1.13's distributions (Jev-labelled rows, 0 = identical) ≈ 0.017
yes/no AUROC 0.995
choice top-1 agreement with Jev (rows with a clear top option) 95.8%
rating error, 0-5 scale (MAE of the expected rating) 0.098
expected calibration error 0.0009
KL to ground-truth labels of unseen task families 0.104

Independent benchmark: decision-models-under-pressure

Human gold labels (CLINC-150, MTOP, GoEmotions, DBpedia), 800 items.

options 2 4 8 16
TypeSafe Jev 1.13 (published) 0.890 0.801 0.782 0.769
JEV-27B / JEV-27B-VL 0.876 0.784 0.767 0.740

Applied tasks (from JEV-27B-DEMO)

task result comparison
Image recommendation, zero-shot (MicroLens, covers only) · JEV-27B-VL AUC 0.727 collaborative filtering from 59,045 users' logs 0.728 · titles only 0.649
Multimodal judge (VL-RewardBench, 1,247 pairs) · JEV-27B-VL 78.3% accuracy, #1 on the leaderboard Skywork-VL-Reward-7B 73.3 · GPT-4o 65.8 · Claude 3.5 Sonnet 55.3
Biomedical research questions (PubMedQA test, 500) · JEV-27B-VL 77.8% accuracy, zero-shot human experts 78.0% · GPT-4 zero-shot 75.2% · BioBERT 68.1%
Response judge (RewardBench, 2,985 pairs) 89.9 Gemini 1.5 Pro 88.2 · GPT-4o 86.7 · Claude 3.5 Sonnet 84.2 as judges
Hallucination guard (TriviaQA, answer the trusted half) 96.4% accuracy answering everything 71.2% · model's own stated confidence 87.2%
News recommendation, zero-shot (MIND) AUC 0.642 best zero-shot baseline 0.606 · LightGBM trained on MIND 0.616
Search re-ranking (TREC-COVID, nDCG@10) 0.858 bge-reranker-v2-m3 0.793 · BM25 0.623
System 1 → System 2 (escalate below 0.70 confidence) accuracy 0.892, 70% answered in 0.11 s System 1 only 0.792 · thinking on everything 0.917

System 2

benchmark result
HumanEval pass@1 (greedy) 78.0%, identical to Qwen3.8-27B

Quick start

hf download autotrust/JEV-27B-VL --local-dir JEV-27B-VL
bash JEV-27B-VL/serve.sh          # vLLM on :8000; one GPU with 80 GB or more

serve.sh runs:

vllm serve JEV-27B-VL --served-model-name autotrust/JEV-27B-VL \
  --enable-lora --max-lora-rank 32 --lora-modules jev-decision=JEV-27B-VL/adapter_vllm \
  --logprobs-mode processed_logprobs --max-model-len 16384 --enable-prefix-caching --mamba-cache-mode align \
  --limit-mm-per-prompt '{"image": 8}' --max-num-seqs 8 --trust-request-chat-template

System 1 on images

The decision prompt is sent verbatim, with images at their place in the state. The client asks only for the verbalizer tokens and applies the bundled bias and temperature:

import base64, json, math, requests

B = "JEV-27B-VL"
DH = json.load(open(f"{B}/adapter_vllm/decision_head.json"))
T = json.load(open(f"{B}/calibration.json"))["per_kind"]
RAW = ("{%- for m in messages -%}{%- for c in m['content'] -%}{%- if c['type'] == 'text' -%}{{ c['text'] }}"
       "{%- else -%}<|vision_start|><|image_pad|><|vision_end|>{%- endif -%}{%- endfor -%}{%- endfor -%}")

def image(path):
    return {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64," + base64.b64encode(open(path, "rb").read()).decode()}}

def decide_mm(kind, state_parts, question, options=None):
    """state_parts: list of str and image paths ({'image': path})."""
    options = {"noul": ["false", "true"], "score": [str(i) for i in range(6)]}.get(kind, options)
    lines = options if kind != "choice" else [f"{'ABCDEFGHIJKLMNOP'[i]}) {o}" for i, o in enumerate(options)]
    content = [{"type": "text", "text": f"[kind] {kind}\n[state] "}]
    content += [image(p["image"]) if isinstance(p, dict) else {"type": "text", "text": p} for p in state_parts]
    content.append({"type": "text", "text": f"\n[question] {question}\n[options]\n" + "\n".join(lines) + "\n[decision]:"})
    s = DH["slots"]["ranges"][kind][0]
    ids = DH["verbalizer_ids"][s: s + len(options)]
    r = requests.post("http://localhost:8000/v1/chat/completions", json={
        "model": "jev-decision", "messages": [{"role": "user", "content": content}], "chat_template": RAW,
        "add_generation_prompt": False, "add_special_tokens": False, "max_tokens": 1, "temperature": 1.0,
        "logprobs": True, "top_logprobs": len(options), "allowed_token_ids": ids, "return_tokens_as_token_ids": True}).json()
    lp = {int(t["token"].split(":")[1]): t["logprob"] for t in r["choices"][0]["logprobs"]["content"][0]["top_logprobs"]}
    z = [(lp.get(t, -1e9) + DH["bias"][s + i]) / T[kind] for i, t in enumerate(ids)]
    e = [math.exp(x - max(z)) for x in z]
    return {o: x / sum(e) for o, x in zip(options, e)}

decide_mm("noul", [{"image": "photo.jpg"}], "Is this scenario one where: the image shows food or cooking?")
# e.g. {'false': 0.00, 'true': 1.00}

Text-only System 1 decisions work exactly as with autotrust/JEV-27B (/v1/completions with the same prompt format), or through decide_mm with text parts only; both give the same result.

System 2 on images

requests.post("http://localhost:8000/v1/chat/completions", json={
    "model": "autotrust/JEV-27B-VL",
    "messages": [{"role": "user", "content": [image("chart.png"), {"type": "text", "text": "What does this chart show?"}]}],
    "max_tokens": 1024, "chat_template_kwargs": {"enable_thinking": False}})

Serving notes

  • --max-num-seqs 8 is required. With more than 8 sequences in one batch, vLLM's LoRA path for this multimodal model class returns wrong System 1 probabilities (the text-only JEV-27B is not affected). With the cap, results match JEV-27B at any client concurrency; requests beyond 8 simply queue.
  • Throughput on one B200 with the cap: 24 short text decisions in about 1.4 s; 4,000 image decisions with 6 images each in 316 s (about 13 per second).
  • --trust-request-chat-template lets the client send the raw decision template shown above. System 2 uses the model's own chat template.
  • Images are resized by the Qwen3.8 processor. Downscaling large images first (for example to at most 448 px) keeps the number of vision tokens and the latency low.

Limitations

  • System 1's decision head was trained on text. Decisions over images are zero-shot: they are well ordered in the experiment above, but their calibration on image tasks has not been measured systematically.
  • The image-recommendation result is one dataset and 200 users.

License

Apache-2.0. This repository contains the weights of Qwen/Qwen3.8-27B (Apache-2.0, see LICENSE) unchanged, plus the JEV System 1 adapter and decision head from autotrust/JEV-27B.

Downloads last month
12
Safetensors
Model size
28B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for autotrust/JEV-27B-VL

Base model

Qwen/Qwen3.8-27B
Adapter
(137)
this model

Article mentioning autotrust/JEV-27B-VL