Jeff-1.0-Large

Jeff-1.0-Large is a Jev-style decision model: a state, a typed question and its options go in, and one forward pass returns a calibrated probability for each option. Nothing is generated. It is Gemma 4 31B (instruction-tuned) with a LoRA adapter merged into the weights and a small decision head, trained with plain cross-entropy for 400 steps.

Use

hf download jgeuter/Jeff-1.0-Large --local-dir /models/Jeff-1.0-Large
git clone https://github.com/j-geuter/jeff-serve && cd jeff-serve
docker build -t jeff . && docker run --gpus '"device=0"' -v /models/Jeff-1.0-Large:/model:ro -p 8013:8013 jeff

One GPU with at least 80 GB (bf16 weights are 62 GB). The model needs to be served through the serving code (see above).

What is in this repository

File Content
model-*.safetensors, config.json, tokenizer files Gemma 4 31B text model (bf16) with merged LoRA adapter
head.safetensors, head_config.json decision head
serving_rule.json per-type temperatures (choice 1.52, yes/no 1.31, score 1.31), yes/no commit
jeff_config.json base model and revision, training run, prompt format, input limit
SHA256SUMS checksums of all files

How it decides

Each question is rendered as a multiple-choice prompt (SemIf's template: a short system message and a JSON object with the state, the question, and the options labeled A, B, C, ...; Gemma's chat template, thinking disabled). The model reads it once. The head maps the last hidden state to one logit per option. The serving rule divides the logits by one temperature per question type (fitted on validation data), and moves yes/no answers with $0.2 < P(\text{yes}) < 0.8$ to just outside the nearer edge (0.801 or 0.199).

Question types: choice (2 to 16 options), yes/no, score (2 to 16 ordered levels). Inputs up to 8,192 tokens.

Training

LoRA rank 32 (alpha = 64) on all attention and MLP projections, trained jointly with the head; AdamW, learning rates 2e-5 (LoRA) and 1e-4 (head), 250 warmup steps, constant to step 300, cosine decay to zero at step 400; 65,536 tokens per step (about 76,000 questions seen); cross-entropy against soft targets where they exist, else the label; option order shuffled. 78 minutes on 4 H200 GPUs. No reinforcement learning, reward model or auxiliary loss.

Training pool (417,825 questions):

Part Rows Labels Licence of the source
decider's public training mixture (79 tasks rebuilt from public datasets with decider's code) 300,042 dataset labels mixed; includes sources with non-commercial terms (e.g. ToxicChat and customer-support-tickets: CC BY-NC 4.0; MS MARCO and RACE: research-only terms)
Civil Comments 30,000 annotator distributions CC0 1.0
Measuring Hate Speech 30,000 annotator distributions CC BY 4.0
STS-B 4,760 similarity distributions see the STS benchmark
ChaosNLI 2,886 100 votes per item CC BY-NC 4.0
synthetic business documents (ours) 30,641 teacher probabilities generated with Qwen models (Apache-2.0)
decider's custom questions, relabeled 19,496 teacher probabilities Apache-2.0 (decider)

Teacher labels come from Qwen3.5-122B-A10B and Qwen3.6-27B (two thinking samples each, averaged).

Disclosure. No JevBench item (public or sealed) was used for training, model selection or calibration. The 231 public JevBench items were used only as an evaluation check. The recipe and the seed were chosen on jb16-dev, our own validation set of 2,389 decisions.

Evaluation

Jeff-1.0-Large Gemma 4 31B, zero-shot letter readout Jev 1.13.0 (API)
jb16-dev Capability, (Intelligence + Calibration) / 2 83.8 78.2 78.7
Seven public held-out decision datasets, macro accuracy (%) 89.0 88.4 88.0
same, mean ECE (lower is better) 0.047 0.077 0.066

jb16-dev is scored without JevBench's yes/no abstention band and without the yes/no commit. With the band and the commit, as served here, the model scores 83.7 (Intelligence 74.0, Calibration 93.4). The seven public sets are TREC, CB, WANLI (256 items), SemIf's authored and perturbation sets, TypeSafe's example set and decider's held-out tasks (9,212 items); none of them was in the training pool. See the paper for intervals.

Limitations

  • At most 16 options per question; inputs above 8,192 tokens are refused.
  • Mostly English training data (jb16-dev is 79% English).
  • On knowledge-heavy multiple choice (MMLU-Pro, MMLU, WinoGrande) Jev is still ahead by 4 to 5 points.
  • Probabilities are calibrated on our validation data; recalibrate for a very different domain.

Licence

The weights are released under CC BY-NC 4.0 (non-commercial). They are derived from google/gemma-4-31B-it (Apache License 2.0), and the training data includes sources with non-commercial terms (see the table above). The serving code is Apache-2.0.

Downloads last month
69
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jgeuter/Jeff-1.0-Large

Finetuned
(296)
this model
Quantizations
1 model