Instructions to use jgeuter/Jeff-1.0-Large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jgeuter/Jeff-1.0-Large with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jgeuter/Jeff-1.0-Large") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("jgeuter/Jeff-1.0-Large") model = AutoModelForCausalLM.from_pretrained("jgeuter/Jeff-1.0-Large", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jgeuter/Jeff-1.0-Large with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jgeuter/Jeff-1.0-Large" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jgeuter/Jeff-1.0-Large", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jgeuter/Jeff-1.0-Large
- SGLang
How to use jgeuter/Jeff-1.0-Large with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jgeuter/Jeff-1.0-Large" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jgeuter/Jeff-1.0-Large", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jgeuter/Jeff-1.0-Large" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jgeuter/Jeff-1.0-Large", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use jgeuter/Jeff-1.0-Large with Docker Model Runner:
docker model run hf.co/jgeuter/Jeff-1.0-Large
Jeff-1.0-Large
Jeff-1.0-Large is a Jev-style decision model: a state, a typed question and its options go in, and one forward pass returns a calibrated probability for each option. Nothing is generated. It is Gemma 4 31B (instruction-tuned) with a LoRA adapter merged into the weights and a small decision head, trained with plain cross-entropy for 400 steps.
- Serving code (Docker, TypeSafe "System One" request format as used by JevBench): https://github.com/j-geuter/jeff-serve
Use
hf download jgeuter/Jeff-1.0-Large --local-dir /models/Jeff-1.0-Large
git clone https://github.com/j-geuter/jeff-serve && cd jeff-serve
docker build -t jeff . && docker run --gpus '"device=0"' -v /models/Jeff-1.0-Large:/model:ro -p 8013:8013 jeff
One GPU with at least 80 GB (bf16 weights are 62 GB). The model needs to be served through the serving code (see above).
What is in this repository
| File | Content |
|---|---|
model-*.safetensors, config.json, tokenizer files |
Gemma 4 31B text model (bf16) with merged LoRA adapter |
head.safetensors, head_config.json |
decision head |
serving_rule.json |
per-type temperatures (choice 1.52, yes/no 1.31, score 1.31), yes/no commit |
jeff_config.json |
base model and revision, training run, prompt format, input limit |
SHA256SUMS |
checksums of all files |
How it decides
Each question is rendered as a multiple-choice prompt (SemIf's template: a short system message and a JSON object with the state, the question, and the options labeled A, B, C, ...; Gemma's chat template, thinking disabled). The model reads it once. The head maps the last hidden state to one logit per option. The serving rule divides the logits by one temperature per question type (fitted on validation data), and moves yes/no answers with $0.2 < P(\text{yes}) < 0.8$ to just outside the nearer edge (0.801 or 0.199).
Question types: choice (2 to 16 options), yes/no, score (2 to 16 ordered levels). Inputs up to 8,192 tokens.
Training
LoRA rank 32 (alpha = 64) on all attention and MLP projections, trained jointly with the head; AdamW, learning rates 2e-5 (LoRA) and 1e-4 (head), 250 warmup steps, constant to step 300, cosine decay to zero at step 400; 65,536 tokens per step (about 76,000 questions seen); cross-entropy against soft targets where they exist, else the label; option order shuffled. 78 minutes on 4 H200 GPUs. No reinforcement learning, reward model or auxiliary loss.
Training pool (417,825 questions):
| Part | Rows | Labels | Licence of the source |
|---|---|---|---|
| decider's public training mixture (79 tasks rebuilt from public datasets with decider's code) | 300,042 | dataset labels | mixed; includes sources with non-commercial terms (e.g. ToxicChat and customer-support-tickets: CC BY-NC 4.0; MS MARCO and RACE: research-only terms) |
| Civil Comments | 30,000 | annotator distributions | CC0 1.0 |
| Measuring Hate Speech | 30,000 | annotator distributions | CC BY 4.0 |
| STS-B | 4,760 | similarity distributions | see the STS benchmark |
| ChaosNLI | 2,886 | 100 votes per item | CC BY-NC 4.0 |
| synthetic business documents (ours) | 30,641 | teacher probabilities | generated with Qwen models (Apache-2.0) |
| decider's custom questions, relabeled | 19,496 | teacher probabilities | Apache-2.0 (decider) |
Teacher labels come from Qwen3.5-122B-A10B and Qwen3.6-27B (two thinking samples each, averaged).
Disclosure. No JevBench item (public or sealed) was used for training, model selection or calibration. The 231 public JevBench items were used only as an evaluation check. The recipe and the seed were chosen on jb16-dev, our own validation set of 2,389 decisions.
Evaluation
| Jeff-1.0-Large | Gemma 4 31B, zero-shot letter readout | Jev 1.13.0 (API) | |
|---|---|---|---|
| jb16-dev Capability, (Intelligence + Calibration) / 2 | 83.8 | 78.2 | 78.7 |
| Seven public held-out decision datasets, macro accuracy (%) | 89.0 | 88.4 | 88.0 |
| same, mean ECE (lower is better) | 0.047 | 0.077 | 0.066 |
jb16-dev is scored without JevBench's yes/no abstention band and without the yes/no commit. With the band and the commit, as served here, the model scores 83.7 (Intelligence 74.0, Calibration 93.4). The seven public sets are TREC, CB, WANLI (256 items), SemIf's authored and perturbation sets, TypeSafe's example set and decider's held-out tasks (9,212 items); none of them was in the training pool. See the paper for intervals.
Limitations
- At most 16 options per question; inputs above 8,192 tokens are refused.
- Mostly English training data (jb16-dev is 79% English).
- On knowledge-heavy multiple choice (MMLU-Pro, MMLU, WinoGrande) Jev is still ahead by 4 to 5 points.
- Probabilities are calibrated on our validation data; recalibrate for a very different domain.
Licence
The weights are released under CC BY-NC 4.0 (non-commercial). They are derived from google/gemma-4-31B-it (Apache License 2.0), and the training data includes sources with non-commercial terms (see the table above). The serving code is Apache-2.0.
- Downloads last month
- 69