Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

jeff-adapter-ground

Passage re-ranking and answer grounding. Picks the passage that answers a question, and checks whether an answer is supported by its sources.

A LoRA adapter for jeff-base v1.3, a small open decision model (a fine-tune of Qwen3.5-0.8B). You send a situation (the state) and questions with named options; Jeff returns a calibrated probability for every option from one forward pass, with no generated text to parse. One Jeff server loads the base once and any number of adapters beside it; each request picks an adapter by name ("model": "ground").

Adapter page, with the full data card: jeffhub.ai/adapters/ground.

Results

On this adapter's held-out test set, never trained on, scored three ways on the same rows: the untrained model Jeff is built from, the Jeff v1.3 base alone, and the base with this adapter. Questions have 4 to 41 options. As of 2026-10-05. All adapters

Test set Test rows Qwen3.5-0.8B untrained Jeff base v1.3 alone Jeff base v1.3 + adapter
test 4,160 28.9% · 0.061 50.9% · 0.079 96.6% · 0.007

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

With llama.cpp (GGUF)

The same test, through llama.cpp: the base GGUF (mstrasser/jeff-base-gguf) plus this adapter's LoRA GGUF (mstrasser/jeff-adapter-ground-gguf), with the temperature refitted for each format. Running Jeff with llama.cpp

Test set Full precision Q8_0 Q4_K_M
test 96.6% · 0.007 96.6% · 0.009 96.4% · 0.010

Not measured yet for v1.3: calibration charts, the commonest confusions and accuracy per answer. External benchmarks: SQuAD 2.0 dev, re-ranking; HotpotQA dev (distractor), re-ranking; FEVER shared-task dev, grounding.

Source of these numbers: results/sources/v1.3/retrained-adapters.table.json in the JeffHub repository, also collected in jeffhub-v1.3.json.

When to use it

  • You run retrieval-augmented generation and want to pick the best of 5 to 40 retrieved passages, or learn that none answers the question.
  • You want to check a generated answer against its sources before showing it, and tell apart supported, partly supported, contradicted and unsupported answers.
  • You want both checks from one small model, fast enough to sit on every request.

When not to use it

  • You need the answer written or the unsupported claim pointed out. Jeff only chooses between options; it does not generate text.
  • Your passages are very long or very many. Training passages were 30 to 300 words, and every request stayed under 8,192 tokens.
  • You need a judgement from outside knowledge. The grounding check is trained to judge only from the sources given.
  • Your texts are mostly not in English. The training data is English.

How to use it

The adapter runs with Jeff's server, on the main branch of firelex/jeff, on the jeff-base v1.3 base.

git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora          # add --extra cuda on NVIDIA GPUs, --extra mac on Apple silicon
uv run --no-default-groups hf download mstrasser/jeff-base --revision v1.3 --local-dir checkpoints/jeff-base
uv run --no-default-groups hf download mstrasser/jeff-adapter-ground --revision v1.3 --local-dir adapters/ground
JEFF_CHECKPOINT=checkpoints/jeff-base JEFF_ADAPTERS=adapters/ PORT=8765 \
  uv run --no-default-groups jeff-serve          # on a Mac, add JEFF_BACKEND=mlx

Every folder in adapters/ is served under its folder name; add or replace adapters while the server runs with curl -X POST http://localhost:8765/v1/adapters/reload. Each adapter records the exact base it was trained on, and the server refuses an adapter trained on a different one, so this adapter loads only on jeff-base v1.3 (a v1.2 adapter does not load on v1.3). For llama.cpp, use mstrasser/jeff-adapter-ground-gguf.

Request format

State (the situation), in this order:

Key Changes per request What it holds
sources no The source passages the answer was written from, one to six in training, each with a short label such as "[1] Title" or "Source 1".
answer yes The answer to check.

Questions:

  • grounding (choice): Whether the answer is supported by the sources, judged only from the sources. Options: Four fixed options: supported, partly_supported, contradicted and unsupported, each with the one-line meaning shown in the example.

Rules:

  • Use the instructions below word for word; the adapter was trained mostly on them. Grounding instructions, "Is the answer supported by the sources? Judge only from the sources, not from outside knowledge."
  • Keep the four grounding option keys and their texts exactly as in the example. Their order does not matter; it was shuffled in training.
  • The adapter also re-ranks passages, with a second request shape. State: one key, question, holding the user's question. Options: none (text "None of these passages answers the question") plus the passages as p1, p2, … with the passage text as the option text, optionally starting with its title in square brackets. Instructions, word for word: "Which passage answers the question? If none of them does, choose none."
  • Re-ranking was trained with 5 to 40 passages of 30 to 300 words each. The passages change on every request, so nothing can be prepared in advance for them.
  • Send one question per request; the two shapes have different state keys and are separate requests.

General rules for every request: the request format guide.

Example

The request below is also in this repository as example.json.

{
  "model": "ground",
  "state": {
    "sources": "[1] Harbour Bridge\nThe Harbour Bridge opened to traffic in March 1932 after eight years of work. It carries eight lanes of road traffic and two railway lines.\n\n[2] Harbour ferries\nFerries have crossed the harbour since the 1840s. The busiest route runs from the Quay to Manly.",
    "answer": "The Harbour Bridge opened in 1934 and carries two railway lines."
  },
  "questions": {
    "grounding": {
      "type": "choice",
      "instructions": "Is the answer supported by the sources? Judge only from the sources, not from outside knowledge.",
      "criteria": {
        "supported": "Everything the answer claims is stated in or follows directly from the sources",
        "partly_supported": "Some claims are supported, but at least one is not in the sources",
        "contradicted": "The sources state something that conflicts with the answer",
        "unsupported": "The answer's main claim is not in the sources at all"
      }
    }
  }
}
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d @adapters/ground/example.json

The answer holds a probability for each option of each question. A recorded response from the v1.3 adapter is not published yet.

Files

  • adapter_model.safetensors, adapter_config.json: the LoRA weights (PEFT format);
  • readout.safetensors: the adapter's own readout over the answer codes;
  • decision_config.json: answer codes, temperature, prompt layout and the checksum of the base it was trained on;
  • test.jsonl: the held-out test set the results below were measured on;
  • calibration.jsonl: the calibration rows the adapter's temperature was fitted on;
  • example.json: the example request above.

adapter_config.json and decision_config.json name the base as mstrasser/jeff-base, revision v1.3; the server checks the base by the checksum of its weights.

Training

Base mstrasser/jeff-base, revision v1.3 (a fine-tune of Qwen3.5-0.8B)
Prompt layout live-last: the fixed part of the request first, the changing state field last
Training code The git_commit recorded in decision_config.json is the training machine's copy and was not published. It builds exactly the same prompt as main of firelex/jeff (from commit 6d0d7da) for a text state and for an object with at least one field; the format is in docs/v1.3-request-format.md
Run 0.8b-ground-20261003-0149, final checkpoint
Adapter files 41.5 MB (adapter_model.safetensors and readout.safetensors)
LoRA GGUF for llama.cpp mstrasser/jeff-adapter-ground-gguf
  • 1.3.0 (2026-10-03): Trained on Jeff v1.3 with the live-last prompt layout (LoRA rank 16, one epoch, about 10% of the base model's own training data mixed in).

Data card

Report attached. The shortcut report and data card are included and pass the JeffHub checks; the numbers are the maintainers’ own. What the levels mean

  • Test set: included in this repository as test.jsonl, so anyone can check the numbers
  • Calibration rows: included in this repository as calibration.jsonl, the rows its threshold is chosen on
  • QA report, sanitised: the data-quality checks run before training

How the test set was held out. Requests from the 10% of source documents and articles that were never trained on, for both re-ranking and grounding.

Training data. Training data not published.

Which models made the data, counted on the 40,960 training rows:

What it did Model Where it ran Training rows
Wrote the text text_teacher Qwen3.8-Max hosted (Alibaba Cloud DashScope) 12,612
Wrote the text (named in the licence field) license Qwen3.8-Flash-Next local (own hardware) 9,873
Wrote the text text_teacher Qwen3.8-Flash-Next local (own hardware) 7,095
Checked the label check_teacher Qwen3.8-Flash hosted (Alibaba Cloud DashScope) 40,550
Gave a second opinion on the label second_opinion DeepSeek-V4-Flash hosted (DeepSeek) 7,483
Checked the label check_teacher Qwen3.8-Flash-Next local (own hardware) 410

Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total. Rows built from SQuAD 2.0, HotpotQA and FEVER keep their public text; the jobs above show which models wrote the generated text and checked the rows.

The attached QA report was written for the data of the previous release; the v1.3 data fixes the notes it left open. The QA report re-run on the v1.3 data is still to be attached.

The terms of the hosted model providers are being checked for training and publication use.

The grounding training rows were chosen so that negation words, digits and source length are spread evenly across the labels, so none of them gives the answer away.

Training mixed in a replay sample of the Jeff base model's own training data: 4,096 rows, about 10% on top of the adapter's 40,960 (inherited from the v1.2 recipe as a precaution; its effect has not been measured).

Data and licence

Adapter licence: Apache-2.0.

Qwen3.5-0.8B notice: these weights were modified from Qwen3.5-0.8B by the Jeff project: jeff-base is a fine-tune of Qwen3.5-0.8B, and this adapter was trained on top of it. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in LICENSE.

It was trained on:

  • SQuAD 2.0 (training split). Licence: CC-BY-SA-4.0 (open, but shared or changed data must keep the same terms) · Not made by a model

    Revision 3ffb306f725f. Used for re-ranking and, with generated answers, for grounding.

  • HotpotQA, distractor setting (training split). Licence: CC-BY-SA-4.0 (open, but shared or changed data must keep the same terms) · Not made by a model

    Revision 1908d6afbbea.

  • FEVER (training split). Licence: CC-BY-SA-3.0 (open, but shared or changed data must keep the same terms) · Not made by a model

    Official release. Annotations include Wikipedia material under the Wikipedia licence terms. Used for grounding.

  • Generated documents, answers and near-miss passages. Licence: Released with the adapter under Apache-2.0 (made for this adapter) · Made by Qwen3.8-Max (hosted) and Qwen3.8-Flash-Next (local); see the data card

    Documents in 12 domains (such as company policies, product manuals and API documentation) with questions and answers, plus answers and near-miss passages for SQuAD and HotpotQA questions, written by language models (which ones, and for how many rows, is in the data card). Every label confirmed by a second blind pass.

Limitations

  • Tied to jeff-base v1.3. It will not load on any other base or version; the server checks the base weights' checksum.
  • Jeff chooses between the options you give it. It does not write text or reason in several steps.
  • Calibration was fitted on this adapter's own calibration rows. On very different data, check it again.
  • Everything listed under When not to use it above.

Links

Jeff is an independent project. It uses the same request format as Jev but is not affiliated with or endorsed by TypeSafe, the makers of Jev.

Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mstrasser/jeff-adapter-ground

Adapter
(30)
this model

Datasets used to train mstrasser/jeff-adapter-ground