Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

jeff-adapter-triage

Support ticket triage. Routes a support ticket or email to a team and scores its urgency, sentiment and need for a person.

A LoRA adapter for jeff-base v1.3, a small open decision model (a fine-tune of Qwen3.5-0.8B). You send a situation (the state) and questions with named options; Jeff returns a calibrated probability for every option from one forward pass, with no generated text to parse. One Jeff server loads the base once and any number of adapters beside it; each request picks an adapter by name ("model": "triage").

Adapter page, with the full data card: jeffhub.ai/adapters/triage.

Results

On this adapter's held-out test set, never trained on, scored three ways on the same rows: the untrained model Jeff is built from, the Jeff v1.3 base alone, and the base with this adapter. Questions have 2 to 36 options. As of 2026-10-05. All adapters

Test set Test rows Qwen3.5-0.8B untrained Jeff base v1.3 alone Jeff base v1.3 + adapter
test 7,256 44.1% · 0.098 47.6% · 0.059 91.5% · 0.015

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

With llama.cpp (GGUF)

The same test, through llama.cpp: the base GGUF (mstrasser/jeff-base-gguf) plus this adapter's LoRA GGUF (mstrasser/jeff-adapter-triage-gguf), with the temperature refitted for each format. Running Jeff with llama.cpp

Test set Full precision Q8_0 Q4_K_M
test 91.5% · 0.015 91.5% · 0.015 91.2% · 0.014

Not measured yet for v1.3: calibration charts, the commonest confusions and accuracy per answer. External benchmarks: typed-decisions, customer_service; Banking77, test split.

Source of these numbers: results/sources/v1.3/retrained-adapters.table.json in the JeffHub repository, also collected in jeffhub-v1.3.json.

When to use it

  • You receive tickets, emails or chat messages and need a team, an urgency and a "does a person need to see this?" flag for each.
  • Your teams are described in a sentence each; the adapter reads the descriptions, so your team list can be your own (up to 40 teams in training).
  • You want all four answers from one request.

When not to use it

  • You need the reply written for you. Jeff chooses between options; it does not generate text.
  • The decision depends on facts outside the message, such as order history or account status. Put them in the state, or decide in code.
  • Messages mostly arrive in languages other than English. Training had some European-language messages, but most are English.

How to use it

The adapter runs with Jeff's server on the jeff-base v1.3 base. Adapter serving arrives with the next Jeff release; until then, these commands need the feat/lora branch of firelex/jeff.

git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora          # add --extra cuda on NVIDIA GPUs, --extra mac on Apple silicon
uv run --no-default-groups hf download mstrasser/jeff-base --revision v1.3 --local-dir checkpoints/jeff-base
uv run --no-default-groups hf download mstrasser/jeff-adapter-triage --revision v1.3 --local-dir adapters/triage
JEFF_CHECKPOINT=checkpoints/jeff-base JEFF_ADAPTERS=adapters/ PORT=8765 \
  uv run --no-default-groups jeff-serve          # on a Mac, add JEFF_BACKEND=mlx

Every folder in adapters/ is served under its folder name; add or replace adapters while the server runs with curl -X POST http://localhost:8765/v1/adapters/reload. Each adapter records the exact base it was trained on, and the server refuses an adapter trained on a different one, so this adapter loads only on jeff-base v1.3 (a v1.2 adapter does not load on v1.3). For llama.cpp, use mstrasser/jeff-adapter-triage-gguf.

Request format

State (the situation), in this order:

Key Changes per request What it holds
company no One sentence about the organisation the message is sent to.
channel no Where the message came from, for example email, web form, chat, app review, social media reply or phone transcript.
message yes The ticket or email as received, including subject line, quoted thread and signature.

Questions:

  • route (choice): Which team should handle the message. If it covers several issues, the team for the most important one. Options: other first (Not for any of these teams), then your teams as k1, k2, … with one line each on what the team handles.
  • urgency (score): How urgently the message needs a response. Options: Five levels, lowest first, from No action needed to Handle immediately.
  • sentiment (score): How the customer feels. Options: Five levels, from Very negative, angry or distressed to Very positive.
  • needs_human (noul): Whether a person must act on it now rather than an automatic reply (escalating complaints, legal or safety issues, requests an automatic system cannot resolve).

Rules:

  • Ask all four questions in one request; they are answered together.
  • Keep other as the first route option, word for word, and list the teams after it in a fixed order so the unchanging part of the request can be prepared in advance.
  • Use the instructions below word for word; the adapter was trained mostly on them.

General rules for every request: the request format guide.

Example

The request below is also in this repository as example.json.

{
  "model": "triage",
  "state": {
    "company": "A mid-sized online furniture retailer shipping across Europe.",
    "channel": "email",
    "message": "Subject: Crushed wardrobe\n\nHi, the wardrobe I ordered (order 55120) arrived today and the box was crushed. Two doors are split. I want a replacement or my money back. This is the second time.\n\nAnna"
  },
  "questions": {
    "route": {
      "type": "choice",
      "instructions": "Which team should handle this message? If it covers several issues, choose the team for the most important one.",
      "criteria": {
        "other": "Not for any of these teams",
        "k1": "Deliveries: late, lost or damaged deliveries, and delivery bookings",
        "k2": "Refunds and payments: refunds, charges and invoices",
        "k3": "Assembly service: booking and complaints about assembly",
        "k4": "Account access: logins, passwords and account details"
      }
    },
    "urgency": {
      "type": "score",
      "instructions": "How urgently does this need a response?",
      "criteria": [
        "No action needed",
        "Can wait a few days",
        "Handle within a day",
        "Handle within hours",
        "Handle immediately"
      ]
    },
    "sentiment": {
      "type": "score",
      "instructions": "How does the customer feel?",
      "criteria": [
        "Very negative, angry or distressed",
        "Negative",
        "Neutral",
        "Positive",
        "Very positive"
      ]
    },
    "needs_human": {
      "type": "noul",
      "instructions": "Does this need a person to act on it now, rather than an automatic reply? Answer yes for complaints that could escalate, legal or safety issues, or requests an automatic system cannot resolve."
    }
  }
}
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d @adapters/triage/example.json

The answer holds a probability for each option of each question. A recorded response from the v1.3 adapter is not published yet.

Files

  • adapter_model.safetensors, adapter_config.json: the LoRA weights (PEFT format);
  • readout.safetensors: the adapter's own readout over the answer codes;
  • decision_config.json: answer codes, temperature, prompt layout and the checksum of the base it was trained on;
  • test.jsonl: the held-out test set the results below were measured on;
  • calibration.jsonl: the calibration rows the adapter's temperature was fitted on;
  • example.json: the example request above.

adapter_config.json and decision_config.json name the base as mstrasser/jeff-base, revision v1.3; the server checks the base by the checksum of its weights.

Training

Base mstrasser/jeff-base, revision v1.3 (a fine-tune of Qwen3.5-0.8B)
Prompt layout live-last: the fixed part of the request first, the changing state field last
Run 0.8b-triage-20261003-0149, final checkpoint
Adapter files 41.5 MB (adapter_model.safetensors and readout.safetensors)
LoRA GGUF for llama.cpp mstrasser/jeff-adapter-triage-gguf
  • 1.3.0 (2026-10-03): Trained on Jeff v1.3 with the live-last prompt layout (LoRA rank 16, one epoch, about 10% of the base model's own training data mixed in).

Data card

Report attached. The shortcut report and data card are included and pass the JeffHub checks; the numbers are the maintainers’ own. What the levels mean

  • Test set: included in this repository as test.jsonl, so anyone can check the numbers
  • Calibration rows: included in this repository as calibration.jsonl, the rows its threshold is chosen on
  • QA report, sanitised: the data-quality checks run before training

How the test set was held out. Messages to the 10% of the roughly 300 generated organisations that were never trained on, scored on all four questions (route, urgency, sentiment and whether a person is needed).

Training data. Training data not published.

Which models made the data, counted on the 65,752 training rows:

What it did Model Where it ran Training rows
Wrote the message teacher.writer Qwen3.8-Max hosted (Alibaba Cloud DashScope) 64,740
Wrote the message teacher.writer Qwen3.8-Flash-Next local (own hardware) 1,012
Edited the message to remove giveaway cues decorrelated_by Qwen3.8-Flash hosted (Alibaba Cloud DashScope) 33,764
Rated urgency and sentiment score_raters Qwen3.8-Max hosted (Alibaba Cloud DashScope) 65,752
Rated urgency and sentiment score_raters DeepSeek-V4-Flash hosted (DeepSeek) 65,752
Checked the labels teacher.checker Qwen3.8-Flash hosted (Alibaba Cloud DashScope) 33,968
Checked the labels teacher.checker Qwen3.8-Flash-Next local (own hardware) 31,784

Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total. Urgency and sentiment targets are the mean of the generation plan and the two model ratings.

The attached QA report was written for the data of the previous release; the v1.3 data fixes the notes it left open. The QA report re-run on the v1.3 data is still to be attached.

The terms of the hosted model providers are being checked for training and publication use.

Training mixed in a replay sample of the Jeff base model's own training data: 6,575 rows, about 10% on top of the adapter's 65,752 (inherited from the v1.2 recipe as a precaution; its effect has not been measured).

Data and licence

Adapter licence: Apache-2.0.

Qwen3.5-0.8B notice: these weights were modified from Qwen3.5-0.8B by the Jeff project: jeff-base is a fine-tune of Qwen3.5-0.8B, and this adapter was trained on top of it. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in LICENSE.

It was trained on:

  • Generated organisations and messages. Licence: Released with the adapter under Apache-2.0 (made for this adapter) · Made by Qwen3.8-Max (hosted), with some rows by Qwen3.8-Flash-Next (local); see the data card

    About 300 organisations and 40,000 messages written by language models and checked by a second blind pass (which models, and for how many rows, is in the data card). Urgency and sentiment targets are the mean of three ratings.

Limitations

  • Tied to jeff-base v1.3. It will not load on any other base or version; the server checks the base weights' checksum.
  • Jeff chooses between the options you give it. It does not write text or reason in several steps.
  • Calibration was fitted on this adapter's own calibration rows. On very different data, check it again.
  • Everything listed under When not to use it above.

Links

Jeff is an independent project. It uses the same request format as Jev but is not affiliated with or endorsed by TypeSafe, the makers of Jev.

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mstrasser/jeff-adapter-triage

Adapter
(30)
this model