Instructions to use mstrasser/jeff-adapter-support-intents with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use mstrasser/jeff-adapter-support-intents with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
jeff-adapter-support-intents
Customer request intents. Names what a customer or assistant user is asking for, from a fixed list of requests such as cancel_order or track_refund.
A LoRA adapter for jeff-base v1.3, a small open decision model
(a fine-tune of Qwen3.5-0.8B). You send a situation (the state) and questions with named options; Jeff returns a
calibrated probability for every option from one forward pass, with no generated text to parse. One Jeff server loads
the base once and any number of adapters beside it; each request picks an adapter by name ("model": "support-intents").
Adapter page, with the full data card: jeffhub.ai/adapters/support-intents.
Results
On this adapter's held-out test set, never trained on, scored three ways on the same rows: the untrained model Jeff is built from, the Jeff v1.3 base alone, and the base with this adapter. Questions have 7 to 64 options. As of 2026-10-05. All adapters
| Test set | Test rows | Qwen3.5-0.8B untrained | Jeff base v1.3 alone | Jeff base v1.3 + adapter |
|---|---|---|---|---|
test |
5,577 | 33.9% · 0.164 | 24.2% · 0.095 | 96.3% · 0.003 |
Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).
By group
| Group | Rows | Qwen3.5-0.8B untrained | Jeff base v1.3 alone | Jeff base v1.3 + adapter |
|---|---|---|---|---|
| bitext | 2,361 | 43.4% | 10.0% | 99.7% |
| hwu64 seen by the base | 897 | 19.7% | 26.8% | 91.6% |
| hwu64 unseen | 1,624 | 17.6% | 26.9% | 93.2% |
| snips | 695 | 57.6% | 62.4% | 98.6% |
With llama.cpp (GGUF)
The same test, through llama.cpp: the base GGUF (mstrasser/jeff-base-gguf) plus this adapter's LoRA GGUF (mstrasser/jeff-adapter-support-intents-gguf), with the temperature refitted for each format. Running Jeff with llama.cpp
| Test set | Full precision | Q8_0 | Q4_K_M |
|---|---|---|---|
test |
96.3% · 0.003 | 96.3% · 0.006 | 96.2% · 0.005 |
Not measured yet for v1.3: calibration charts, the commonest confusions and accuracy per answer.
Source of these numbers: results/sources/v1.3/retrained-adapters.table.json in the JeffHub repository, also collected in jeffhub-v1.3.json.
When to use it
- You run an online shop chat or a voice assistant and want each message matched to one request from a fixed list.
- Your requests are close to the ones in training (27 online-shop requests, 64 home-assistant requests, 7 voice-assistant requests), each described in one line.
- Messages are short and in English.
When not to use it
- Your list of requests is very different from the training lists. Try it, but measure first; the triage adapter reads your own team descriptions.
- A message holds several requests and you need all of them. The question picks one.
- You need the reply written. Jeff chooses between options; it does not generate text.
- Your messages are long emails or tickets. Training messages are short single requests.
How to use it
The adapter runs with Jeff's server, on the main branch of firelex/jeff, on the
jeff-base v1.3 base.
git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora # add --extra cuda on NVIDIA GPUs, --extra mac on Apple silicon
uv run --no-default-groups hf download mstrasser/jeff-base --revision v1.3 --local-dir checkpoints/jeff-base
uv run --no-default-groups hf download mstrasser/jeff-adapter-support-intents --revision v1.3 --local-dir adapters/support-intents
JEFF_CHECKPOINT=checkpoints/jeff-base JEFF_ADAPTERS=adapters/ PORT=8765 \
uv run --no-default-groups jeff-serve # on a Mac, add JEFF_BACKEND=mlx
Every folder in adapters/ is served under its folder name; add or replace adapters while the server runs with
curl -X POST http://localhost:8765/v1/adapters/reload. Each adapter records the exact base it was trained on, and
the server refuses an adapter trained on a different one, so this adapter loads only on jeff-base v1.3 (a v1.2 adapter
does not load on v1.3). For llama.cpp, use mstrasser/jeff-adapter-support-intents-gguf.
Request format
State (the situation), in this order:
| Key | Changes per request | What it holds |
|---|---|---|
service |
no | One short phrase about the service the message is sent to. Training used three, one per data set, for example "Customer support chat of an online shop". |
message |
yes | The customer's or user's message as received. |
Questions:
intent(choice): What the customer wants, as the request that best matches their message. Options: In training, each message's options were all the requests of its own data set: 27 for the online shop (Bitext), 64 for the home assistant (HWU64) and 7 for the voice assistant (SNIPS). Keys are snake_case names such as track_order, each with a one-line description. The full lists are BITEXT, HWU64 and SNIPS in descriptions.py in the source.
Rules:
- Use the option keys and descriptions from descriptions.py where they fit your service; the adapter was trained on them.
- Use the instructions below word for word; the adapter was trained mostly on them.
General rules for every request: the request format guide.
Example
The request below is also in this repository as example.json.
{
"model": "support-intents",
"state": {
"service": "Customer support chat of an online shop",
"message": "hi, I sent back the jacket two weeks ago and still haven't seen the money. where is it?"
},
"questions": {
"intent": {
"type": "choice",
"instructions": "What does the customer want? Choose the request that best matches what the customer is asking for in their message.",
"criteria": {
"track_refund": "Check the status of a refund they are expecting.",
"get_refund": "Get their money back for a purchase.",
"check_refund_policy": "Learn the refund policy and whether they qualify for a refund.",
"track_order": "Find out where their order is or its current status.",
"cancel_order": "Cancel an order they placed.",
"change_order": "Change an existing order (for example add, remove or swap items).",
"payment_issue": "Report or solve a problem with a payment.",
"complaint": "Make a complaint about the product, service or company.",
"contact_human_agent": "Talk to a human agent instead of an automated assistant.",
"delivery_period": "Find out when an order will arrive or how long delivery takes."
}
}
}
}
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d @adapters/support-intents/example.json
The answer holds a probability for each option of each question. A recorded response from the v1.3 adapter is not published yet.
Files
adapter_model.safetensors,adapter_config.json: the LoRA weights (PEFT format);readout.safetensors: the adapter's own readout over the answer codes;decision_config.json: answer codes, temperature, prompt layout and the checksum of the base it was trained on;test.jsonl: the held-out test set the results below were measured on;calibration.jsonl: the calibration rows the adapter's temperature was fitted on;example.json: the example request above.
adapter_config.json and decision_config.json name the base as mstrasser/jeff-base, revision v1.3; the server
checks the base by the checksum of its weights.
Training
| Base | mstrasser/jeff-base, revision v1.3 (a fine-tune of Qwen3.5-0.8B) |
| Prompt layout | live-last: the fixed part of the request first, the changing state field last |
| Training code | The git_commit recorded in decision_config.json is the training machine's copy and was not published. It builds exactly the same prompt as main of firelex/jeff (from commit 6d0d7da) for a text state and for an object with at least one field; the format is in docs/v1.3-request-format.md |
| Run | 0.8b-support-intents-20261003-0149, final checkpoint |
| Adapter files | 41.5 MB (adapter_model.safetensors and readout.safetensors) |
| LoRA GGUF for llama.cpp | mstrasser/jeff-adapter-support-intents-gguf |
- 1.3.0 (2026-10-03): Trained on Jeff v1.3 with the live-last prompt layout (LoRA rank 16, one epoch, about 10% of the base model's own training data mixed in).
Data card
Report attached. The shortcut report and data card are included and pass the JeffHub checks; the numbers are the maintainers’ own. What the levels mean
- Test set: included in this repository as
test.jsonl, so anyone can check the numbers - Calibration rows: included in this repository as
calibration.jsonl, the rows its threshold is chosen on - QA report, sanitised: the data-quality checks run before training
How the test set was held out. 10% of the Bitext and HWU64 messages, held out by a stable hash of the text, plus the official SNIPS June 2017 held-out files; never trained on.
Training data. Built from public data sets, listed under Data and licence.
The attached QA report was written for the data of the previous release; the v1.3 data fixes the notes it left open. The QA report re-run on the v1.3 data is still to be attached.
The source data sets are public (listed under Data and licence). A script to rebuild our rows from them will follow.
Order and invoice numbers in the Bitext messages are drawn from one shared set of formats, so a number's format does not give the intent away.
Training mixed in a replay sample of the Jeff base model's own training data: 5,485 rows, about 10% on top of the adapter's 54,849 (inherited from the v1.2 recipe as a precaution; its effect has not been measured).
Data and licence
Adapter licence: Apache-2.0.
Qwen3.5-0.8B notice: these weights were modified from Qwen3.5-0.8B by the Jeff project: jeff-base is a fine-tune of Qwen3.5-0.8B, and this adapter was trained on top of it. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in LICENSE.
It was trained on:
Bitext customer support training data set (27 intents). Licence: CDLA-Sharing-1.0 (open, but shared or changed data must keep the same terms)
Only the customer message and intent columns are used; the assistant response column is never read. Under CDLA, models trained on the data are Results, not Data.
HWU64 (Liu et al. 2019, NLU-Evaluation-Data). Licence: CC-BY-4.0 (open licence)
Home-assistant requests. Four very small intents outside the usual 64 are left out.
SNIPS custom intent engines benchmark, June 2017 (Coucke et al. 2018). Licence: CC0-1.0 (open licence)
Seven voice-assistant intents.
Limitations
- Tied to jeff-base v1.3. It will not load on any other base or version; the server checks the base weights' checksum.
- Jeff chooses between the options you give it. It does not write text or reason in several steps.
- Calibration was fitted on this adapter's own calibration rows. On very different data, check it again.
- Everything listed under When not to use it above.
Links
- Adapter page: jeffhub.ai/adapters/support-intents
- Base model: mstrasser/jeff-base (revision v1.3)
- LoRA GGUF for llama.cpp: mstrasser/jeff-adapter-support-intents-gguf
- What changed in v1.3: release notes
- Code and server: github.com/firelex/jeff
Jeff is an independent project. It uses the same request format as Jev but is not affiliated with or endorsed by TypeSafe, the makers of Jev.
- Downloads last month
- 11