Instructions to use mstrasser/jeff-adapter-spam with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use mstrasser/jeff-adapter-spam with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
jeff-adapter-spam
Spam and phishing in SMS and email. Says whether a text message or email is spam or phishing, and whether it is legitimate, spam or a phishing scam.
A LoRA adapter for jeff-base v1.3, a small open decision model
(a fine-tune of Qwen3.5-0.8B). You send a situation (the state) and questions with named options; Jeff returns a
calibrated probability for every option from one forward pass, with no generated text to parse. One Jeff server loads
the base once and any number of adapters beside it; each request picks an adapter by name ("model": "spam").
Adapter page, with the full data card: jeffhub.ai/adapters/spam.
Results
On this adapter's held-out test sets, never trained on, scored three ways on the same rows: the untrained model Jeff is built from, the Jeff v1.3 base alone, and the base with this adapter. Questions have 2 to 3 options. As of 2026-10-05. All adapters
| Test set | Test rows | Qwen3.5-0.8B untrained | Jeff base v1.3 alone | Jeff base v1.3 + adapter |
|---|---|---|---|---|
test |
3,897 | 58.1% · 0.047 | 69.8% · 0.046 | 98.1% · 0.006 |
test-v12 |
3,603 | 59.3% · 0.060 | 72.2% · 0.060 | 98.1% · 0.007 |
Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).
test.jsonl in this repository is the test set; test-v12 is not included.
By group (test set test)
| Group | Rows | Qwen3.5-0.8B untrained | Jeff base v1.3 alone | Jeff base v1.3 + adapter |
|---|---|---|---|---|
| choice: legitimate / phishing / spam (email) | 172 | 57.0% | 58.1% | 97.7% |
| choice: legitimate / phishing / spam (sms) | 82 | 82.9% | 62.2% | 97.6% |
| choice: legitimate / spam / phishing (email) | 163 | 33.1% | 68.1% | 99.4% |
| choice: legitimate / spam / phishing (sms) | 108 | 25.0% | 70.4% | 93.5% |
| choice: phishing / legitimate / spam (email) | 163 | 57.7% | 64.4% | 98.2% |
| choice: phishing / legitimate / spam (sms) | 102 | 87.3% | 55.9% | 98.0% |
| choice: phishing / spam / legitimate (email) | 149 | 53.0% | 56.4% | 99.3% |
| choice: phishing / spam / legitimate (sms) | 93 | 79.6% | 78.5% | 98.9% |
| choice: spam / legitimate / phishing (email) | 178 | 61.8% | 59.6% | 97.2% |
| choice: spam / legitimate / phishing (sms) | 110 | 83.6% | 55.5% | 98.2% |
| choice: spam / phishing / legitimate (email) | 168 | 61.9% | 59.5% | 98.2% |
| choice: spam / phishing / legitimate (sms) | 78 | 84.6% | 60.3% | 96.2% |
| yes/no question (email) | 1,739 | 64.6% | 72.0% | 98.3% |
| yes/no question (sms) | 592 | 31.8% | 84.1% | 98.3% |
By group (test set test-v12)
| Group | Rows | Qwen3.5-0.8B untrained | Jeff base v1.3 alone | Jeff base v1.3 + adapter |
|---|---|---|---|---|
| choice: legitimate / phishing / spam (email) | 144 | 58.3% | 66.0% | 97.2% |
| choice: legitimate / phishing / spam (sms) | 82 | 82.9% | 62.2% | 97.6% |
| choice: legitimate / spam / phishing (email) | 149 | 28.9% | 73.2% | 99.3% |
| choice: legitimate / spam / phishing (sms) | 108 | 25.0% | 69.4% | 93.5% |
| choice: phishing / legitimate / spam (email) | 134 | 70.1% | 72.4% | 98.5% |
| choice: phishing / legitimate / spam (sms) | 102 | 87.3% | 57.8% | 98.0% |
| choice: phishing / spam / legitimate (email) | 123 | 64.2% | 65.9% | 99.2% |
| choice: phishing / spam / legitimate (sms) | 93 | 79.6% | 78.5% | 98.9% |
| choice: spam / legitimate / phishing (email) | 157 | 68.2% | 61.8% | 96.8% |
| choice: spam / legitimate / phishing (sms) | 110 | 83.6% | 54.5% | 98.2% |
| choice: spam / phishing / legitimate (email) | 139 | 72.7% | 64.7% | 97.8% |
| choice: spam / phishing / legitimate (sms) | 78 | 84.6% | 59.0% | 96.2% |
| yes/no question (email) | 1,592 | 64.3% | 73.4% | 98.3% |
| yes/no question (sms) | 592 | 32.3% | 84.1% | 98.3% |
With llama.cpp (GGUF)
The same test, through llama.cpp: the base GGUF (mstrasser/jeff-base-gguf) plus this adapter's LoRA GGUF (mstrasser/jeff-adapter-spam-gguf), with the temperature refitted for each format. Running Jeff with llama.cpp
| Test set | Full precision | Q8_0 | Q4_K_M |
|---|---|---|---|
test |
98.1% · 0.006 | 98.1% · 0.004 | 97.8% · 0.005 |
test-v12 |
98.1% · 0.007 | 98.0% · 0.005 | 97.6% · 0.007 |
Not measured yet for v1.3: calibration charts, the commonest confusions and accuracy per answer.
Source of these numbers: results/sources/v1.3/retrained-adapters.table.json in the JeffHub repository, also collected in jeffhub-v1.3.json.
When to use it
- You receive text messages or emails and want a spam or phishing probability for each.
- You want to tell phishing (a scam after details, passwords or money) apart from ordinary advertising spam.
- You set your own threshold on the yes probability, rather than taking a hard yes or no.
When not to use it
- You need an email's headers, links or attachments judged. Only the subject and readable body were used, and bodies over 6,000 characters were cut.
- Your messages are not in English. The data sets are English.
- You need to catch the newest scams. Much of the data is old (SMS from 2011 and 2022, most email from 2002 to 2005, phishing email up to 2025), so recent tricks may be missed.
- You block messages with no human check. Treat the answer as a signal, not a verdict.
How to use it
The adapter runs with Jeff's server, on the main branch of firelex/jeff, on the
jeff-base v1.3 base.
git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora # add --extra cuda on NVIDIA GPUs, --extra mac on Apple silicon
uv run --no-default-groups hf download mstrasser/jeff-base --revision v1.3 --local-dir checkpoints/jeff-base
uv run --no-default-groups hf download mstrasser/jeff-adapter-spam --revision v1.3 --local-dir adapters/spam
JEFF_CHECKPOINT=checkpoints/jeff-base JEFF_ADAPTERS=adapters/ PORT=8765 \
uv run --no-default-groups jeff-serve # on a Mac, add JEFF_BACKEND=mlx
Every folder in adapters/ is served under its folder name; add or replace adapters while the server runs with
curl -X POST http://localhost:8765/v1/adapters/reload. Each adapter records the exact base it was trained on, and
the server refuses an adapter trained on a different one, so this adapter loads only on jeff-base v1.3 (a v1.2 adapter
does not load on v1.3). For llama.cpp, use mstrasser/jeff-adapter-spam-gguf.
Request format
State (the situation), in this order:
| Key | Changes per request | What it holds |
|---|---|---|
channel |
no | Where the message came from: "sms" or "email", as in training. |
message |
yes | The text of the message as received. For an email, its subject and readable body (no other headers, no attachments). |
Questions:
is_spam(noul): Whether the message is spam (unwanted bulk or advertising) or phishing (a scam after personal details or money), rather than a normal message.message_type(choice): What kind of message it is. Options: Three options: legitimate, spam and phishing, each with a one-line description (SPAM_CHOICE in descriptions.py in the source).
Rules:
- Ask both questions in one request if you need both; they are answered together.
- For is_spam, give the false and true descriptions below, in that order; the adapter was trained with them.
- The yes/no question was trained on every message. The three-way question was trained only where the source separates phishing from spam, the SMS phishing set and the SpamAssassin and Nazario emails.
- Use the instructions below word for word; the adapter was trained mostly on them.
General rules for every request: the request format guide.
Example
The request below is also in this repository as example.json.
{
"model": "spam",
"state": {
"channel": "sms",
"message": "Your parcel could not be delivered. Pay the 1.99 redelivery fee within 24 hours at parcel-redeliver-help.example to avoid return."
},
"questions": {
"is_spam": {
"type": "noul",
"instructions": "Is this message spam or phishing? Answer yes if it is unwanted bulk or advertising, or a scam trying to get personal details or money; answer no if it is a normal message.",
"criteria": {
"false": "The message is a normal message: not unwanted bulk or advertising, and not a scam.",
"true": "The message is spam (unwanted bulk or advertising) or phishing (a scam trying to get personal details, passwords or money)."
}
},
"message_type": {
"type": "choice",
"instructions": "What kind of message is this: a legitimate message, spam (unwanted bulk or advertising), or phishing (a scam trying to get personal details, passwords or money)?",
"criteria": {
"legitimate": "A normal message from a person or a genuine organisation, not spam and not phishing.",
"spam": "Unwanted bulk or advertising message (for example prizes, offers, premium-rate services), but not trying to steal personal or account details.",
"phishing": "A scam message that tries to trick the reader into giving personal details, passwords or money, often by pretending to be a bank, company or authority and asking them to click a link or call a number."
}
}
}
}
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d @adapters/spam/example.json
The answer holds a probability for each option of each question. A recorded response from the v1.3 adapter is not published yet.
Files
adapter_model.safetensors,adapter_config.json: the LoRA weights (PEFT format);readout.safetensors: the adapter's own readout over the answer codes;decision_config.json: answer codes, temperature, prompt layout and the checksum of the base it was trained on;test.jsonl: the held-out test set the results below were measured on;calibration.jsonl: the calibration rows the adapter's temperature was fitted on;example.json: the example request above.
adapter_config.json and decision_config.json name the base as mstrasser/jeff-base, revision v1.3; the server
checks the base by the checksum of its weights.
Training
| Base | mstrasser/jeff-base, revision v1.3 (a fine-tune of Qwen3.5-0.8B) |
| Prompt layout | live-last: the fixed part of the request first, the changing state field last |
| Training code | The git_commit recorded in decision_config.json is the training machine's copy and was not published. It builds exactly the same prompt as main of firelex/jeff (from commit 6d0d7da) for a text state and for an object with at least one field; the format is in docs/v1.3-request-format.md |
| Run | 0.8b-spam-20261003-0235, final checkpoint |
| Adapter files | 41.5 MB (adapter_model.safetensors and readout.safetensors) |
| LoRA GGUF for llama.cpp | mstrasser/jeff-adapter-spam-gguf |
- 1.3.0 (2026-10-03): Trained on Jeff v1.3 with the live-last prompt layout (LoRA rank 16, one epoch, about 10% of the base model's own training data mixed in).
Data card
Report attached. The shortcut report and data card are included and pass the JeffHub checks; the numbers are the maintainers’ own. What the levels mean
- Test set: included in this repository as
test.jsonl, so anyone can check the numbers - Calibration rows: included in this repository as
calibration.jsonl, the rows its threshold is chosen on - QA report, sanitised: the data-quality checks run before training
How the test set was held out. 10% of the text messages and 10% of the emails, held out by a stable hash of the text (no source set has an official test split); never trained on. It includes 147 generated phishing and spam emails (294 rows). Scored on both questions.
Training data. Training data not published.
Which models made the data, counted on the 30,251 training rows:
| What it did | Model | Where it ran | Training rows |
|---|---|---|---|
Wrote the email teacher_model |
Qwen3.8-Max | hosted (Alibaba Cloud DashScope) | 7,702 |
Checked the email blind (named its kind) check_model |
Qwen3.8-Flash | hosted (Alibaba Cloud DashScope) | 7,702 |
Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total. 5,364 of these rows are legitimate emails and 2,338 are phishing or spam emails written in the same polished style, so a tidy, signed email does not mean legitimate. All other rows come from the public collections.
The attached QA report was written for the data of the previous release; the v1.3 data fixes the notes it left open. The QA report re-run on the v1.3 data is still to be attached.
The public sources are listed under Data and licence; a script to rebuild those rows from them will follow. The generated emails are not published.
The terms of the hosted model providers are being checked for training and publication use.
Training mixed in a replay sample of the Jeff base model's own training data: 3,025 rows, about 10% on top of the adapter's 30,251 (inherited from the v1.2 recipe as a precaution; its effect has not been measured).
Data and licence
Adapter licence: Apache-2.0.
Qwen3.5-0.8B notice: these weights were modified from Qwen3.5-0.8B by the Jeff project: jeff-base is a fine-tune of Qwen3.5-0.8B, and this adapter was trained on top of it. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in LICENSE.
It was trained on:
UCI SMS Spam Collection (Almeida and Hidalgo 2011). Licence: CC-BY-4.0 (open licence)
Labels ham and spam.
SMS Phishing Dataset for Machine Learning and Pattern Recognition (Mishra and Soni 2022), version 1. Licence: CC-BY-4.0 (open licence)
Labels ham, spam and smishing (SMS phishing). Messages shared with the UCI set were merged by text.
Phishing Email Dataset (zefang-liu, a copy of the Kaggle set "Phishing Email Detection"). Licence: No licence given by the source (no licence given by the source)
Tagged LGPL-3.0 by the uploader; the Kaggle original does not say where its emails come from. Labels safe and phishing.
ealvaradob/phishing-dataset (emails only). Licence: No licence given by the source (no licence given by the source)
Tagged Apache-2.0 by the compiler; its emails are the same Kaggle set as above, so it adds only a handful of messages.
SpamAssassin public mail corpus (ham, hard ham, spam). Licence: No licence given by the source (no licence given by the source)
Real mail received 2002 to 2005, sorted by hand into spam and non-spam. No licence is given; the corpus says copyright in the messages stays with their senders.
Generated emails. Licence: Released with the adapter under Apache-2.0 (made for this adapter) · Made by Qwen3.8-Max (hosted); see the data card
Legitimate, phishing and spam emails in a polished modern style, written by a language model and kept only when a second model, shown each email alone, named its intended kind (which models, and for how many rows, is in the data card).
Nazario phishing corpus. Licence: CC-BY-4.0, as the corpus states (open licence)
Phishing mail collected and sorted by hand by Jose Nazario.
Limitations
- Tied to jeff-base v1.3. It will not load on any other base or version; the server checks the base weights' checksum.
- Jeff chooses between the options you give it. It does not write text or reason in several steps.
- Calibration was fitted on this adapter's own calibration rows. On very different data, check it again.
- Everything listed under When not to use it above.
Links
- Adapter page: jeffhub.ai/adapters/spam
- Base model: mstrasser/jeff-base (revision v1.3)
- LoRA GGUF for llama.cpp: mstrasser/jeff-adapter-spam-gguf
- What changed in v1.3: release notes
- Code and server: github.com/firelex/jeff
Jeff is an independent project. It uses the same request format as Jev but is not affiliated with or endorsed by TypeSafe, the makers of Jev.
- Downloads last month
- 9