Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

jeff-adapter-soc

Security alert triage. Applies a security team's playbook to a new alert - ignore, investigate, isolate, block or escalate.

A LoRA adapter for jeff-base v1.3, a small open decision model (a fine-tune of Qwen3.5-0.8B). You send a situation (the state) and questions with named options; Jeff returns a calibrated probability for every option from one forward pass, with no generated text to parse. One Jeff server loads the base once and any number of adapters beside it; each request picks an adapter by name ("model": "soc").

Adapter page, with the full data card: jeffhub.ai/adapters/soc.

Results

On this adapter's held-out test set, never trained on, scored three ways on the same rows: the untrained model Jeff is built from, the Jeff v1.3 base alone, and the base with this adapter. Questions have 4 to 5 options. As of 2026-10-05. All adapters

Test set Test rows Qwen3.5-0.8B untrained Jeff base v1.3 alone Jeff base v1.3 + adapter
test 4,929 22.1% · 0.011 33.8% · 0.032 94.1% · 0.014

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

By group

Group Rows Qwen3.5-0.8B untrained Jeff base v1.3 alone Jeff base v1.3 + adapter
alert source: endpoint 1,500 23.0% 34.7% 96.3%
alert source: insider 1,429 22.3% 29.7% 87.8%
alert source: network 2,000 21.4% 36.2% 97.0%
label: block 673 34.5% 34.5% 95.4%
label: escalate 1,085 32.4% 15.9% 95.2%
label: ignore 1,165 6.9% 63.8% 92.0%
label: investigate 1,237 18.6% 15.2% 92.5%
label: isolate 769 25.6% 43.3% 97.3%

With llama.cpp (GGUF)

The same test, through llama.cpp: the base GGUF (mstrasser/jeff-base-gguf) plus this adapter's LoRA GGUF (mstrasser/jeff-adapter-soc-gguf), with the temperature refitted for each format. Running Jeff with llama.cpp

Test set Full precision Q8_0 Q4_K_M
test 94.1% · 0.014 94.0% · 0.013 93.7% · 0.009

Not measured yet for v1.3: calibration charts, the commonest confusions and accuracy per answer.

Source of these numbers: results/sources/v1.3/new-adapters.table.json in the JeffHub repository, also collected in jeffhub-v1.3.json.

When to use it

  • Your security operations team has a written playbook (rules in order, thresholds, an asset inventory, allowed actions) and you want each new alert sorted in milliseconds.
  • You want the playbook's own answer with a calibrated probability, so the unsure alerts go to an analyst first.
  • Your alerts are network detections, endpoint file alerts or user-behaviour alerts; the playbook and the alert can be your own wording.

When not to use it

  • You want the model to find threats your playbook does not describe. It applies the rules you give it.
  • Alerts span several events that must be read together (a scan followed by a login burst). Training alerts are single events or single working days.
  • You plan commercial use or to publish data built on the insider alerts. Two sources restrict this (see Data and licence).

How to use it

The adapter runs with Jeff's server on the jeff-base v1.3 base. Adapter serving arrives with the next Jeff release; until then, these commands need the feat/lora branch of firelex/jeff.

git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora          # add --extra cuda on NVIDIA GPUs, --extra mac on Apple silicon
uv run --no-default-groups hf download mstrasser/jeff-base --revision v1.3 --local-dir checkpoints/jeff-base
uv run --no-default-groups hf download mstrasser/jeff-adapter-soc --revision v1.3 --local-dir adapters/soc
JEFF_CHECKPOINT=checkpoints/jeff-base JEFF_ADAPTERS=adapters/ PORT=8765 \
  uv run --no-default-groups jeff-serve          # on a Mac, add JEFF_BACKEND=mlx

Every folder in adapters/ is served under its folder name; add or replace adapters while the server runs with curl -X POST http://localhost:8765/v1/adapters/reload. Each adapter records the exact base it was trained on, and the server refuses an adapter trained on a different one, so this adapter loads only on jeff-base v1.3 (a v1.2 adapter does not load on v1.3). For llama.cpp, use mstrasser/jeff-adapter-soc-gguf.

Request format

State (the situation), in this order:

Key Changes per request What it holds
organisation no One line on the organisation, its sector and which alert source this queue handles.
playbook no The playbook: numbered rules (the first that matches decides), thresholds, severity table, asset inventory or office policy, allowed actions, and what to do when a rule calls for an action the team may not take.
alert yes The new alert as your tools report it - what fired, on which host or user, the evidence and any lookups.

Questions:

  • action (choice): Which action the playbook requires. Options: The playbook's allowed actions: ignore, investigate, escalate always; isolate and block when the team may take them.

Rules:

  • Offer only the actions your playbook allows; the playbook's own fallback rule says what to do when a rule calls for another.
  • Use the instructions below word for word; every training row used them.
  • Keep the organisation line and playbook the same across alerts, and put the alert last, so the unchanging part can be prepared in advance.

General rules for every request: the request format guide.

Example

The request below is also in this repository as example.json.

{
  "model": "soc",
  "state": {
    "organisation": "Kestrel Stores, a retail organisation. This alert queue receives endpoint alerts on suspicious files.",
    "playbook": "Kestrel Stores endpoint alert card. Covers suspicious files found on company computers. Use it for every new alert from the endpoint protection agents.\nDefinitions: The file reputation service returns one of: malicious (with the malware family), known good, unknown (the file has not been rated yet), or lookup failed. A file that ran was executed on the computer; otherwise it was only written to disk.\nAsset inventory (host, role, criticality): BKP-40 (backup server, critical); MTG-PC-34 (meeting-room computer, important); ERP-DB-19 (ERP database server, critical); PKI-04 (certificate authority server, important); DEV-WS-35 (developer workstation, standard); TEST-WEB-29 (test web server, standard); CRM-38 (CRM server, important); EXEC-LT-05 (executive laptop, important); INTRA-31 (intranet web server, important); HR-APP-37 (HR application server, important).\nSuspicious traits: A file is suspicious if it has no digital signature. Also if a section has entropy above 7.18. Also if it imports any of these: AdjustTokenPrivileges, CryptAcquireContextA, CryptEncrypt, InternetOpenUrlW, SetWindowsHookExW, URLDownloadToFileA, WinExec. Also if it imports no functions at all.\nRules in order, first match decides:\n(1) Unknown hosts: if the host is not in the asset inventory, escalate the alert.\n(2) Malicious files. Reputation is malicious? Check the family. Family is delf, downloadguide, high, kovter, vittalia or zamg? Escalate. File ran? Isolate the host. File did not run? Block the file. Critical hosts are never isolated automatically. Rule says isolate a critical host? Escalate instead.\n(3) Failed lookups: if the reputation lookup failed, escalate when the host is critical. Otherwise investigate.\n(4) Unknown files: reputation unknown. Count the suspicious traits. With 2 or more traits, block the file if it did not run. If it ran, investigate. With fewer traits, investigate when the host is critical. Otherwise ignore the alert.\n(5) Known good files: reputation known good. Ignore the alert.\nActions: This team may ignore, investigate, isolate, block or escalate alerts. If a rule calls for an action that is not allowed here, escalate instead.",
    "alert": "Sunday 14:48 EDR ALERT host=CRM-38 user=f.jensen file=\"D:\\Shared\\Apps\\invoice_viewer.exe\" sha256=53eb6d2d440384ac885525f1af04c3688f7080a5f034e54325bfbeea21922e7b size=139264 first_seen=2018-11 execution=\"ran\" signed=no sections=\".text 5.80, .rdata 5.02, .data 3.98, .rsrc 7.43\" imports_total=54 notable_imports=\"none of note\" urls=0 reputation=\"malicious (family zbot)\""
  },
  "questions": {
    "action": {
      "type": "choice",
      "instructions": "You are the first-line analyst in a security operations centre. Apply the organisation's playbook to the new alert and choose the action the playbook requires. The playbook's rules are checked in order and the first rule that matches decides.",
      "criteria": {
        "isolate": "Isolate: cut the affected internal computer off the network.",
        "investigate": "Investigate: open a case for an analyst; no containment now.",
        "escalate": "Escalate: hand the alert at once to the senior responder.",
        "ignore": "Ignore: close the alert because no action at all is needed.",
        "block": "Block: block the outside address, file or account involved."
      }
    }
  }
}
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d @adapters/soc/example.json

The answer holds a probability for each option of each question. A recorded response from the v1.3 adapter is not published yet.

Files

  • adapter_model.safetensors, adapter_config.json: the LoRA weights (PEFT format);
  • readout.safetensors: the adapter's own readout over the answer codes;
  • decision_config.json: answer codes, temperature, prompt layout and the checksum of the base it was trained on;
  • example.json: the example request above.

adapter_config.json and decision_config.json name the base as mstrasser/jeff-base, revision v1.3; the server checks the base by the checksum of its weights.

Training

Base mstrasser/jeff-base, revision v1.3 (a fine-tune of Qwen3.5-0.8B)
Prompt layout live-last: the fixed part of the request first, the changing state field last
Run 0.8b-soc-20261003-1942, final checkpoint
Adapter files 41.5 MB (adapter_model.safetensors and readout.safetensors)
LoRA GGUF for llama.cpp mstrasser/jeff-adapter-soc-gguf
  • 1.3.0 (2026-10-03): First release, trained on Jeff v1.3 with the live-last prompt layout (LoRA rank 16, one epoch, about 10% of the base model's own training data mixed in).

Data card

Self-reported. The numbers come from the adapter’s own maintainers and have not been re-run by anyone else. What the levels mean

  • Test set: not attached yet
  • QA report: not available here yet; it will be added once sanitised

How the test set was held out. Whole playbooks held out (36 organisations never trained on), and the underlying data too - the network flows of UNSW-NB15's official testing set, EMBER's test-set files, and CERT employees who never appear in training.

Training data. Training data not published.

Which models made the data, counted on the 43,639 training rows:

What it did Model Where it ran Training rows
Rewrote the playbook rules in one of five styles (checked by code to keep every number, name and action word) playbook_text.model GLM 5.3 own hardware (DGX B200) 28,485
Rewrote a network connection summary (checked by code) flow_text GLM 5.3 own hardware (DGX B200) 5,016
Rewrote an endpoint or insider alert in one of four styles (checked by code) alert_text GLM 5.3 own hardware (DGX B200) 4,449

Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total. Every label is computed by code by applying the playbook's rules, in order, to the facts the alert shows. GLM never sees or sets a label.

Training mixed in a replay sample of the Jeff base model's own training data: 4,364 rows, about 10% on top of the adapter's 43,639 (inherited from the v1.2 recipe as a precaution; its effect has not been measured).

Three alert sources: network (UNSW-NB15 flows; the true attack category enters as the detection), endpoint (EMBER 2018 files; the true label enters as the file reputation) and insider (CERT r4.2 user-days; the answer key chooses the scenario days).

About 55% of rows are hard cases built on purpose - values just either side of a threshold, conflicting signals, missing information, and rules that call for an action the team may not take.

The adapter learns to apply a stated policy to the facts shown. It does not judge on its own whether traffic, a file or an employee is malicious.

In a blind check, GLM 5.3 answered 200 random test rows without the label and agreed on all 200.

Training mixes in about 10% of the base model's own training data.

Data and licence

Adapter licence: CC BY-NC 4.0 (Creative Commons Attribution-NonCommercial 4.0: free to use and share with credit, but not for commercial purposes). The JeffHub page still lists the adapter's licence as to be decided; this release uses CC BY-NC 4.0.

Qwen3.5-0.8B notice: these weights were modified from Qwen3.5-0.8B by the Jeff project: jeff-base is a fine-tune of Qwen3.5-0.8B, and this adapter was trained on top of it. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in LICENSE.

To confirm: the adapter's licence and whether it may be released, given the UNSW-NB15 and CERT terms.

It was trained on:

  • UNSW-NB15 (network alerts). Licence: Free use for academic research; commercial use prohibited (ReadMe.pdf, copyright Nour Moustafa) (non-commercial use only) · Not made by a model

    Moustafa and Slay, "UNSW-NB15 - a comprehensive data set for network intrusion detection systems", MilCIS 2015.

    To confirm: whether an adapter trained on UNSW-NB15 may be published or used commercially

  • EMBER 2018 (endpoint alerts). Licence: MIT (data files; the code in the same repository is AGPL-3.0) (open licence) · Not made by a model

    Anderson and Roth, "EMBER - An Open Dataset for Training Static PE Malware Machine Learning Models", arXiv:1804.04637, 2018.

  • CERT Insider Threat Test Dataset r4.2 (insider alerts). Licence: ExactData end-user agreement, "Copyright 2011 ExactData, LLC, All Rights Reserved"; no redistribution, derivative works only to the minimal extent necessary (all rights reserved; publishing anything built on it needs the owner’s permission) · Not made by a model

    Lindauer, Insider Threat Test Dataset, Carnegie Mellon University, 2020. The data is synthetic (CERT's generator), from 2010-2011.

    To confirm: with the licence owner, whether anything built on the insider rows (the adapter, the test set) may be published

  • Synthetic playbooks and alert context. Licence: Built for this adapter; terms follow the adapter's licence (made for this adapter) · Made by GLM 5.3 (own hardware), for the reworded texts; everything else by code

    285 fictional organisations with their playbooks, inventories and thresholds, and the alert context (hosts, directions, lookups), generated by code with fixed seeds. Playbooks and some alerts were reworded by GLM 5.3 and checked by code.

Limitations

  • Tied to jeff-base v1.3. It will not load on any other base or version; the server checks the base weights' checksum.
  • Jeff chooses between the options you give it. It does not write text or reason in several steps.
  • Calibration was fitted on this adapter's own calibration rows. On very different data, check it again.
  • Everything listed under When not to use it above.

Links

Jeff is an independent project. It uses the same request format as Jev but is not affiliated with or endorsed by TypeSafe, the makers of Jev.

Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mstrasser/jeff-adapter-soc

Adapter
(30)
this model

Paper for mstrasser/jeff-adapter-soc