Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

jeff-adapter-code

Jeff-Code: information-gathering steps ahead of the coding model. In the Jeff-Code agent, takes the information-gathering steps ahead of Qwen3.8-27B (reads, listings, searches) and hands over when unsure.

A LoRA adapter for jeff-base v1.3, a small open decision model (a fine-tune of Qwen3.5-0.8B). You send a situation (the state) and questions with named options; Jeff returns a calibrated probability for every option from one forward pass, with no generated text to parse. One Jeff server loads the base once and any number of adapters beside it; each request picks an adapter by name ("model": "code").

Results

  • 62.4% vs 62.8%: pass rate of Jeff-Code vs Qwen3.8-27B alone; paired difference −0.2 points (95% interval −2.6 to +2.1), 1,242 paired tasks on 6 benchmarks
  • 47% faster (32% less time) per task: 0.68× Qwen alone's time on average (95% interval 0.64–0.72; median task 0.70×)
  • 14% less total time over all tasks combined (0.86×, 0.80–0.93)
  • 92% / 68%: offline accuracy on held-out scoring tasks: code (steps) / code-router (thinking)

Jeff-Code runs two adapters on the fixed Jeff v1.3 base: code takes the information-gathering steps (step threshold 0.40), and code-router decides whether Qwen thinks hard on a turn (thinking off unless P(xhigh), out of the four levels off/low/medium/xhigh, ≥ 0.6). The baseline is Qwen3.8-27B alone in the same Jeff-Code build with every Jeff feature off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it. Each task ran in both settings side by side, at the same time on the same Qwen server, and every comparison is paired by task. Only tasks Jeff never saw in training were used: the held-out splits of six benchmarks.

Benchmark Paired tasks Qwen alone Jeff-Code Difference, points (95% interval) Time per task
SWE-bench Verified 486 70.6% 70.8% +0.4 (−3.5 to +4.3) 0.63× (0.57–0.69)
SWE-rebench (2 rounds) 370 58.9% 58.1% −0.8 (−5.7 to +3.5) 0.66× (0.61–0.73)
Terminal-Bench Pro (2 rounds) 195 61.2% 62.8% +1.5 (−4.6 to +7.7) 0.64× (0.55–0.76)
Terminal-Bench 2.0 (40 tasks, 3 attempts each) 108 75.9% 70.0% −4.6 (−12.1 to +3.7) 0.96× (0.78–1.16)
SkillsBench 42 28.6% 31.0% +2.4 (−11.9 to +16.7) 0.91× (0.68–1.20)
Harbor Index 41 12.2% 9.8% −2.4 (−12.2 to +7.3) 0.71× (0.51–0.99)
All six, pooled 1,242 62.8% 62.4% −0.2 (−2.6 to +2.1) 0.68× (0.64–0.72)
  • No benchmark shows a clear pass-rate difference: every interval includes zero. SWE-rebench and Terminal-Bench Pro ran twice; their two rounds are combined, with intervals computed over both. Time per task is the geometric mean of the per-task time ratios (Jeff-Code's time divided by Qwen alone's); below 1 is faster. The pass rates count every finished session; the difference counts only tasks finished in both settings, so it is not exactly the gap between the two pass rates.
  • Total time drops less than time per task: in about 5% of tasks, Jeff-Code runs more than 30 minutes longer than Qwen alone, because it keeps going where Qwen alone gives up after a few minutes. The good news is, sometimes that pays off: in those tasks Jeff-Code solved 26 to Qwen's 24.
  • Thinking off on every turn (with the same safeguards and thinking limit) is faster still but clearly worse: −7.6 points (−10.6 to −4.5), −13.5 on Terminal-Bench 2.0. A less cautious router threshold (0.7) loses 4.5 points. Jeff's decisions are what keep the quality.
  • Terminal-Bench (original) and Terminal-Bench Science also ran, but both settings solve 0% of their tasks, so they are left out. 27 task pairs that hit an infrastructure failure twice are left out for both sides; about 10 long re-runs were still running when these numbers were taken.

The agent: the Jeff-Code repository (github.com/firelex/jeff-code). Offline accuracy reported by the Jeff-Code session: practically the same as the full fine-tune compared offline (step 92%, router 68% for both). Numbers from the evaluation report of 2026-10-05 18:23. Source: results/sources/v1.3/jeff-code-owner-supplied.json in the JeffHub repository (supplied by the maintainers).

Jeff-Code: the whole story

What it is. Jeff-Code is a coding agent based on Pi, with two Jeff v1.3 adapters trained specifically for Qwen3.8-27B: this one and its partner (code for steps, code-router for thinking). We forked Pi because its extension framework doesn't currently let a fast decision model sit deep enough inside the agent loop. Besides being useful to people who run Qwen3.8-27B locally as their daily coding model, it is an experiment: how far can a small, fast "System 1" model go inside a coding agent?

How it works. Jeff makes two kinds of decisions around every Qwen turn, each by a small adapter on the same Jeff base, in about 0.2 s each:

  1. Jeff works ahead of Qwen (code). If it can take the next information-gathering step itself (read a file, list a folder, search the code, check which tools are installed), it does. It picks the tool first and then its argument, and can take several steps in a row while it is confident. Qwen then starts its turn with those results already in front of it, rather than spending a slow turn fetching them. It can also run the tests or a build, repeat Qwen's last command and, with the run-approval setting the evaluation used, run a script Qwen wrote or install a missing package. Writing and editing files always stay with Qwen, and whenever Jeff is unsure, it hands over to Qwen. So Jeff does not only pick a tool; it also fills in the tool's argument.
  2. Jeff decides whether Qwen needs to think hard on its next turn (code-router). Thinking stays off unless Jeff is confident the turn needs it. In the evaluation, about three quarters of Qwen's turns ran with thinking off.

Because Jeff returns a calibrated probability for every choice, each behaviour is controlled by a single setting: the step threshold (0.40) and the thinking threshold (0.6).

How we trained it: Jeff predicts what Qwen would do next.

  • Steps: each training label is simply the step Qwen actually took next. If Jeff can take that step early, Qwen gets the result without spending a turn on it. These labels are built from Qwen sessions by code, with no other model involved.
  • Thinking: each recorded Qwen turn at full thinking was asked again with thinking off, then low, then medium. The label is the cheapest level whose action was as good as the original, or full thinking if none was. "As good" is decided by code wherever possible (the same kind of step on the same target); otherwise Qwen3.8-Max, with thinking off, judges whether the cheaper step would serve the task just as well at that moment. About 27% of turns didn't need thinking at all.
  • These labels are deliberately strict, which made the router cautious, and that is what preserved the pass rate. We also tried the looser question, "Is the full-thinking step materially better?". On a single turn the judge can't reliably tell thinking-off from a second full-thinking answer either, so a router trained that way would switch thinking off almost everywhere, and the thinking-off run shows what that costs over a whole task.

Adapters, not a fine-tune. Everything measured above used adapters: small LoRA files on the fixed Jeff v1.3 base. We also trained one full fine-tune to make both decisions and compared it offline; its accuracy was practically identical (step 92%, router 68% for both). So we release the adapters, which keep the multi-adapter design without giving up accuracy.

What changed from Pi.

  • Thinking per turn: Pi only switches thinking on or off, so Qwen always thought at its highest level (in our tests, Qwen's low and medium levels thought about as long as the highest one, so they saved no time). Jeff-Code sets Qwen's thinking level for each turn, fixed or decided by Jeff.
  • Safeguards for thinking off: a loop guard catches repeated or near-identical actions up to six steps back (a file write only counts as progress if it changes the file); a caught repeat is thrown away and that turn is asked again with full thinking, and near-identical outputs or two failed commands in a row send the next turn to full thinking. The thinking-off comparison had exactly the same safeguards, the same thinking limit and the same escalations (about 3% of its turns ended up thinking); the only difference from Jeff-Code is Jeff's decisions.
  • Runaway cut-off: if Qwen's thinking or text keeps repeating itself, the reply is stopped and asked again with full thinking. This was on in every run, including the baseline (it caught 2 replies there).
  • Thinking limit: at 8,000 thinking tokens, Qwen answers from what it has thought so far. That also rescues replies that would otherwise hit the 32K output limit, which ends a Pi session. The baseline ran without it, as plain Pi does.
  • Jeff steps: before each Qwen turn, Jeff-Code builds a menu of concrete next steps from what is already known, and Jeff takes them when it is confident.
  • A pool of Jeff servers, one per GPU, keeps each decision at about 0.2 s, and every Qwen request and Jeff decision is logged.

With llama.cpp (GGUF)

The same rows, through llama.cpp: the base GGUF (mstrasser/jeff-base-gguf) plus this adapter's LoRA GGUF (mstrasser/jeff-adapter-code-gguf), with the temperature refitted for each format. Running Jeff with llama.cpp

Test set Full precision Q8_0 Q4_K_M
development 90.4% · 0.023 90.4% · 0.030 91.0% · 0.029

Run-time rule: the same action (act on the top option at 0.40, otherwise hand over). The GGUF takes the same decision as full precision on 99.9% of the 991 development rows at Q8_0 and 97.9% at Q4_K_M.

Full precision: full precision from the trainer's own evaluation on the development split.

When to use it

  • You run the Jeff-Code agent with Qwen3.8-27B locally and want the information-gathering steps (reading a file, listing a folder, searching the code, checking which tools are installed) taken without a turn of the large model.
  • You want a single setting (the step threshold) that decides how much Jeff does on its own.

When not to use it

  • You use another coding agent or another large model. The adapter imitates Qwen3.8-27B's own next steps in a bash-only agent.
  • You want Jeff to write or edit files. Writing and editing always stay with Qwen; Jeff takes information steps, can run the tests or a build and repeat Qwen's last command, and, with the run-approval setting the evaluation used, can run a script Qwen wrote or install a missing package.

How to use it

The adapter runs with Jeff's server, on the main branch of firelex/jeff, on the jeff-base v1.3 base.

git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora          # add --extra cuda on NVIDIA GPUs, --extra mac on Apple silicon
uv run --no-default-groups hf download mstrasser/jeff-base --revision v1.3 --local-dir checkpoints/jeff-base
uv run --no-default-groups hf download mstrasser/jeff-adapter-code --revision v1.3 --local-dir adapters/code
JEFF_CHECKPOINT=checkpoints/jeff-base JEFF_ADAPTERS=adapters/ PORT=8765 \
  uv run --no-default-groups jeff-serve          # on a Mac, add JEFF_BACKEND=mlx

Every folder in adapters/ is served under its folder name; add or replace adapters while the server runs with curl -X POST http://localhost:8765/v1/adapters/reload. Each adapter records the exact base it was trained on, and the server refuses an adapter trained on a different one, so this adapter loads only on jeff-base v1.3 (a v1.2 adapter does not load on v1.3). For llama.cpp, use mstrasser/jeff-adapter-code-gguf.

Request format

State (the situation), in this order:

Key Changes per request What it holds
task no The coding task, as given to the agent.
recent_steps yes The last few steps (each tool call and its output), trimmed to about 2,000 tokens. For the argument question, also the tool chosen first.

Questions:

  • tool (choice): Which tool to use next, or hand over to the large model. Options: Read, List, Search, Check (which programs or packages are installed), and Hand over; a tool with nothing to choose from is left out
  • argument (choice): The argument for the chosen tool, from a list built by code. It only names files, folders, programs or packages already seen in the task or in earlier output. At most 25 options.

Rules:

  • Jeff-Code builds these requests itself; you do not write them by hand.
  • The request format may still change before release.

General rules for every request: the request format guide.

Example

The request below is also in this repository as example.json.

{
  "model": "code",
  "state": {
    "task": "The unit tests in tests/test_parser.py fail after the last change. Find out why and fix it.",
    "recent_steps": "bash: python -m pytest tests/test_parser.py -q -> FAILED tests/test_parser.py::test_dates - AttributeError: 'NoneType' object has no attribute 'group' (src/parser.py:42)"
  },
  "questions": {
    "tool": {
      "type": "choice",
      "instructions": "Gather what the coding model will need for its next step, and hand over as soon as more looking would not help. Which tool should run next?",
      "criteria": {
        "read": "Read: read a file, or a slice of it around a line named in an error",
        "search": "Search: search the code for a name or an error message",
        "check": "Check: see whether a program or package is installed",
        "hand_over": "Hand over: let the coding model take its turn"
      }
    }
  }
}
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d @adapters/code/example.json

The answer holds a probability for each option of each question. A recorded response from the v1.3 adapter is not published yet.

Files

  • adapter_model.safetensors, adapter_config.json: the LoRA weights (PEFT format);
  • readout.safetensors: the adapter's own readout over the answer codes;
  • decision_config.json: answer codes, temperature, prompt layout and the checksum of the base it was trained on;
  • example.json: the example request above.

adapter_config.json and decision_config.json name the base as mstrasser/jeff-base, revision v1.3; the server checks the base by the checksum of its weights.

Training

Base mstrasser/jeff-base, revision v1.3 (a fine-tune of Qwen3.5-0.8B)
Prompt layout live-last. Jeff-Code sends the state as plain text, and under live-last a text state keeps the original order: State, then Question, then Options, then the line asking for the letter code
Training code The git_commit recorded in decision_config.json is the training machine's copy and was not published. It builds exactly the same prompt as main of firelex/jeff (from commit 6d0d7da) for a text state and for an object with at least one field; the format is in docs/v1.3-request-format.md
LoRA GGUF for llama.cpp mstrasser/jeff-adapter-code-gguf
  • 1.3.0 (2026-10-03): First release, trained on Jeff v1.3 with the live-last prompt layout. Used in Jeff-Code's measured runs with the step threshold at 0.40.

Data card

Self-reported. The numbers come from the adapter’s own maintainers and have not been re-run by anyone else. What the levels mean

  • Test set: not attached yet
  • QA report: not available here yet; it will be added once sanitised

How the test set was held out. Whole tasks are held out: Jeff-Code was measured only on tasks never used for training. SWE-bench Verified ran in full (all 500 tasks; none of its repositories were used for training). For the benchmarks also used in training, the tasks were split and every held-out task was run: Terminal-Bench 2.0 (40 frozen tasks; four near-twins of them were also kept out of training), SWE-rebench (189), Terminal-Bench Pro (100), SkillsBench (44) and Harbor Index (41). A leak check compares every training task with the evaluation tasks.

Training data. Training data not published.

The row counts of the final training data are still to be added to this card.

How it works: Jeff works ahead of the large model. Before each of Qwen's turns, Jeff decides whether it can already take the next information-gathering step itself: read a file, list a folder, search the code, check which tools are installed. It picks the tool first, then its argument, and can take up to 8 steps in a row. Qwen then starts its turn with those results already in front of it. Whenever Jeff is unsure, or the next step changes something (writing, editing, running or installing), it hands over. Because Jeff gives a calibrated probability for every choice, its step threshold is a single setting: Jeff takes a page's top option only when its probability is at least the threshold (0.40 in the measured runs), and hands over otherwise. A second adapter, code-router, decides on every turn whether Qwen should think hard.

How we trained it: Jeff predicts what Qwen would do next. Each training label is the information-gathering step Qwen actually took next, built by code from Qwen sessions; if Jeff takes that step early, Qwen gets the result without spending a turn on it.

Training order: the adapter was trained in curriculum order, stage 1, then stage 2, then stage 3, with stage 1 sampled down to half the size of stage 3.

Questions longer than 8,192 tokens are cut to fit, keeping the head and tail of the state; the question and its options are never cut. The same rule applies at run time.

The data comes in three stages: (1) about 14,400 public Qwen3.8-27B agent sessions (ukisai/Qwen3.8-27B-multi-turn-agent-sft), with Jeff's option lists rebuilt from each transcript; (2) public Terminal-Bench sessions of Qwen3.8-27B, replayed in each task's container; (3) our own sessions of Qwen3.8-27B in Jeff-Code, where Jeff logs its full option list before every turn without acting.

Jeff only ever offers files, folders, programs or packages that have already appeared in the task or in earlier output, the same things Qwen could act on.

To keep training and real use identical, Qwen works with bash only in Jeff-Code.

Jeff-Code is a fork of the Pi coding agent (MIT), with its own changes such as a loop guard and Jeff acting first.

Jeff-Code's measured results come entirely from LoRA adapters on the fixed Jeff v1.3 base - this one and code-router - both loaded beside the other adapters on one jeff-base.

Data and licence

Adapter licence: Apache-2.0.

Qwen3.5-0.8B notice: these weights were modified from Qwen3.5-0.8B by the Jeff project: jeff-base is a fine-tune of Qwen3.5-0.8B, and this adapter was trained on top of it. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in LICENSE.

To confirm: that the licences of the session data and of Terminal-Bench 2.0 allow training and publishing the adapter.

It was trained on:

  • ukisai/Qwen3.8-27B-multi-turn-agent-sft (about 14,400 public Qwen3.8-27B agent sessions). Licence: To confirm (licence not confirmed yet) · Made by Qwen3.8-27B (the sessions); labels built by code

    Jeff's option lists are rebuilt from each transcript; the label is the step Qwen took next.

    To confirm: the licence of the ukisai data set on Hugging Face

  • openguardrails Terminal-Bench sessions of Qwen3.8-27B. Licence: To confirm (licence not confirmed yet) · Made by Qwen3.8-27B (the sessions); labels built by code

    Public sessions, replayed in each task's container.

    To confirm: the licence of the openguardrails Terminal-Bench sessions

  • Terminal-Bench 2.0 tasks (outside the 40 evaluation tasks and their four near-twins). Licence: To confirm (licence not confirmed yet) · Not made by a model

    The tasks our own Jeff-Code sessions and the replays run on.

    To confirm: the Terminal-Bench 2.0 licence

  • Our own Jeff-Code sessions with Qwen3.8-27B. Licence: Built for this adapter (made for this adapter) · Made by Qwen3.8-27B (local), the sessions; labels built by code

    Qwen3.8-27B in Jeff-Code (bash only); Jeff logs its full option list before every turn without acting.

Limitations

  • Tied to jeff-base v1.3. It will not load on any other base or version; the server checks the base weights' checksum.
  • Jeff chooses between the options you give it. It does not write text or reason in several steps.
  • Calibration was fitted on this adapter's own calibration rows. On very different data, check it again.
  • Everything listed under When not to use it above.

Links

Jeff is an independent project. It uses the same request format as Jev but is not affiliated with or endorsed by TypeSafe, the makers of Jev.

Downloads last month
19
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mstrasser/jeff-adapter-code

Adapter
(30)
this model

Dataset used to train mstrasser/jeff-adapter-code