mstrasser commited on
Commit
e6c013d
Β·
verified Β·
1 Parent(s): 41321ce

Card: final Jeff-Code results, per-benchmark table, full narrative

Browse files
Files changed (1) hide show
  1. README.md +52 -8
README.md CHANGED
@@ -29,14 +29,59 @@ Adapter page, with the full data card: [jeffhub.ai/adapters/code](https://jeffhu
29
 
30
  ## Results
31
 
32
- - **53.4% vs 53.1%**: pass rate of Jeff-Code vs Qwen3.8-27B alone, 493 tasks on 5 benchmarks
33
- - **about 20% less**: time per session on average on SWE-rebench and Terminal-Bench 2.0
34
- - **30–45% faster**: the typical session
35
  - **92% / 68%**: offline accuracy on held-out scoring tasks: code (steps) / code-router (thinking)
36
 
37
- Jeff-Code runs two adapters on the fixed Jeff v1.3 base: [`code`](https://jeffhub.ai/adapters/code) takes the information-gathering steps (step threshold 0.40), and [`code-router`](https://jeffhub.ai/adapters/code-router) decides whether Qwen thinks hard on a turn (thinking off unless P(xhigh), out of the four levels off/low/medium/xhigh, β‰₯ 0.6). With Jeff's thinking threshold at 0.6, it matches Qwen3.8-27B's pass rate on five benchmarks: pooled 53.4% against 53.1% over 493 tasks, run side by side in paired blocks. It needs about 20% less time per session on average on the big benchmarks, SWE-rebench and Terminal-Bench 2.0; the typical session is 30–45% faster. Thinking off throughout is fast but clearly worse (βˆ’12 to βˆ’6 points): Jeff's decisions are what keep the quality.
38
 
39
- The agent: [the Jeff-Code repository](https://github.com/firelex/jeff-code) (github.com/firelex/jeff-code). Offline accuracy reported by the Jeff-Code session: practically the same as the full fine-tune compared offline (step 92%, router 68% for both). Numbers supplied by the maintainers on 2026-10-05; the detailed result files (pass rate per benchmark, time per session) are still to come. Source: `results/sources/v1.3/jeff-code-owner-supplied.json` in the JeffHub repository (supplied by the maintainers).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
40
 
41
  ### With llama.cpp (GGUF)
42
 
@@ -59,7 +104,6 @@ Full precision: full precision from the trainer's own evaluation on the developm
59
 
60
  - You use another coding agent or another large model. The adapter imitates Qwen3.8-27B's own next steps in a bash-only agent.
61
  - You want Jeff to write, edit, run or install anything. It only takes information-gathering steps and hands over before anything changes.
62
- - You need per-benchmark numbers now. The results on this page are the summary; the detailed result files are still to come.
63
 
64
  ## How to use it
65
 
@@ -152,7 +196,7 @@ checks the base by the checksum of its weights.
152
  | Prompt layout | live-last: the fixed part of the request first, the changing state field last |
153
  | LoRA GGUF for llama.cpp | [mstrasser/jeff-adapter-code-gguf](https://huggingface.co/mstrasser/jeff-adapter-code-gguf) |
154
 
155
- - **1.3.0** (2026-10-03): First release, trained on Jeff v1.3 with the live-last prompt layout. Results with Qwen3.8-27B supplied by the maintainers; the detailed result files are still to come.
156
 
157
  ## Data card
158
 
@@ -161,7 +205,7 @@ checks the base by the checksum of its weights.
161
  - Test set: not attached yet
162
  - QA report: not available here yet; it will be added once sanitised
163
 
164
- **How the test set was held out.** Whole tasks are held out: the evaluation uses 40 frozen Terminal-Bench 2.0 tasks, never used for training. Training uses the remaining Terminal-Bench 2.0 tasks minus four near-twins of evaluation tasks, and a leak check compares every training task with the evaluation tasks.
165
 
166
  **Training data.** Training data not published.
167
 
 
29
 
30
  ## Results
31
 
32
+ - **62.4% vs 62.8%**: pass rate of Jeff-Code vs Qwen3.8-27B alone; paired difference βˆ’0.2 points (95% interval βˆ’2.6 to +2.1), 1,242 paired tasks on 6 benchmarks
33
+ - **47% faster (32% less time) per task**: 0.68Γ— Qwen alone's time on average (95% interval 0.64–0.72; median task 0.70Γ—)
34
+ - **14% less total time** over all tasks combined (0.86Γ—, 0.80–0.93)
35
  - **92% / 68%**: offline accuracy on held-out scoring tasks: code (steps) / code-router (thinking)
36
 
37
+ Jeff-Code runs two adapters on the fixed Jeff v1.3 base: [`code`](https://jeffhub.ai/adapters/code) takes the information-gathering steps (step threshold 0.40), and [`code-router`](https://jeffhub.ai/adapters/code-router) decides whether Qwen thinks hard on a turn (thinking off unless P(xhigh), out of the four levels off/low/medium/xhigh, β‰₯ 0.6). The baseline is Qwen3.8-27B alone in the same Jeff-Code build with every Jeff feature off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it. Each task ran in both settings side by side, at the same time on the same Qwen server, and every comparison is paired by task. Only tasks Jeff never saw in training were used.
38
 
39
+ | Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
40
+ |---|---:|---:|---:|---|---:|
41
+ | SWE-bench Verified | 486 | 70.6% | 70.8% | +0.4 (βˆ’3.5 to +4.3) | 0.63Γ— (0.57–0.69) |
42
+ | SWE-rebench, round 1 | 184 | 57.1% | 59.8% | +2.7 (βˆ’3.8 to +8.7) | 0.63Γ— (0.55–0.71) |
43
+ | SWE-rebench, round 2 | 186 | 60.8% | 56.5% | βˆ’4.3 (βˆ’10.2 to +2.2) | 0.70Γ— (0.62–0.80) |
44
+ | Terminal-Bench Pro, round 1 | 98 | 60.6% | 60.2% | 0.0 (βˆ’9.2 to +9.2) | 0.63Γ— (0.52–0.77) |
45
+ | Terminal-Bench Pro, round 2 | 97 | 61.9% | 65.3% | +3.1 (βˆ’3.1 to +10.3) | 0.66Γ— (0.52–0.83) |
46
+ | Terminal-Bench 2.0 (40 tasks, 3 attempts each) | 108 | 75.9% | 70.0% | βˆ’4.6 (βˆ’12.1 to +3.7) | 0.96Γ— (0.78–1.16) |
47
+ | SkillsBench | 42 | 28.6% | 31.0% | +2.4 (βˆ’11.9 to +16.7) | 0.91Γ— (0.68–1.20) |
48
+ | Harbor Index | 41 | 12.2% | 9.8% | βˆ’2.4 (βˆ’12.2 to +7.3) | 0.71Γ— (0.51–0.99) |
49
+ | **All six, pooled** | **1,242** | **62.8%** | **62.4%** | **βˆ’0.2 (βˆ’2.6 to +2.1)** | **0.68Γ— (0.64–0.72)** |
50
+
51
+ - No benchmark shows a clear pass-rate difference: every interval includes zero. Time per task is the geometric mean of the per-task time ratios (Jeff-Code's time divided by Qwen alone's); below 1 is faster. The pass rates count every finished session; the difference counts only tasks finished in both settings, so it is not exactly the gap between the two pass rates.
52
+ - Total time drops less than time per task: in about 5% of tasks, Jeff-Code runs more than 30 minutes longer than Qwen alone, because it keeps going where Qwen alone gives up after a few minutes. The good news is, sometimes that pays off: in those tasks Jeff-Code solved 26 to Qwen's 24.
53
+ - Thinking off on every turn (with the same safeguards and thinking limit) is faster still but clearly worse: βˆ’7.6 points (βˆ’10.6 to βˆ’4.5), βˆ’13.5 on Terminal-Bench 2.0. A less cautious router threshold (0.7) loses 4.5 points. Jeff's decisions are what keep the quality.
54
+ - Terminal-Bench (original) and Terminal-Bench Science also ran, but both settings solve 0% of their tasks, so they are left out. 27 task pairs that hit an infrastructure failure twice are left out for both sides; about 10 long re-runs were still running when these numbers were taken.
55
+
56
+ The agent: [the Jeff-Code repository](https://github.com/firelex/jeff-code) (github.com/firelex/jeff-code). Offline accuracy reported by the Jeff-Code session: practically the same as the full fine-tune compared offline (step 92%, router 68% for both). Numbers from the evaluation report of 2026-10-05 18:23. Source: `results/sources/v1.3/jeff-code-owner-supplied.json` in the JeffHub repository (supplied by the maintainers).
57
+
58
+ ## Jeff-Code: the whole story
59
+
60
+ **What it is.** [Jeff-Code](https://github.com/firelex/jeff-code) is a coding agent based on [Pi](https://pi.dev), with two Jeff v1.3 adapters trained specifically for Qwen3.8-27B: this one and its partner ([`code`](https://huggingface.co/mstrasser/jeff-adapter-code) for steps, [`code-router`](https://huggingface.co/mstrasser/jeff-adapter-code-router) for thinking). We forked Pi because its extension framework doesn't currently let a fast decision model sit deep enough inside the agent loop. Besides being useful to people who run Qwen3.8-27B locally as their daily coding model, it is an experiment: how far can a small, fast "System 1" model go inside a coding agent?
61
+
62
+ **How it works.** Jeff makes two kinds of decisions around every Qwen turn, each by a small adapter on the same Jeff base, in about 0.2 s each:
63
+
64
+ 1. **Jeff works ahead of Qwen** (`code`). If it can take the next information-gathering step itself (read a file, list a folder, search the code, check which tools are installed), it does. It picks the tool first and then its argument, and can take several steps in a row while it is confident. Qwen then starts its turn with those results already in front of it, rather than spending a slow turn fetching them. Whenever Jeff is unsure, or the next step would change something (writing, editing, running or installing), it hands over to Qwen. So Jeff does not only pick a tool; it also fills in the tool's argument.
65
+ 2. **Jeff decides whether Qwen needs to think hard on its next turn** (`code-router`). Thinking stays off unless Jeff is confident the turn needs it. In the evaluation, about three quarters of Qwen's turns ran with thinking off.
66
+
67
+ Because Jeff returns a calibrated probability for every choice, each behaviour is controlled by a single setting: the step threshold (0.40) and the thinking threshold (0.6).
68
+
69
+ **How we trained it: Jeff predicts what Qwen would do next.**
70
+
71
+ - *Steps:* each training label is simply the step Qwen actually took next. If Jeff can take that step early, Qwen gets the result without spending a turn on it. These labels are built from Qwen sessions by code, with no other model involved.
72
+ - *Thinking:* each recorded Qwen turn at full thinking was asked again with thinking off, then low, then medium. The label is the cheapest level whose action was as good as the original, or full thinking if none was. "As good" is decided by code wherever possible (the same kind of step on the same target); otherwise Qwen3.8-Max, with thinking off, judges whether the cheaper step would serve the task just as well at that moment. About 27% of turns didn't need thinking at all.
73
+ - These labels are deliberately strict, which made the router cautious, and that is what preserved the pass rate. We also tried the looser question, "Is the full-thinking step materially better?". On a single turn the judge can't reliably tell thinking-off from a second full-thinking answer either, so a router trained that way would switch thinking off almost everywhere, and the thinking-off run shows what that costs over a whole task.
74
+
75
+ **Adapters, not a fine-tune.** Everything measured above used adapters: small LoRA files on the fixed Jeff v1.3 base. We also trained one full fine-tune to make both decisions and compared it offline; its accuracy was practically identical (step 92%, router 68% for both). So we release the adapters, which keep the multi-adapter design without giving up accuracy.
76
+
77
+ **What changed from Pi.**
78
+
79
+ - *Thinking per turn:* Pi only switches thinking on or off, so Qwen always thought at its highest level (in our tests, Qwen's low and medium levels thought about as long as the highest one, so they saved no time). Jeff-Code sets Qwen's thinking level for each turn, fixed or decided by Jeff.
80
+ - *Safeguards for thinking off:* a loop guard catches repeated or near-identical actions up to six steps back (a file write only counts as progress if it changes the file); a caught repeat is thrown away and that turn is asked again with full thinking, and near-identical outputs or two failed commands in a row send the next turn to full thinking. The thinking-off comparison had exactly the same safeguards, the same thinking limit and the same escalations (about 3% of its turns ended up thinking); the only difference from Jeff-Code is Jeff's decisions.
81
+ - *Runaway cut-off:* if Qwen's thinking or text keeps repeating itself, the reply is stopped and asked again with full thinking. This was on in every run, including the baseline (it caught 2 replies there).
82
+ - *Thinking limit:* at 8,000 thinking tokens, Qwen answers from what it has thought so far. That also rescues replies that would otherwise hit the 32K output limit, which ends a Pi session. The baseline ran without it, as plain Pi does.
83
+ - *Jeff steps:* before each Qwen turn, Jeff-Code builds a menu of concrete next steps from what is already known, and Jeff takes them when it is confident.
84
+ - *A pool of Jeff servers*, one per GPU, keeps each decision at about 0.2 s, and every Qwen request and Jeff decision is logged.
85
 
86
  ### With llama.cpp (GGUF)
87
 
 
104
 
105
  - You use another coding agent or another large model. The adapter imitates Qwen3.8-27B's own next steps in a bash-only agent.
106
  - You want Jeff to write, edit, run or install anything. It only takes information-gathering steps and hands over before anything changes.
 
107
 
108
  ## How to use it
109
 
 
196
  | Prompt layout | live-last: the fixed part of the request first, the changing state field last |
197
  | LoRA GGUF for llama.cpp | [mstrasser/jeff-adapter-code-gguf](https://huggingface.co/mstrasser/jeff-adapter-code-gguf) |
198
 
199
+ - **1.3.0** (2026-10-03): First release, trained on Jeff v1.3 with the live-last prompt layout. Used in Jeff-Code's measured runs with the step threshold at 0.40.
200
 
201
  ## Data card
202
 
 
205
  - Test set: not attached yet
206
  - QA report: not available here yet; it will be added once sanitised
207
 
208
+ **How the test set was held out.** Whole tasks are held out: Jeff-Code was measured only on tasks never used for training. SWE-bench Verified ran in full (all 500 tasks; none of its repositories were used for training). For the benchmarks also used in training, the tasks were split and every held-out task was run: Terminal-Bench 2.0 (40 frozen tasks; four near-twins of them were also kept out of training), SWE-rebench (189), Terminal-Bench Pro (100), SkillsBench (44) and Harbor Index (41). A leak check compares every training task with the evaluation tasks.
209
 
210
  **Training data.** Training data not published.
211