mstrasser commited on
Commit
9655922
·
verified ·
1 Parent(s): ef5e707

Card: Qwen3.5-0.8B licence notice for the v1.3 release

Browse files
Files changed (1) hide show
  1. README.md +5 -3
README.md CHANGED
@@ -34,7 +34,7 @@ Adapter page, with the full data card: [jeffhub.ai/adapters/code](https://jeffhu
34
  - **14% less total time** over all tasks combined (0.86×, 0.80–0.93)
35
  - **92% / 68%**: offline accuracy on held-out scoring tasks: code (steps) / code-router (thinking)
36
 
37
- Jeff-Code runs two adapters on the fixed Jeff v1.3 base: [`code`](https://jeffhub.ai/adapters/code) takes the information-gathering steps (step threshold 0.40), and [`code-router`](https://jeffhub.ai/adapters/code-router) decides whether Qwen thinks hard on a turn (thinking off unless P(xhigh), out of the four levels off/low/medium/xhigh, ≥ 0.6). The baseline is Qwen3.8-27B alone in the same Jeff-Code build with every Jeff feature off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it. Each task ran in both settings side by side, at the same time on the same Qwen server, and every comparison is paired by task. Only tasks Jeff never saw in training were used.
38
 
39
  | Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
40
  |---|---:|---:|---:|---|---:|
@@ -59,7 +59,7 @@ The agent: [the Jeff-Code repository](https://github.com/firelex/jeff-code) (git
59
 
60
  **How it works.** Jeff makes two kinds of decisions around every Qwen turn, each by a small adapter on the same Jeff base, in about 0.2 s each:
61
 
62
- 1. **Jeff works ahead of Qwen** (`code`). If it can take the next information-gathering step itself (read a file, list a folder, search the code, check which tools are installed), it does. It picks the tool first and then its argument, and can take several steps in a row while it is confident. Qwen then starts its turn with those results already in front of it, rather than spending a slow turn fetching them. Whenever Jeff is unsure, or the next step would change something (writing, editing, running or installing), it hands over to Qwen. So Jeff does not only pick a tool; it also fills in the tool's argument.
63
  2. **Jeff decides whether Qwen needs to think hard on its next turn** (`code-router`). Thinking stays off unless Jeff is confident the turn needs it. In the evaluation, about three quarters of Qwen's turns ran with thinking off.
64
 
65
  Because Jeff returns a calibrated probability for every choice, each behaviour is controlled by a single setting: the step threshold (0.40) and the thinking threshold (0.6).
@@ -101,7 +101,7 @@ Full precision: full precision from the trainer's own evaluation on the developm
101
  ## When not to use it
102
 
103
  - You use another coding agent or another large model. The adapter imitates Qwen3.8-27B's own next steps in a bash-only agent.
104
- - You want Jeff to write, edit, run or install anything. It only takes information-gathering steps and hands over before anything changes.
105
 
106
  ## How to use it
107
 
@@ -231,6 +231,8 @@ Jeff-Code's measured results come entirely from LoRA adapters on the fixed Jeff
231
 
232
  **Adapter licence: Apache-2.0.**
233
 
 
 
234
  **To confirm:** that the licences of the session data and of Terminal-Bench 2.0 allow training and publishing the adapter.
235
 
236
  It was trained on:
 
34
  - **14% less total time** over all tasks combined (0.86×, 0.80–0.93)
35
  - **92% / 68%**: offline accuracy on held-out scoring tasks: code (steps) / code-router (thinking)
36
 
37
+ Jeff-Code runs two adapters on the fixed Jeff v1.3 base: [`code`](https://jeffhub.ai/adapters/code) takes the information-gathering steps (step threshold 0.40), and [`code-router`](https://jeffhub.ai/adapters/code-router) decides whether Qwen thinks hard on a turn (thinking off unless P(xhigh), out of the four levels off/low/medium/xhigh, ≥ 0.6). The baseline is Qwen3.8-27B alone in the same Jeff-Code build with every Jeff feature off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it. Each task ran in both settings side by side, at the same time on the same Qwen server, and every comparison is paired by task. Only tasks Jeff never saw in training were used: the held-out splits of six benchmarks.
38
 
39
  | Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
40
  |---|---:|---:|---:|---|---:|
 
59
 
60
  **How it works.** Jeff makes two kinds of decisions around every Qwen turn, each by a small adapter on the same Jeff base, in about 0.2 s each:
61
 
62
+ 1. **Jeff works ahead of Qwen** (`code`). If it can take the next information-gathering step itself (read a file, list a folder, search the code, check which tools are installed), it does. It picks the tool first and then its argument, and can take several steps in a row while it is confident. Qwen then starts its turn with those results already in front of it, rather than spending a slow turn fetching them. It can also run the tests or a build, repeat Qwen's last command and, with the run-approval setting the evaluation used, run a script Qwen wrote or install a missing package. Writing and editing files always stay with Qwen, and whenever Jeff is unsure, it hands over to Qwen. So Jeff does not only pick a tool; it also fills in the tool's argument.
63
  2. **Jeff decides whether Qwen needs to think hard on its next turn** (`code-router`). Thinking stays off unless Jeff is confident the turn needs it. In the evaluation, about three quarters of Qwen's turns ran with thinking off.
64
 
65
  Because Jeff returns a calibrated probability for every choice, each behaviour is controlled by a single setting: the step threshold (0.40) and the thinking threshold (0.6).
 
101
  ## When not to use it
102
 
103
  - You use another coding agent or another large model. The adapter imitates Qwen3.8-27B's own next steps in a bash-only agent.
104
+ - You want Jeff to write or edit files. Writing and editing always stay with Qwen; Jeff takes information steps, can run the tests or a build and repeat Qwen's last command, and, with the run-approval setting the evaluation used, can run a script Qwen wrote or install a missing package.
105
 
106
  ## How to use it
107
 
 
231
 
232
  **Adapter licence: Apache-2.0.**
233
 
234
+ **Qwen3.5-0.8B notice:** these weights were modified from [Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B) by the Jeff project: [jeff-base](https://huggingface.co/mstrasser/jeff-base) is a fine-tune of Qwen3.5-0.8B, and this adapter was trained on top of it. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in [`LICENSE`](LICENSE).
235
+
236
  **To confirm:** that the licences of the session data and of Terminal-Bench 2.0 allow training and publishing the adapter.
237
 
238
  It was trained on: