Instructions to use mstrasser/jeff-adapter-code with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use mstrasser/jeff-adapter-code with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Card: Qwen3.5-0.8B licence notice for the v1.3 release
Browse files
README.md
CHANGED
|
@@ -34,7 +34,7 @@ Adapter page, with the full data card: [jeffhub.ai/adapters/code](https://jeffhu
|
|
| 34 |
- **14% less total time** over all tasks combined (0.86×, 0.80–0.93)
|
| 35 |
- **92% / 68%**: offline accuracy on held-out scoring tasks: code (steps) / code-router (thinking)
|
| 36 |
|
| 37 |
-
Jeff-Code runs two adapters on the fixed Jeff v1.3 base: [`code`](https://jeffhub.ai/adapters/code) takes the information-gathering steps (step threshold 0.40), and [`code-router`](https://jeffhub.ai/adapters/code-router) decides whether Qwen thinks hard on a turn (thinking off unless P(xhigh), out of the four levels off/low/medium/xhigh, ≥ 0.6). The baseline is Qwen3.8-27B alone in the same Jeff-Code build with every Jeff feature off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it. Each task ran in both settings side by side, at the same time on the same Qwen server, and every comparison is paired by task. Only tasks Jeff never saw in training were used.
|
| 38 |
|
| 39 |
| Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
|
| 40 |
|---|---:|---:|---:|---|---:|
|
|
@@ -59,7 +59,7 @@ The agent: [the Jeff-Code repository](https://github.com/firelex/jeff-code) (git
|
|
| 59 |
|
| 60 |
**How it works.** Jeff makes two kinds of decisions around every Qwen turn, each by a small adapter on the same Jeff base, in about 0.2 s each:
|
| 61 |
|
| 62 |
-
1. **Jeff works ahead of Qwen** (`code`). If it can take the next information-gathering step itself (read a file, list a folder, search the code, check which tools are installed), it does. It picks the tool first and then its argument, and can take several steps in a row while it is confident. Qwen then starts its turn with those results already in front of it, rather than spending a slow turn fetching them.
|
| 63 |
2. **Jeff decides whether Qwen needs to think hard on its next turn** (`code-router`). Thinking stays off unless Jeff is confident the turn needs it. In the evaluation, about three quarters of Qwen's turns ran with thinking off.
|
| 64 |
|
| 65 |
Because Jeff returns a calibrated probability for every choice, each behaviour is controlled by a single setting: the step threshold (0.40) and the thinking threshold (0.6).
|
|
@@ -101,7 +101,7 @@ Full precision: full precision from the trainer's own evaluation on the developm
|
|
| 101 |
## When not to use it
|
| 102 |
|
| 103 |
- You use another coding agent or another large model. The adapter imitates Qwen3.8-27B's own next steps in a bash-only agent.
|
| 104 |
-
- You want Jeff to write
|
| 105 |
|
| 106 |
## How to use it
|
| 107 |
|
|
@@ -231,6 +231,8 @@ Jeff-Code's measured results come entirely from LoRA adapters on the fixed Jeff
|
|
| 231 |
|
| 232 |
**Adapter licence: Apache-2.0.**
|
| 233 |
|
|
|
|
|
|
|
| 234 |
**To confirm:** that the licences of the session data and of Terminal-Bench 2.0 allow training and publishing the adapter.
|
| 235 |
|
| 236 |
It was trained on:
|
|
|
|
| 34 |
- **14% less total time** over all tasks combined (0.86×, 0.80–0.93)
|
| 35 |
- **92% / 68%**: offline accuracy on held-out scoring tasks: code (steps) / code-router (thinking)
|
| 36 |
|
| 37 |
+
Jeff-Code runs two adapters on the fixed Jeff v1.3 base: [`code`](https://jeffhub.ai/adapters/code) takes the information-gathering steps (step threshold 0.40), and [`code-router`](https://jeffhub.ai/adapters/code-router) decides whether Qwen thinks hard on a turn (thinking off unless P(xhigh), out of the four levels off/low/medium/xhigh, ≥ 0.6). The baseline is Qwen3.8-27B alone in the same Jeff-Code build with every Jeff feature off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it. Each task ran in both settings side by side, at the same time on the same Qwen server, and every comparison is paired by task. Only tasks Jeff never saw in training were used: the held-out splits of six benchmarks.
|
| 38 |
|
| 39 |
| Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
|
| 40 |
|---|---:|---:|---:|---|---:|
|
|
|
|
| 59 |
|
| 60 |
**How it works.** Jeff makes two kinds of decisions around every Qwen turn, each by a small adapter on the same Jeff base, in about 0.2 s each:
|
| 61 |
|
| 62 |
+
1. **Jeff works ahead of Qwen** (`code`). If it can take the next information-gathering step itself (read a file, list a folder, search the code, check which tools are installed), it does. It picks the tool first and then its argument, and can take several steps in a row while it is confident. Qwen then starts its turn with those results already in front of it, rather than spending a slow turn fetching them. It can also run the tests or a build, repeat Qwen's last command and, with the run-approval setting the evaluation used, run a script Qwen wrote or install a missing package. Writing and editing files always stay with Qwen, and whenever Jeff is unsure, it hands over to Qwen. So Jeff does not only pick a tool; it also fills in the tool's argument.
|
| 63 |
2. **Jeff decides whether Qwen needs to think hard on its next turn** (`code-router`). Thinking stays off unless Jeff is confident the turn needs it. In the evaluation, about three quarters of Qwen's turns ran with thinking off.
|
| 64 |
|
| 65 |
Because Jeff returns a calibrated probability for every choice, each behaviour is controlled by a single setting: the step threshold (0.40) and the thinking threshold (0.6).
|
|
|
|
| 101 |
## When not to use it
|
| 102 |
|
| 103 |
- You use another coding agent or another large model. The adapter imitates Qwen3.8-27B's own next steps in a bash-only agent.
|
| 104 |
+
- You want Jeff to write or edit files. Writing and editing always stay with Qwen; Jeff takes information steps, can run the tests or a build and repeat Qwen's last command, and, with the run-approval setting the evaluation used, can run a script Qwen wrote or install a missing package.
|
| 105 |
|
| 106 |
## How to use it
|
| 107 |
|
|
|
|
| 231 |
|
| 232 |
**Adapter licence: Apache-2.0.**
|
| 233 |
|
| 234 |
+
**Qwen3.5-0.8B notice:** these weights were modified from [Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B) by the Jeff project: [jeff-base](https://huggingface.co/mstrasser/jeff-base) is a fine-tune of Qwen3.5-0.8B, and this adapter was trained on top of it. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in [`LICENSE`](LICENSE).
|
| 235 |
+
|
| 236 |
**To confirm:** that the licences of the session data and of Terminal-Bench 2.0 allow training and publishing the adapter.
|
| 237 |
|
| 238 |
It was trained on:
|