mstrasser commited on
Commit
cf80fd3
·
verified ·
1 Parent(s): 3eb1cef

Card: link the code adapters to their Hugging Face cards

Browse files
Files changed (1) hide show
  1. README.md +1 -4
README.md CHANGED
@@ -25,8 +25,6 @@ A LoRA adapter for [jeff-base](https://huggingface.co/mstrasser/jeff-base) **v1.
25
  calibrated probability for every option from one forward pass, with no generated text to parse. One Jeff server loads
26
  the base once and any number of adapters beside it; each request picks an adapter by name (`"model": "code"`).
27
 
28
- Adapter page, with the full data card: [jeffhub.ai/adapters/code](https://jeffhub.ai/adapters/code).
29
-
30
  ## Results
31
 
32
  - **62.4% vs 62.8%**: pass rate of Jeff-Code vs Qwen3.8-27B alone; paired difference −0.2 points (95% interval −2.6 to +2.1), 1,242 paired tasks on 6 benchmarks
@@ -34,7 +32,7 @@ Adapter page, with the full data card: [jeffhub.ai/adapters/code](https://jeffhu
34
  - **14% less total time** over all tasks combined (0.86×, 0.80–0.93)
35
  - **92% / 68%**: offline accuracy on held-out scoring tasks: code (steps) / code-router (thinking)
36
 
37
- Jeff-Code runs two adapters on the fixed Jeff v1.3 base: [`code`](https://jeffhub.ai/adapters/code) takes the information-gathering steps (step threshold 0.40), and [`code-router`](https://jeffhub.ai/adapters/code-router) decides whether Qwen thinks hard on a turn (thinking off unless P(xhigh), out of the four levels off/low/medium/xhigh, ≥ 0.6). The baseline is Qwen3.8-27B alone in the same Jeff-Code build with every Jeff feature off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it. Each task ran in both settings side by side, at the same time on the same Qwen server, and every comparison is paired by task. Only tasks Jeff never saw in training were used: the held-out splits of six benchmarks.
38
 
39
  | Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
40
  |---|---:|---:|---:|---|---:|
@@ -268,7 +266,6 @@ It was trained on:
268
 
269
  ## Links
270
 
271
- - Adapter page: [jeffhub.ai/adapters/code](https://jeffhub.ai/adapters/code)
272
  - Base model: [mstrasser/jeff-base](https://huggingface.co/mstrasser/jeff-base) (revision v1.3)
273
  - LoRA GGUF for llama.cpp: [mstrasser/jeff-adapter-code-gguf](https://huggingface.co/mstrasser/jeff-adapter-code-gguf)
274
  - What changed in v1.3: [release notes](https://jeffhub.ai/docs/release-notes-v1-3)
 
25
  calibrated probability for every option from one forward pass, with no generated text to parse. One Jeff server loads
26
  the base once and any number of adapters beside it; each request picks an adapter by name (`"model": "code"`).
27
 
 
 
28
  ## Results
29
 
30
  - **62.4% vs 62.8%**: pass rate of Jeff-Code vs Qwen3.8-27B alone; paired difference −0.2 points (95% interval −2.6 to +2.1), 1,242 paired tasks on 6 benchmarks
 
32
  - **14% less total time** over all tasks combined (0.86×, 0.80–0.93)
33
  - **92% / 68%**: offline accuracy on held-out scoring tasks: code (steps) / code-router (thinking)
34
 
35
+ Jeff-Code runs two adapters on the fixed Jeff v1.3 base: [`code`](https://huggingface.co/mstrasser/jeff-adapter-code) takes the information-gathering steps (step threshold 0.40), and [`code-router`](https://huggingface.co/mstrasser/jeff-adapter-code-router) decides whether Qwen thinks hard on a turn (thinking off unless P(xhigh), out of the four levels off/low/medium/xhigh, ≥ 0.6). The baseline is Qwen3.8-27B alone in the same Jeff-Code build with every Jeff feature off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it. Each task ran in both settings side by side, at the same time on the same Qwen server, and every comparison is paired by task. Only tasks Jeff never saw in training were used: the held-out splits of six benchmarks.
36
 
37
  | Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
38
  |---|---:|---:|---:|---|---:|
 
266
 
267
  ## Links
268
 
 
269
  - Base model: [mstrasser/jeff-base](https://huggingface.co/mstrasser/jeff-base) (revision v1.3)
270
  - LoRA GGUF for llama.cpp: [mstrasser/jeff-adapter-code-gguf](https://huggingface.co/mstrasser/jeff-adapter-code-gguf)
271
  - What changed in v1.3: [release notes](https://jeffhub.ai/docs/release-notes-v1-3)