Instructions to use mstrasser/jeff-adapter-code with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use mstrasser/jeff-adapter-code with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Card: benchmark rounds combined
Browse files
README.md
CHANGED
|
@@ -39,16 +39,14 @@ Jeff-Code runs two adapters on the fixed Jeff v1.3 base: [`code`](https://jeffhu
|
|
| 39 |
| Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
|
| 40 |
|---|---:|---:|---:|---|---:|
|
| 41 |
| SWE-bench Verified | 486 | 70.6% | 70.8% | +0.4 (β3.5 to +4.3) | 0.63Γ (0.57β0.69) |
|
| 42 |
-
| SWE-rebench
|
| 43 |
-
|
|
| 44 |
-
| Terminal-Bench Pro, round 1 | 98 | 60.6% | 60.2% | 0.0 (β9.2 to +9.2) | 0.63Γ (0.52β0.77) |
|
| 45 |
-
| Terminal-Bench Pro, round 2 | 97 | 61.9% | 65.3% | +3.1 (β3.1 to +10.3) | 0.66Γ (0.52β0.83) |
|
| 46 |
| Terminal-Bench 2.0 (40 tasks, 3 attempts each) | 108 | 75.9% | 70.0% | β4.6 (β12.1 to +3.7) | 0.96Γ (0.78β1.16) |
|
| 47 |
| SkillsBench | 42 | 28.6% | 31.0% | +2.4 (β11.9 to +16.7) | 0.91Γ (0.68β1.20) |
|
| 48 |
| Harbor Index | 41 | 12.2% | 9.8% | β2.4 (β12.2 to +7.3) | 0.71Γ (0.51β0.99) |
|
| 49 |
| **All six, pooled** | **1,242** | **62.8%** | **62.4%** | **β0.2 (β2.6 to +2.1)** | **0.68Γ (0.64β0.72)** |
|
| 50 |
|
| 51 |
-
- No benchmark shows a clear pass-rate difference: every interval includes zero. Time per task is the geometric mean of the per-task time ratios (Jeff-Code's time divided by Qwen alone's); below 1 is faster. The pass rates count every finished session; the difference counts only tasks finished in both settings, so it is not exactly the gap between the two pass rates.
|
| 52 |
- Total time drops less than time per task: in about 5% of tasks, Jeff-Code runs more than 30 minutes longer than Qwen alone, because it keeps going where Qwen alone gives up after a few minutes. The good news is, sometimes that pays off: in those tasks Jeff-Code solved 26 to Qwen's 24.
|
| 53 |
- Thinking off on every turn (with the same safeguards and thinking limit) is faster still but clearly worse: β7.6 points (β10.6 to β4.5), β13.5 on Terminal-Bench 2.0. A less cautious router threshold (0.7) loses 4.5 points. Jeff's decisions are what keep the quality.
|
| 54 |
- Terminal-Bench (original) and Terminal-Bench Science also ran, but both settings solve 0% of their tasks, so they are left out. 27 task pairs that hit an infrastructure failure twice are left out for both sides; about 10 long re-runs were still running when these numbers were taken.
|
|
|
|
| 39 |
| Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
|
| 40 |
|---|---:|---:|---:|---|---:|
|
| 41 |
| SWE-bench Verified | 486 | 70.6% | 70.8% | +0.4 (β3.5 to +4.3) | 0.63Γ (0.57β0.69) |
|
| 42 |
+
| SWE-rebench (2 rounds) | 370 | 58.9% | 58.1% | β0.8 (β5.7 to +3.5) | 0.66Γ (0.61β0.73) |
|
| 43 |
+
| Terminal-Bench Pro (2 rounds) | 195 | 61.2% | 62.8% | +1.5 (β4.6 to +7.7) | 0.64Γ (0.55β0.76) |
|
|
|
|
|
|
|
| 44 |
| Terminal-Bench 2.0 (40 tasks, 3 attempts each) | 108 | 75.9% | 70.0% | β4.6 (β12.1 to +3.7) | 0.96Γ (0.78β1.16) |
|
| 45 |
| SkillsBench | 42 | 28.6% | 31.0% | +2.4 (β11.9 to +16.7) | 0.91Γ (0.68β1.20) |
|
| 46 |
| Harbor Index | 41 | 12.2% | 9.8% | β2.4 (β12.2 to +7.3) | 0.71Γ (0.51β0.99) |
|
| 47 |
| **All six, pooled** | **1,242** | **62.8%** | **62.4%** | **β0.2 (β2.6 to +2.1)** | **0.68Γ (0.64β0.72)** |
|
| 48 |
|
| 49 |
+
- No benchmark shows a clear pass-rate difference: every interval includes zero. SWE-rebench and Terminal-Bench Pro ran twice; their two rounds are combined, with intervals computed over both. Time per task is the geometric mean of the per-task time ratios (Jeff-Code's time divided by Qwen alone's); below 1 is faster. The pass rates count every finished session; the difference counts only tasks finished in both settings, so it is not exactly the gap between the two pass rates.
|
| 50 |
- Total time drops less than time per task: in about 5% of tasks, Jeff-Code runs more than 30 minutes longer than Qwen alone, because it keeps going where Qwen alone gives up after a few minutes. The good news is, sometimes that pays off: in those tasks Jeff-Code solved 26 to Qwen's 24.
|
| 51 |
- Thinking off on every turn (with the same safeguards and thinking limit) is faster still but clearly worse: β7.6 points (β10.6 to β4.5), β13.5 on Terminal-Bench 2.0. A less cautious router threshold (0.7) loses 4.5 points. Jeff's decisions are what keep the quality.
|
| 52 |
- Terminal-Bench (original) and Terminal-Bench Science also ran, but both settings solve 0% of their tasks, so they are left out. 27 task pairs that hit an infrastructure failure twice are left out for both sides; about 10 long re-runs were still running when these numbers were taken.
|