mstrasser commited on
Commit
ef5e707
Β·
verified Β·
1 Parent(s): e6c013d

Card: benchmark rounds combined

Browse files
Files changed (1) hide show
  1. README.md +3 -5
README.md CHANGED
@@ -39,16 +39,14 @@ Jeff-Code runs two adapters on the fixed Jeff v1.3 base: [`code`](https://jeffhu
39
  | Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
40
  |---|---:|---:|---:|---|---:|
41
  | SWE-bench Verified | 486 | 70.6% | 70.8% | +0.4 (βˆ’3.5 to +4.3) | 0.63Γ— (0.57–0.69) |
42
- | SWE-rebench, round 1 | 184 | 57.1% | 59.8% | +2.7 (βˆ’3.8 to +8.7) | 0.63Γ— (0.55–0.71) |
43
- | SWE-rebench, round 2 | 186 | 60.8% | 56.5% | βˆ’4.3 (βˆ’10.2 to +2.2) | 0.70Γ— (0.62–0.80) |
44
- | Terminal-Bench Pro, round 1 | 98 | 60.6% | 60.2% | 0.0 (βˆ’9.2 to +9.2) | 0.63Γ— (0.52–0.77) |
45
- | Terminal-Bench Pro, round 2 | 97 | 61.9% | 65.3% | +3.1 (βˆ’3.1 to +10.3) | 0.66Γ— (0.52–0.83) |
46
  | Terminal-Bench 2.0 (40 tasks, 3 attempts each) | 108 | 75.9% | 70.0% | βˆ’4.6 (βˆ’12.1 to +3.7) | 0.96Γ— (0.78–1.16) |
47
  | SkillsBench | 42 | 28.6% | 31.0% | +2.4 (βˆ’11.9 to +16.7) | 0.91Γ— (0.68–1.20) |
48
  | Harbor Index | 41 | 12.2% | 9.8% | βˆ’2.4 (βˆ’12.2 to +7.3) | 0.71Γ— (0.51–0.99) |
49
  | **All six, pooled** | **1,242** | **62.8%** | **62.4%** | **βˆ’0.2 (βˆ’2.6 to +2.1)** | **0.68Γ— (0.64–0.72)** |
50
 
51
- - No benchmark shows a clear pass-rate difference: every interval includes zero. Time per task is the geometric mean of the per-task time ratios (Jeff-Code's time divided by Qwen alone's); below 1 is faster. The pass rates count every finished session; the difference counts only tasks finished in both settings, so it is not exactly the gap between the two pass rates.
52
  - Total time drops less than time per task: in about 5% of tasks, Jeff-Code runs more than 30 minutes longer than Qwen alone, because it keeps going where Qwen alone gives up after a few minutes. The good news is, sometimes that pays off: in those tasks Jeff-Code solved 26 to Qwen's 24.
53
  - Thinking off on every turn (with the same safeguards and thinking limit) is faster still but clearly worse: βˆ’7.6 points (βˆ’10.6 to βˆ’4.5), βˆ’13.5 on Terminal-Bench 2.0. A less cautious router threshold (0.7) loses 4.5 points. Jeff's decisions are what keep the quality.
54
  - Terminal-Bench (original) and Terminal-Bench Science also ran, but both settings solve 0% of their tasks, so they are left out. 27 task pairs that hit an infrastructure failure twice are left out for both sides; about 10 long re-runs were still running when these numbers were taken.
 
39
  | Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
40
  |---|---:|---:|---:|---|---:|
41
  | SWE-bench Verified | 486 | 70.6% | 70.8% | +0.4 (βˆ’3.5 to +4.3) | 0.63Γ— (0.57–0.69) |
42
+ | SWE-rebench (2 rounds) | 370 | 58.9% | 58.1% | βˆ’0.8 (βˆ’5.7 to +3.5) | 0.66Γ— (0.61–0.73) |
43
+ | Terminal-Bench Pro (2 rounds) | 195 | 61.2% | 62.8% | +1.5 (βˆ’4.6 to +7.7) | 0.64Γ— (0.55–0.76) |
 
 
44
  | Terminal-Bench 2.0 (40 tasks, 3 attempts each) | 108 | 75.9% | 70.0% | βˆ’4.6 (βˆ’12.1 to +3.7) | 0.96Γ— (0.78–1.16) |
45
  | SkillsBench | 42 | 28.6% | 31.0% | +2.4 (βˆ’11.9 to +16.7) | 0.91Γ— (0.68–1.20) |
46
  | Harbor Index | 41 | 12.2% | 9.8% | βˆ’2.4 (βˆ’12.2 to +7.3) | 0.71Γ— (0.51–0.99) |
47
  | **All six, pooled** | **1,242** | **62.8%** | **62.4%** | **βˆ’0.2 (βˆ’2.6 to +2.1)** | **0.68Γ— (0.64–0.72)** |
48
 
49
+ - No benchmark shows a clear pass-rate difference: every interval includes zero. SWE-rebench and Terminal-Bench Pro ran twice; their two rounds are combined, with intervals computed over both. Time per task is the geometric mean of the per-task time ratios (Jeff-Code's time divided by Qwen alone's); below 1 is faster. The pass rates count every finished session; the difference counts only tasks finished in both settings, so it is not exactly the gap between the two pass rates.
50
  - Total time drops less than time per task: in about 5% of tasks, Jeff-Code runs more than 30 minutes longer than Qwen alone, because it keeps going where Qwen alone gives up after a few minutes. The good news is, sometimes that pays off: in those tasks Jeff-Code solved 26 to Qwen's 24.
51
  - Thinking off on every turn (with the same safeguards and thinking limit) is faster still but clearly worse: βˆ’7.6 points (βˆ’10.6 to βˆ’4.5), βˆ’13.5 on Terminal-Bench 2.0. A less cautious router threshold (0.7) loses 4.5 points. Jeff's decisions are what keep the quality.
52
  - Terminal-Bench (original) and Terminal-Bench Science also ran, but both settings solve 0% of their tasks, so they are left out. 27 task pairs that hit an infrastructure failure twice are left out for both sides; about 10 long re-runs were still running when these numbers were taken.