File size: 897 Bytes
17b1b6d
7fd448e
 
17b1b6d
 
7fd448e
17b1b6d
7fd448e
17b1b6d
7fd448e
 
17b1b6d
7fd448e
17b1b6d
7fd448e
 
 
 
17b1b6d
7fd448e
 
17b1b6d
7fd448e
17b1b6d
7fd448e
17b1b6d
7fd448e
17b1b6d
7fd448e
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
---
license: other
base_model: Qwen/Qwen3.5-35B-A3B
---

# diffrecon-rl-iter34

Co-evolve **iter-1 RL** checkpoint **iter_34** (Qwen3.5-35B-A3B).

- Run: iter-1 RL (GRPO + dynamic sampling), harness = `diff-reconcile`, init = vanilla Qwen3.5-35B-A3B
- This is checkpoint **34 of 50** from that run.

## Held-out SWE-bench Verified (/500, temp/top_p 0.95, 3-pass)

| harness | score | n |
|---|---|---|
| diff-reconcile | 67.9 +/- 1.5 | 5 |
| combo_fb | 68.1 +/- 1.7 | 3 |

For reference, iter_49 from the same run: 68.2 +/- 0.8 (diff-reconcile, n=6),
68.8 +/- 1.2 (combo_fb, n=10 pooled). The checkpoints are statistically tied on SWE-V.

## In-training eval at this step

swe_val 65.3% (128/196), swe_multi 63.4% (single-pass proxy sets).

## Not to be confused with

`sweagent/rl-combo-iter34` -- that is **combo2's** iter_34, a different run.
This repo is the **iter-1 diff-reconcile** RL run.