Ariful1904129 commited on
Commit
669a784
·
verified ·
1 Parent(s): 0337b7f

Model card: replace results table with training configurations

Browse files
Files changed (1) hide show
  1. README.md +35 -50
README.md CHANGED
@@ -17,8 +17,8 @@ tags:
17
  Classifies a Java/Kotlin test method into one of six categories: five kinds of flaky
18
  test plus non-flaky.
19
 
20
- **This is a partial, single-fold reproduction of published work — not the authors' model
21
- and not a match for their reported results.** Please read the limitations before using it.
22
 
23
  ## What this is
24
 
@@ -27,48 +27,31 @@ A fine-tune of `microsoft/codebert-base` on the FlakeBench dataset from
27
  (OOPSLA 2025), trained as a reproduction exercise on a single 8 GB consumer GPU.
28
 
29
  Trained on **one** of the paper's four project-disjoint folds (project group 2), so it is
30
- **not** comparable to the paper's four-fold headline number and should not be described
31
- as reproducing it.
32
-
33
- ## Results
34
-
35
- Held-out test set: 2,181 tests from 25 projects unseen during training.
36
-
37
- | Category | This model | Paper (4-fold) |
38
- |---|---|---|
39
- | Async Wait | 55.38% | 58.37% |
40
- | Concurrency | 10.53% | 35.92% |
41
- | Time | 57.14% | 72.73% |
42
- | Unordered Collections | 54.55% | 73.63% |
43
- | Test Order Dependency | **77.11%** | 64.35% |
44
- | Non-flaky | 100.00% | 100.00% |
45
- | **Macro F1** | **59.12%** | **65.79%** |
46
-
47
- Macro-F1 is **6.67 points below** the paper. Order-Dependency is the one category that
48
- exceeds it. Concurrency is weak (10.53%): 12 of 17 Concurrency tests are misclassified as
49
- Async Wait — the two categories share `Thread`/`await` vocabulary and the model does not
50
- separate them.
51
-
52
- Against a faithful baseline trained under the same compute budget without rebalancing
53
- (45.59%), this model is **+13.53 points** better — bootstrap 95% CI [+4.48, +21.23],
54
- p = 0.999, 2,000 resamples.
55
-
56
- ## Training
57
-
58
- | | |
59
- |---|---|
60
- | Base | `microsoft/codebert-base` |
61
- | Head | Linear(768→512) → ReLU → Dropout(0.3) → Linear(512→6) |
62
- | Loss | Focal loss, γ=2.0, balanced class weights |
63
- | Optimiser | AdamW, lr 1e-5, weight decay 0.01 |
64
- | Batch / max length | 8 / 512 |
65
- | Precision | fp16 |
66
- | Epochs | 18 (best at 6, selected on validation macro-F1) |
67
- | Rebalancing | non-flaky undersampled 4,972→800; each flaky class duplicated to 160 |
68
- | Hardware | 1× RTX 4060 Laptop (8 GB) |
69
-
70
- Class rebalancing is the one deviation from the paper's method, which trains on the raw
71
- distribution (97% non-flaky). It is what produced the gain over the baseline.
72
 
73
  ## Usage
74
 
@@ -100,14 +83,16 @@ Labels: `0` async wait, `1` concurrency, `2` time, `3` unordered collections,
100
 
101
  ## Limitations
102
 
103
- - **One fold, not four.** No evidence it generalises to folds 1, 3 or 4.
104
- - **Below the published result.** 59.12% vs 65.79%.
105
- - **Concurrency is unreliable** (10.53% F1) — it mostly predicts Async Wait instead.
106
- - **Noise floor is wide.** The test fold has 2,181 tests but only 103 flaky ones, so
 
 
107
  per-category F1 carries roughly ±6 points of uncertainty. Treat small differences as
108
  meaningless.
109
- - **Non-flaky dominates.** 95.3% of the test set is non-flaky and scores 100%, so overall
110
- accuracy (97%) is not informative — macro-F1 is the metric that matters.
111
  - Java/Kotlin test methods only; inputs longer than 512 tokens are truncated.
112
 
113
  ## Citation
 
17
  Classifies a Java/Kotlin test method into one of six categories: five kinds of flaky
18
  test plus non-flaky.
19
 
20
+ **This is a partial, single-fold reproduction of published work — not the authors' model.
21
+ It scores below the published results.** Please read the limitations before using it.
22
 
23
  ## What this is
24
 
 
27
  (OOPSLA 2025), trained as a reproduction exercise on a single 8 GB consumer GPU.
28
 
29
  Trained on **one** of the paper's four project-disjoint folds (project group 2), so it is
30
+ **not** comparable to the paper's four-fold results and should not be described as
31
+ reproducing them.
32
+
33
+ The uploaded weights are the "Balanced" configuration below.
34
+
35
+ ## Training configurations
36
+
37
+ | Parameter | Baseline | lr 2e-5 | Balanced | Augmented | Paper |
38
+ | --- | --- | --- | --- | --- | --- |
39
+ | Encoder | codebert-base | codebert-base | codebert-base | codebert-base | codebert-base |
40
+ | Learning rate | 1e-5 | 2e-5 | 1e-5 | 1e-5 | 1e-5 |
41
+ | Batch size | 8 | 8 | 8 | 8 | 8 |
42
+ | Max length | 512 | 512 | 512 | 512 | 512 |
43
+ | Loss | focal γ=2.0 | focal γ=2.0 | focal γ=2.0 | focal γ=2.0 | focal γ=2.0 |
44
+ | Class weights | balanced | balanced | balanced | balanced | balanced |
45
+ | Optimizer | AdamW wd 0.01 | AdamW wd 0.01 | AdamW wd 0.01 | AdamW wd 0.01 | AdamW wd 0.01 |
46
+ | Precision | fp16 | fp16 | fp16 | fp16 | fp32 |
47
+ | Non-flaky rows | 4,972 | 4,972 | 800 | 800 | full |
48
+ | Minority handling | none | none | ×160 copies | ×200 variants | none |
49
+ | Train rows | 5,114 | 5,114 | 1,600 | 1,800 | 5,114 |
50
+ | Epochs run | 8 | 8 | 18 | 13 | 40 |
51
+ | Dynamic padding | no | no | no | no | no |
52
+
53
+ Hardware: 1× RTX 4060 Laptop (8 GB). Class rebalancing is the one deviation from the
54
+ paper's method, which trains on the raw distribution (97% non-flaky).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
55
 
56
  ## Usage
57
 
 
83
 
84
  ## Limitations
85
 
86
+ - **One fold, not four.** No evidence it generalises to folds 1, 3 or 4. Evaluated across
87
+ all four folds, this configuration averages several points lower than on fold 2 alone.
88
+ - **Macro-F1 is below the published result** on every configuration tried.
89
+ - **Concurrency is unreliable.** Concurrency tests are frequently misclassified as
90
+ Async Wait; the two categories share `Thread`/`await` vocabulary.
91
+ - **Wide noise floor.** The test fold has 2,181 tests but only 103 flaky ones, so
92
  per-category F1 carries roughly ±6 points of uncertainty. Treat small differences as
93
  meaningless.
94
+ - **Non-flaky dominates.** 95.3% of the test set is non-flaky and is classified almost
95
+ perfectly, so overall accuracy is not informative — macro-F1 is the metric that matters.
96
  - Java/Kotlin test methods only; inputs longer than 512 tokens are truncated.
97
 
98
  ## Citation