Model card: replace results table with training configurations
Browse files
README.md
CHANGED
|
@@ -17,8 +17,8 @@ tags:
|
|
| 17 |
Classifies a Java/Kotlin test method into one of six categories: five kinds of flaky
|
| 18 |
test plus non-flaky.
|
| 19 |
|
| 20 |
-
**This is a partial, single-fold reproduction of published work — not the authors' model
|
| 21 |
-
|
| 22 |
|
| 23 |
## What this is
|
| 24 |
|
|
@@ -27,48 +27,31 @@ A fine-tune of `microsoft/codebert-base` on the FlakeBench dataset from
|
|
| 27 |
(OOPSLA 2025), trained as a reproduction exercise on a single 8 GB consumer GPU.
|
| 28 |
|
| 29 |
Trained on **one** of the paper's four project-disjoint folds (project group 2), so it is
|
| 30 |
-
**not** comparable to the paper's four-fold
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
|
| 38 |
-
|---|---|---|
|
| 39 |
-
|
|
| 40 |
-
|
|
| 41 |
-
|
|
| 42 |
-
|
|
| 43 |
-
|
|
| 44 |
-
|
|
| 45 |
-
|
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
## Training
|
| 57 |
-
|
| 58 |
-
| | |
|
| 59 |
-
|---|---|
|
| 60 |
-
| Base | `microsoft/codebert-base` |
|
| 61 |
-
| Head | Linear(768→512) → ReLU → Dropout(0.3) → Linear(512→6) |
|
| 62 |
-
| Loss | Focal loss, γ=2.0, balanced class weights |
|
| 63 |
-
| Optimiser | AdamW, lr 1e-5, weight decay 0.01 |
|
| 64 |
-
| Batch / max length | 8 / 512 |
|
| 65 |
-
| Precision | fp16 |
|
| 66 |
-
| Epochs | 18 (best at 6, selected on validation macro-F1) |
|
| 67 |
-
| Rebalancing | non-flaky undersampled 4,972→800; each flaky class duplicated to 160 |
|
| 68 |
-
| Hardware | 1× RTX 4060 Laptop (8 GB) |
|
| 69 |
-
|
| 70 |
-
Class rebalancing is the one deviation from the paper's method, which trains on the raw
|
| 71 |
-
distribution (97% non-flaky). It is what produced the gain over the baseline.
|
| 72 |
|
| 73 |
## Usage
|
| 74 |
|
|
@@ -100,14 +83,16 @@ Labels: `0` async wait, `1` concurrency, `2` time, `3` unordered collections,
|
|
| 100 |
|
| 101 |
## Limitations
|
| 102 |
|
| 103 |
-
- **One fold, not four.** No evidence it generalises to folds 1, 3 or 4.
|
| 104 |
-
|
| 105 |
-
- **
|
| 106 |
-
- **
|
|
|
|
|
|
|
| 107 |
per-category F1 carries roughly ±6 points of uncertainty. Treat small differences as
|
| 108 |
meaningless.
|
| 109 |
-
- **Non-flaky dominates.** 95.3% of the test set is non-flaky and
|
| 110 |
-
|
| 111 |
- Java/Kotlin test methods only; inputs longer than 512 tokens are truncated.
|
| 112 |
|
| 113 |
## Citation
|
|
|
|
| 17 |
Classifies a Java/Kotlin test method into one of six categories: five kinds of flaky
|
| 18 |
test plus non-flaky.
|
| 19 |
|
| 20 |
+
**This is a partial, single-fold reproduction of published work — not the authors' model.
|
| 21 |
+
It scores below the published results.** Please read the limitations before using it.
|
| 22 |
|
| 23 |
## What this is
|
| 24 |
|
|
|
|
| 27 |
(OOPSLA 2025), trained as a reproduction exercise on a single 8 GB consumer GPU.
|
| 28 |
|
| 29 |
Trained on **one** of the paper's four project-disjoint folds (project group 2), so it is
|
| 30 |
+
**not** comparable to the paper's four-fold results and should not be described as
|
| 31 |
+
reproducing them.
|
| 32 |
+
|
| 33 |
+
The uploaded weights are the "Balanced" configuration below.
|
| 34 |
+
|
| 35 |
+
## Training configurations
|
| 36 |
+
|
| 37 |
+
| Parameter | Baseline | lr 2e-5 | Balanced | Augmented | Paper |
|
| 38 |
+
| --- | --- | --- | --- | --- | --- |
|
| 39 |
+
| Encoder | codebert-base | codebert-base | codebert-base | codebert-base | codebert-base |
|
| 40 |
+
| Learning rate | 1e-5 | 2e-5 | 1e-5 | 1e-5 | 1e-5 |
|
| 41 |
+
| Batch size | 8 | 8 | 8 | 8 | 8 |
|
| 42 |
+
| Max length | 512 | 512 | 512 | 512 | 512 |
|
| 43 |
+
| Loss | focal γ=2.0 | focal γ=2.0 | focal γ=2.0 | focal γ=2.0 | focal γ=2.0 |
|
| 44 |
+
| Class weights | balanced | balanced | balanced | balanced | balanced |
|
| 45 |
+
| Optimizer | AdamW wd 0.01 | AdamW wd 0.01 | AdamW wd 0.01 | AdamW wd 0.01 | AdamW wd 0.01 |
|
| 46 |
+
| Precision | fp16 | fp16 | fp16 | fp16 | fp32 |
|
| 47 |
+
| Non-flaky rows | 4,972 | 4,972 | 800 | 800 | full |
|
| 48 |
+
| Minority handling | none | none | ×160 copies | ×200 variants | none |
|
| 49 |
+
| Train rows | 5,114 | 5,114 | 1,600 | 1,800 | 5,114 |
|
| 50 |
+
| Epochs run | 8 | 8 | 18 | 13 | 40 |
|
| 51 |
+
| Dynamic padding | no | no | no | no | no |
|
| 52 |
+
|
| 53 |
+
Hardware: 1× RTX 4060 Laptop (8 GB). Class rebalancing is the one deviation from the
|
| 54 |
+
paper's method, which trains on the raw distribution (97% non-flaky).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 55 |
|
| 56 |
## Usage
|
| 57 |
|
|
|
|
| 83 |
|
| 84 |
## Limitations
|
| 85 |
|
| 86 |
+
- **One fold, not four.** No evidence it generalises to folds 1, 3 or 4. Evaluated across
|
| 87 |
+
all four folds, this configuration averages several points lower than on fold 2 alone.
|
| 88 |
+
- **Macro-F1 is below the published result** on every configuration tried.
|
| 89 |
+
- **Concurrency is unreliable.** Concurrency tests are frequently misclassified as
|
| 90 |
+
Async Wait; the two categories share `Thread`/`await` vocabulary.
|
| 91 |
+
- **Wide noise floor.** The test fold has 2,181 tests but only 103 flaky ones, so
|
| 92 |
per-category F1 carries roughly ±6 points of uncertainty. Treat small differences as
|
| 93 |
meaningless.
|
| 94 |
+
- **Non-flaky dominates.** 95.3% of the test set is non-flaky and is classified almost
|
| 95 |
+
perfectly, so overall accuracy is not informative — macro-F1 is the metric that matters.
|
| 96 |
- Java/Kotlin test methods only; inputs longer than 512 tokens are truncated.
|
| 97 |
|
| 98 |
## Citation
|