Update README.md
Browse files
README.md
CHANGED
|
@@ -22,7 +22,6 @@ test plus non-flaky.
|
|
| 22 |
A fine-tune of `microsoft/codebert-base` on the FlakeBench dataset from
|
| 23 |
[*Understanding and Improving Flaky Test Classification*](https://utexas.app.box.com/v/august-shi-OOPSLA2025)
|
| 24 |
(OOPSLA 2025), trained as a reproduction exercise on a single 8 GB consumer GPU.
|
| 25 |
-
|
| 26 |
The uploaded weights are the "Balanced" configuration below.
|
| 27 |
|
| 28 |
## Training configurations
|
|
@@ -46,6 +45,32 @@ The uploaded weights are the "Balanced" configuration below.
|
|
| 46 |
Hardware: 1× RTX 4060 Laptop (8 GB). Class rebalancing is the one deviation from the
|
| 47 |
paper's method, which trains on the raw distribution (97% non-flaky).
|
| 48 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 49 |
## Usage
|
| 50 |
|
| 51 |
```python
|
|
@@ -75,6 +100,9 @@ Labels: `0` async wait, `1` concurrency, `2` time, `3` unordered collections,
|
|
| 75 |
`4` test order dependency, `5` non-flaky.
|
| 76 |
|
| 77 |
Scope: Java/Kotlin test methods; inputs longer than 512 tokens are truncated.
|
|
|
|
|
|
|
|
|
|
| 78 |
|
| 79 |
## Citation
|
| 80 |
|
|
@@ -89,4 +117,4 @@ endorsed by its authors.
|
|
| 89 |
}
|
| 90 |
```
|
| 91 |
|
| 92 |
-
Dataset and method: [UT-SE-Research/FlakyLens](https://github.com/UT-SE-Research/FlakyLens).
|
|
|
|
| 22 |
A fine-tune of `microsoft/codebert-base` on the FlakeBench dataset from
|
| 23 |
[*Understanding and Improving Flaky Test Classification*](https://utexas.app.box.com/v/august-shi-OOPSLA2025)
|
| 24 |
(OOPSLA 2025), trained as a reproduction exercise on a single 8 GB consumer GPU.
|
|
|
|
| 25 |
The uploaded weights are the "Balanced" configuration below.
|
| 26 |
|
| 27 |
## Training configurations
|
|
|
|
| 45 |
Hardware: 1× RTX 4060 Laptop (8 GB). Class rebalancing is the one deviation from the
|
| 46 |
paper's method, which trains on the raw distribution (97% non-flaky).
|
| 47 |
|
| 48 |
+
## Results (per-category F1, fold 2 test set)
|
| 49 |
+
|
| 50 |
+
| Category | Baseline | lr 2e-5 | Balanced | Augmented | Paper |
|
| 51 |
+
| ---------------- | ---------: | ---------: | ---------: | ---------: | ---------: |
|
| 52 |
+
| Async Wait | 76.92% | 78.26% | 74.07% | 64.52% | 58.37% |
|
| 53 |
+
| Concurrency | 0.00% | 0.00% | 0.00% | 0.00% | 35.92% |
|
| 54 |
+
| Time | 57.14% | 66.67% | 66.67% | 40.00% | 72.73% |
|
| 55 |
+
| Unordered Coll. | 75.00% | 83.33% | 83.33% | 72.73% | 73.63% |
|
| 56 |
+
| Order Dep. | 82.35% | 86.96% | 95.24% | 73.68% | 64.35% |
|
| 57 |
+
| Non-flaky | 100.00% | 99.92% | 99.51% | 100.00% | 100.00% |
|
| 58 |
+
| **Macro F1** | **65.24%** | **69.19%** | **69.89%** | **58.49%** | **65.79%** |
|
| 59 |
+
|
| 60 |
+
The **Balanced** configuration (uploaded weights) achieves the best macro-F1 of
|
| 61 |
+
**69.89%** on this fold, outperforming the paper's reported 65.79% overall, but at
|
| 62 |
+
the cost of a **0.00% F1 on Concurrency** across every configuration tried — this
|
| 63 |
+
category collapses entirely regardless of rebalancing strategy, likely due to its
|
| 64 |
+
very small sample size (37 tests total across the whole dataset) and single-fold
|
| 65 |
+
evaluation rather than the paper's 4-fold average. The **Augmented** configuration
|
| 66 |
+
underperforms the paper on most categories, suggesting the synthetic variants used
|
| 67 |
+
for minority-class expansion may not be adding genuinely useful signal.
|
| 68 |
+
|
| 69 |
+
**Caveat**: these numbers come from a **single fold**, not the paper's 4-fold
|
| 70 |
+
average, so they are not directly comparable to the paper's reported scores without
|
| 71 |
+
that caveat in mind — a single lucky/unlucky project split can swing per-category
|
| 72 |
+
F1 substantially, especially for categories with as few as 33–41 test examples.
|
| 73 |
+
|
| 74 |
## Usage
|
| 75 |
|
| 76 |
```python
|
|
|
|
| 100 |
`4` test order dependency, `5` non-flaky.
|
| 101 |
|
| 102 |
Scope: Java/Kotlin test methods; inputs longer than 512 tokens are truncated.
|
| 103 |
+
**Known limitation**: this checkpoint does not predict Concurrency (label `1`) at
|
| 104 |
+
all on the fold-2 test set — treat any Concurrency-relevant use case with caution
|
| 105 |
+
until retrained with more Concurrency examples or evaluated across all 4 folds.
|
| 106 |
|
| 107 |
## Citation
|
| 108 |
|
|
|
|
| 117 |
}
|
| 118 |
```
|
| 119 |
|
| 120 |
+
Dataset and method: [UT-SE-Research/FlakyLens](https://github.com/UT-SE-Research/FlakyLens).
|