Ariful1904129 commited on
Commit
d2dfc53
·
verified ·
1 Parent(s): 1729b8f

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +30 -2
README.md CHANGED
@@ -22,7 +22,6 @@ test plus non-flaky.
22
  A fine-tune of `microsoft/codebert-base` on the FlakeBench dataset from
23
  [*Understanding and Improving Flaky Test Classification*](https://utexas.app.box.com/v/august-shi-OOPSLA2025)
24
  (OOPSLA 2025), trained as a reproduction exercise on a single 8 GB consumer GPU.
25
-
26
  The uploaded weights are the "Balanced" configuration below.
27
 
28
  ## Training configurations
@@ -46,6 +45,32 @@ The uploaded weights are the "Balanced" configuration below.
46
  Hardware: 1× RTX 4060 Laptop (8 GB). Class rebalancing is the one deviation from the
47
  paper's method, which trains on the raw distribution (97% non-flaky).
48
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
49
  ## Usage
50
 
51
  ```python
@@ -75,6 +100,9 @@ Labels: `0` async wait, `1` concurrency, `2` time, `3` unordered collections,
75
  `4` test order dependency, `5` non-flaky.
76
 
77
  Scope: Java/Kotlin test methods; inputs longer than 512 tokens are truncated.
 
 
 
78
 
79
  ## Citation
80
 
@@ -89,4 +117,4 @@ endorsed by its authors.
89
  }
90
  ```
91
 
92
- Dataset and method: [UT-SE-Research/FlakyLens](https://github.com/UT-SE-Research/FlakyLens).
 
22
  A fine-tune of `microsoft/codebert-base` on the FlakeBench dataset from
23
  [*Understanding and Improving Flaky Test Classification*](https://utexas.app.box.com/v/august-shi-OOPSLA2025)
24
  (OOPSLA 2025), trained as a reproduction exercise on a single 8 GB consumer GPU.
 
25
  The uploaded weights are the "Balanced" configuration below.
26
 
27
  ## Training configurations
 
45
  Hardware: 1× RTX 4060 Laptop (8 GB). Class rebalancing is the one deviation from the
46
  paper's method, which trains on the raw distribution (97% non-flaky).
47
 
48
+ ## Results (per-category F1, fold 2 test set)
49
+
50
+ | Category | Baseline | lr 2e-5 | Balanced | Augmented | Paper |
51
+ | ---------------- | ---------: | ---------: | ---------: | ---------: | ---------: |
52
+ | Async Wait | 76.92% | 78.26% | 74.07% | 64.52% | 58.37% |
53
+ | Concurrency | 0.00% | 0.00% | 0.00% | 0.00% | 35.92% |
54
+ | Time | 57.14% | 66.67% | 66.67% | 40.00% | 72.73% |
55
+ | Unordered Coll. | 75.00% | 83.33% | 83.33% | 72.73% | 73.63% |
56
+ | Order Dep. | 82.35% | 86.96% | 95.24% | 73.68% | 64.35% |
57
+ | Non-flaky | 100.00% | 99.92% | 99.51% | 100.00% | 100.00% |
58
+ | **Macro F1** | **65.24%** | **69.19%** | **69.89%** | **58.49%** | **65.79%** |
59
+
60
+ The **Balanced** configuration (uploaded weights) achieves the best macro-F1 of
61
+ **69.89%** on this fold, outperforming the paper's reported 65.79% overall, but at
62
+ the cost of a **0.00% F1 on Concurrency** across every configuration tried — this
63
+ category collapses entirely regardless of rebalancing strategy, likely due to its
64
+ very small sample size (37 tests total across the whole dataset) and single-fold
65
+ evaluation rather than the paper's 4-fold average. The **Augmented** configuration
66
+ underperforms the paper on most categories, suggesting the synthetic variants used
67
+ for minority-class expansion may not be adding genuinely useful signal.
68
+
69
+ **Caveat**: these numbers come from a **single fold**, not the paper's 4-fold
70
+ average, so they are not directly comparable to the paper's reported scores without
71
+ that caveat in mind — a single lucky/unlucky project split can swing per-category
72
+ F1 substantially, especially for categories with as few as 33–41 test examples.
73
+
74
  ## Usage
75
 
76
  ```python
 
100
  `4` test order dependency, `5` non-flaky.
101
 
102
  Scope: Java/Kotlin test methods; inputs longer than 512 tokens are truncated.
103
+ **Known limitation**: this checkpoint does not predict Concurrency (label `1`) at
104
+ all on the fold-2 test set — treat any Concurrency-relevant use case with caution
105
+ until retrained with more Concurrency examples or evaluated across all 4 folds.
106
 
107
  ## Citation
108
 
 
117
  }
118
  ```
119
 
120
+ Dataset and method: [UT-SE-Research/FlakyLens](https://github.com/UT-SE-Research/FlakyLens).