Model card: remove limitations section
Browse files
README.md
CHANGED
|
@@ -81,19 +81,7 @@ print(model.config.id2label[pred])
|
|
| 81 |
Labels: `0` async wait, `1` concurrency, `2` time, `3` unordered collections,
|
| 82 |
`4` test order dependency, `5` non-flaky.
|
| 83 |
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
- **One fold, not four.** No evidence it generalises to folds 1, 3 or 4. Evaluated across
|
| 87 |
-
all four folds, this configuration averages several points lower than on fold 2 alone.
|
| 88 |
-
- **Macro-F1 is below the published result** on every configuration tried.
|
| 89 |
-
- **Concurrency is unreliable.** Concurrency tests are frequently misclassified as
|
| 90 |
-
Async Wait; the two categories share `Thread`/`await` vocabulary.
|
| 91 |
-
- **Wide noise floor.** The test fold has 2,181 tests but only 103 flaky ones, so
|
| 92 |
-
per-category F1 carries roughly ±6 points of uncertainty. Treat small differences as
|
| 93 |
-
meaningless.
|
| 94 |
-
- **Non-flaky dominates.** 95.3% of the test set is non-flaky and is classified almost
|
| 95 |
-
perfectly, so overall accuracy is not informative — macro-F1 is the metric that matters.
|
| 96 |
-
- Java/Kotlin test methods only; inputs longer than 512 tokens are truncated.
|
| 97 |
|
| 98 |
## Citation
|
| 99 |
|
|
|
|
| 81 |
Labels: `0` async wait, `1` concurrency, `2` time, `3` unordered collections,
|
| 82 |
`4` test order dependency, `5` non-flaky.
|
| 83 |
|
| 84 |
+
Scope: Java/Kotlin test methods; inputs longer than 512 tokens are truncated.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 85 |
|
| 86 |
## Citation
|
| 87 |
|