Ariful1904129 commited on
Commit
0337b7f
·
verified ·
1 Parent(s): 53a9e54

CodeBERT flaky-test classifier, FlakeBench fold 2 (59.12% macro-F1, single-fold reproduction)

Browse files
README.md ADDED
@@ -0,0 +1,126 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ language:
4
+ - code
5
+ base_model: microsoft/codebert-base
6
+ pipeline_tag: text-classification
7
+ tags:
8
+ - flaky-tests
9
+ - software-testing
10
+ - code
11
+ - codebert
12
+ - reproduction
13
+ ---
14
+
15
+ # CodeBERT for Flaky Test Categorisation (FlakeBench, fold 2)
16
+
17
+ Classifies a Java/Kotlin test method into one of six categories: five kinds of flaky
18
+ test plus non-flaky.
19
+
20
+ **This is a partial, single-fold reproduction of published work — not the authors' model
21
+ and not a match for their reported results.** Please read the limitations before using it.
22
+
23
+ ## What this is
24
+
25
+ A fine-tune of `microsoft/codebert-base` on the FlakeBench dataset from
26
+ [*Understanding and Improving Flaky Test Classification*](https://utexas.app.box.com/v/august-shi-OOPSLA2025)
27
+ (OOPSLA 2025), trained as a reproduction exercise on a single 8 GB consumer GPU.
28
+
29
+ Trained on **one** of the paper's four project-disjoint folds (project group 2), so it is
30
+ **not** comparable to the paper's four-fold headline number and should not be described
31
+ as reproducing it.
32
+
33
+ ## Results
34
+
35
+ Held-out test set: 2,181 tests from 25 projects unseen during training.
36
+
37
+ | Category | This model | Paper (4-fold) |
38
+ |---|---|---|
39
+ | Async Wait | 55.38% | 58.37% |
40
+ | Concurrency | 10.53% | 35.92% |
41
+ | Time | 57.14% | 72.73% |
42
+ | Unordered Collections | 54.55% | 73.63% |
43
+ | Test Order Dependency | **77.11%** | 64.35% |
44
+ | Non-flaky | 100.00% | 100.00% |
45
+ | **Macro F1** | **59.12%** | **65.79%** |
46
+
47
+ Macro-F1 is **6.67 points below** the paper. Order-Dependency is the one category that
48
+ exceeds it. Concurrency is weak (10.53%): 12 of 17 Concurrency tests are misclassified as
49
+ Async Wait — the two categories share `Thread`/`await` vocabulary and the model does not
50
+ separate them.
51
+
52
+ Against a faithful baseline trained under the same compute budget without rebalancing
53
+ (45.59%), this model is **+13.53 points** better — bootstrap 95% CI [+4.48, +21.23],
54
+ p = 0.999, 2,000 resamples.
55
+
56
+ ## Training
57
+
58
+ | | |
59
+ |---|---|
60
+ | Base | `microsoft/codebert-base` |
61
+ | Head | Linear(768→512) → ReLU → Dropout(0.3) → Linear(512→6) |
62
+ | Loss | Focal loss, γ=2.0, balanced class weights |
63
+ | Optimiser | AdamW, lr 1e-5, weight decay 0.01 |
64
+ | Batch / max length | 8 / 512 |
65
+ | Precision | fp16 |
66
+ | Epochs | 18 (best at 6, selected on validation macro-F1) |
67
+ | Rebalancing | non-flaky undersampled 4,972→800; each flaky class duplicated to 160 |
68
+ | Hardware | 1× RTX 4060 Laptop (8 GB) |
69
+
70
+ Class rebalancing is the one deviation from the paper's method, which trains on the raw
71
+ distribution (97% non-flaky). It is what produced the gain over the baseline.
72
+
73
+ ## Usage
74
+
75
+ ```python
76
+ from transformers import AutoTokenizer, AutoModelForSequenceClassification
77
+ import torch
78
+
79
+ name = "Ariful1904129/codebert-flakytest-fold2"
80
+ tok = AutoTokenizer.from_pretrained(name)
81
+ model = AutoModelForSequenceClassification.from_pretrained(name, trust_remote_code=True).eval()
82
+
83
+ code = """@Test
84
+ public void testConnect() throws Exception {
85
+ Thread.sleep(1000);
86
+ assertTrue(client.isConnected());
87
+ }"""
88
+
89
+ x = tok(code, return_tensors="pt", truncation=True, max_length=512)
90
+ with torch.no_grad():
91
+ pred = model(**x).logits.argmax(-1).item()
92
+ print(model.config.id2label[pred])
93
+ ```
94
+
95
+ `trust_remote_code=True` is required — the MLP head is a custom architecture defined in
96
+ `modeling_flakylens.py`.
97
+
98
+ Labels: `0` async wait, `1` concurrency, `2` time, `3` unordered collections,
99
+ `4` test order dependency, `5` non-flaky.
100
+
101
+ ## Limitations
102
+
103
+ - **One fold, not four.** No evidence it generalises to folds 1, 3 or 4.
104
+ - **Below the published result.** 59.12% vs 65.79%.
105
+ - **Concurrency is unreliable** (10.53% F1) — it mostly predicts Async Wait instead.
106
+ - **Noise floor is wide.** The test fold has 2,181 tests but only 103 flaky ones, so
107
+ per-category F1 carries roughly ±6 points of uncertainty. Treat small differences as
108
+ meaningless.
109
+ - **Non-flaky dominates.** 95.3% of the test set is non-flaky and scores 100%, so overall
110
+ accuracy (97%) is not informative — macro-F1 is the metric that matters.
111
+ - Java/Kotlin test methods only; inputs longer than 512 tokens are truncated.
112
+
113
+ ## Citation
114
+
115
+ Please cite the original paper. This model is a third-party reproduction and is not
116
+ endorsed by its authors.
117
+
118
+ ```bibtex
119
+ @inproceedings{flakylens2025,
120
+ title = {Understanding and Improving Flaky Test Classification},
121
+ booktitle = {OOPSLA},
122
+ year = {2025}
123
+ }
124
+ ```
125
+
126
+ Dataset and method: [UT-SE-Research/FlakyLens](https://github.com/UT-SE-Research/FlakyLens).
config.json ADDED
@@ -0,0 +1,49 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "FlakyLensForTestClassification"
4
+ ],
5
+ "attention_probs_dropout_prob": 0.1,
6
+ "auto_map": {
7
+ "AutoConfig": "modeling_flakylens.FlakyLensConfig",
8
+ "AutoModelForSequenceClassification": "modeling_flakylens.FlakyLensForTestClassification"
9
+ },
10
+ "bos_token_id": 0,
11
+ "classifier_dropout": null,
12
+ "eos_token_id": 2,
13
+ "head_dropout": 0.3,
14
+ "head_hidden": 512,
15
+ "hidden_act": "gelu",
16
+ "hidden_dropout_prob": 0.1,
17
+ "hidden_size": 768,
18
+ "id2label": {
19
+ "0": "async wait",
20
+ "1": "concurrency",
21
+ "2": "time",
22
+ "3": "unordered collections",
23
+ "4": "test order dependency",
24
+ "5": "non-flaky"
25
+ },
26
+ "initializer_range": 0.02,
27
+ "intermediate_size": 3072,
28
+ "label2id": {
29
+ "async wait": 0,
30
+ "concurrency": 1,
31
+ "non-flaky": 5,
32
+ "test order dependency": 4,
33
+ "time": 2,
34
+ "unordered collections": 3
35
+ },
36
+ "layer_norm_eps": 1e-05,
37
+ "max_position_embeddings": 514,
38
+ "model_type": "flakylens",
39
+ "num_attention_heads": 12,
40
+ "num_hidden_layers": 12,
41
+ "output_past": true,
42
+ "pad_token_id": 1,
43
+ "position_embedding_type": "absolute",
44
+ "torch_dtype": "float32",
45
+ "transformers_version": "4.40.1",
46
+ "type_vocab_size": 1,
47
+ "use_cache": true,
48
+ "vocab_size": 50265
49
+ }
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:dd7e599447b3fe15ac3ed2c09cf6c7b51343f0cb2bcfcf591cbbb8320f9c2c33
3
+ size 500194032
modeling_flakylens.py ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """CodeBERT + MLP head for flaky-test category classification.
2
+
3
+ Mirrors BERT_Arch from the FlakyLens artifact: RoBERTa pooled output -> Linear(768,512)
4
+ -> ReLU -> Dropout -> Linear(512,6). The original applies LogSoftmax to the final layer;
5
+ this returns raw logits instead, following the HF convention. argmax is identical either
6
+ way, so predictions match exactly -- apply log_softmax if you need the original's values.
7
+ """
8
+ import torch.nn as nn
9
+ from transformers import RobertaConfig, RobertaModel, RobertaPreTrainedModel
10
+ from transformers.modeling_outputs import SequenceClassifierOutput
11
+
12
+
13
+ class FlakyLensConfig(RobertaConfig):
14
+ model_type = "flakylens"
15
+
16
+ def __init__(self, head_hidden=512, head_dropout=0.3, **kwargs):
17
+ super().__init__(**kwargs)
18
+ self.head_hidden = head_hidden
19
+ self.head_dropout = head_dropout
20
+
21
+
22
+ class FlakyLensForTestClassification(RobertaPreTrainedModel):
23
+ config_class = FlakyLensConfig
24
+
25
+ def __init__(self, config):
26
+ super().__init__(config)
27
+ self.roberta = RobertaModel(config, add_pooling_layer=True)
28
+ self.fc1 = nn.Linear(config.hidden_size, config.head_hidden)
29
+ self.relu = nn.ReLU()
30
+ self.dropout = nn.Dropout(config.head_dropout)
31
+ self.fc2 = nn.Linear(config.head_hidden, config.num_labels)
32
+ self.post_init()
33
+
34
+ def forward(self, input_ids=None, attention_mask=None, labels=None, **kwargs):
35
+ outputs = self.roberta(input_ids=input_ids, attention_mask=attention_mask)
36
+ pooled = outputs[1]
37
+ logits = self.fc2(self.dropout(self.relu(self.fc1(pooled))))
38
+ loss = None
39
+ if labels is not None:
40
+ loss = nn.functional.cross_entropy(logits, labels)
41
+ return SequenceClassifierOutput(loss=loss, logits=logits)
special_tokens_map.json ADDED
@@ -0,0 +1,51 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": {
3
+ "content": "<s>",
4
+ "lstrip": false,
5
+ "normalized": true,
6
+ "rstrip": false,
7
+ "single_word": false
8
+ },
9
+ "cls_token": {
10
+ "content": "<s>",
11
+ "lstrip": false,
12
+ "normalized": true,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "eos_token": {
17
+ "content": "</s>",
18
+ "lstrip": false,
19
+ "normalized": true,
20
+ "rstrip": false,
21
+ "single_word": false
22
+ },
23
+ "mask_token": {
24
+ "content": "<mask>",
25
+ "lstrip": true,
26
+ "normalized": false,
27
+ "rstrip": false,
28
+ "single_word": false
29
+ },
30
+ "pad_token": {
31
+ "content": "<pad>",
32
+ "lstrip": false,
33
+ "normalized": true,
34
+ "rstrip": false,
35
+ "single_word": false
36
+ },
37
+ "sep_token": {
38
+ "content": "</s>",
39
+ "lstrip": false,
40
+ "normalized": true,
41
+ "rstrip": false,
42
+ "single_word": false
43
+ },
44
+ "unk_token": {
45
+ "content": "<unk>",
46
+ "lstrip": false,
47
+ "normalized": true,
48
+ "rstrip": false,
49
+ "single_word": false
50
+ }
51
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,57 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "added_tokens_decoder": {
4
+ "0": {
5
+ "content": "<s>",
6
+ "lstrip": false,
7
+ "normalized": true,
8
+ "rstrip": false,
9
+ "single_word": false,
10
+ "special": true
11
+ },
12
+ "1": {
13
+ "content": "<pad>",
14
+ "lstrip": false,
15
+ "normalized": true,
16
+ "rstrip": false,
17
+ "single_word": false,
18
+ "special": true
19
+ },
20
+ "2": {
21
+ "content": "</s>",
22
+ "lstrip": false,
23
+ "normalized": true,
24
+ "rstrip": false,
25
+ "single_word": false,
26
+ "special": true
27
+ },
28
+ "3": {
29
+ "content": "<unk>",
30
+ "lstrip": false,
31
+ "normalized": true,
32
+ "rstrip": false,
33
+ "single_word": false,
34
+ "special": true
35
+ },
36
+ "50264": {
37
+ "content": "<mask>",
38
+ "lstrip": true,
39
+ "normalized": false,
40
+ "rstrip": false,
41
+ "single_word": false,
42
+ "special": true
43
+ }
44
+ },
45
+ "bos_token": "<s>",
46
+ "clean_up_tokenization_spaces": true,
47
+ "cls_token": "<s>",
48
+ "eos_token": "</s>",
49
+ "errors": "replace",
50
+ "mask_token": "<mask>",
51
+ "model_max_length": 512,
52
+ "pad_token": "<pad>",
53
+ "sep_token": "</s>",
54
+ "tokenizer_class": "RobertaTokenizer",
55
+ "trim_offsets": true,
56
+ "unk_token": "<unk>"
57
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff