thealper2's picture
Update README.md
608f394 verified
|
Raw History Blame Contribute Delete
4.92 kB
metadata
license: bsd-3-clause
base_model: Salesforce/codet5-base
datasets:
  - google/code_x_glue_cc_code_refinement
language:
  - code
library_name: transformers
pipeline_tag: text-generation
tags:
  - code
  - codet5
  - code-repair
  - program-repair
  - bug-fixing
  - java
  - seq2seq
model-index:
  - name: codet5-base-code-repair
    results:
      - task:
          type: text2text-generation
          name: Automated Program Repair
        dataset:
          name: CodeXGLUE code-refinement (small)
          type: google/code_x_glue_cc_code_refinement
          config: small
          split: test
        metrics:
          - type: exact_match
            value: 22.43
            name: Exact Match (%)
          - type: bleu
            value: 80.13
            name: BLEU

codet5-base-code-repair

Salesforce/codet5-base fine-tuned on the CodeXGLUE code-refinement small split for automated program repair: given a buggy Java method, the model generates the fixed version.

Inputs and outputs follow the dataset's abstracted Java style, where identifiers are normalised to tokens such as METHOD_1, VAR_1, TYPE_1 and STRING_1.

Usage

from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

model_id = "MODEL_ID"  # <- repo id
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)

buggy = (
    "public int METHOD_1 ( int VAR_1 ) { if ( VAR_1 = 0 ) { return 1 ; } "
    "return ( VAR_1 * ( METHOD_1 ( ( VAR_1 - 1 ) ) ) ) ; }"
)

inputs = tokenizer(buggy, max_length=256, truncation=True, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=256, num_beams=4, early_stopping=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

No task prefix is needed — the buggy snippet is fed in as-is. Beam search with num_beams=4 is the decoding setting used for all numbers reported below, and it is already stored in the repo's generation_config.json.

Results

Full splits (5,835 examples each), beam search with 4 beams:

Split Exact Match BLEU Loss
Validation 21.29% 80.27 0.1272
Test 22.43% 80.13 0.1257

Breakdown of the test-set predictions:

Outcome Share
Exact fix 22.43%
Partial fix (changed, closer but not exact) 13.49%
Input copied unchanged 3.38%
Incorrect 64.08%

The high BLEU next to the modest exact-match rate is expected for this task: the fixed method is usually a near-copy of the buggy one, so most generated tokens are correct even when the actual bug is not fixed. Exact match is the metric that matters here; BLEU mostly measures how well the model preserves the surrounding code.

Validation exact match by epoch (1,000-example in-training subset):

Epoch 1 2 3 4 5 6 7 8 9 10
EM 10.5 15.2 17.5 19.5 19.8 20.6 20.9 20.6 21.2 21.0

The epoch-9 checkpoint scored best and is the one published here.

Training

Base model Salesforce/codet5-base (~223M params)
Dataset google/code_x_glue_cc_code_refinement, config small
Train / validation / test 46,680 / 5,835 / 5,835
Epochs 10 (best checkpoint by exact match kept)
Learning rate 5e-5, linear decay, 5% warmup
Batch size 16 × 2 gradient accumulation (effective 32)
Weight decay 0.01
Max grad norm 1.0
Max source / target length 256 / 256 tokens
Precision bf16
Seed 42
Training time ~1h25m on a single GPU

No example in any split was truncated at 256 tokens (longest source: 132 tokens), and the dataset contains no identical buggy/fixed pairs.

Limitations

  • Trained only on abstracted Java methods from CodeXGLUE. Real-world code with actual identifier names, or any other language, is out of distribution and will perform much worse.
  • Handles single, self-contained methods — no cross-file or repository-level context.
  • Roughly two thirds of test inputs are still not repaired correctly. Treat outputs as suggestions to review, not as verified fixes, and always re-run your tests.
  • The model can return the input unchanged (3.4% of the test set) when it finds no fix.

License

Released under BSD-3-Clause, following the Salesforce/codet5-base base model. The training data, CodeXGLUE code-refinement, is distributed under the Computational Use of Data Agreement (C-UDA).