FL-7B-3.1

A Qwen2.5-Coder-7B model adapted for COBOL, mainframe knowledge, and legacy-code modernization.

FL-7B-3.1 starts from Qwen/Qwen2.5-Coder-7B and was trained in two stages: continued pretraining (CPT) on COBOL source material, followed by assistant-only supervised fine-tuning (SFT) on COBOL and mainframe-oriented instructions.

The model is intended for:

  • generating and completing GnuCOBOL programs;
  • translating COBOL into Java;
  • answering mainframe and legacy-system questions;
  • explaining and summarizing COBOL source code.

Highlights

Benchmark Result
COBOLEval pass@1 17.81% (26/146)
COBOLEval test compilation rate 57.73% (474/821)
COBOLEval test pass rate 30.82% (253/821)
COBOL-to-Java CSR 61.54% (88/143)
COBOL-to-Java pass@1 48.25% (69/143)
MainframeBench MCQ accuracy 80.84% (1,561/1,931)

Usage with Transformers

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "FLs-AI/FL-7B-3.1"

tokenizer = AutoTokenizer.from_pretrained(
    model_id,
    fix_mistral_regex=True,
)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    dtype=torch.bfloat16,
    device_map="auto",
)

messages = [
    {
        "role": "user",
        "content": (
            "Write a complete GnuCOBOL 3.2 program that reads signed integers "
            "until EOF and prints their sum. Return only COBOL source code."
        ),
    }
]

inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=1536,
    do_sample=False,
    repetition_penalty=1.05,
)

generated = outputs[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(generated, skip_special_tokens=True))

Recommended generation settings

Use case Temperature Max new tokens Repetition penalty
COBOL generation 0 1536 1.05
COBOL to Java 0 4096 1.0
Mainframe MCQ 0 16 1.0
Mainframe QA / summarization 0 512 1.0

For the final COBOLEval run, repetition_penalty=1.05 produced the strongest measured result. COBOL relies heavily on repeated identifiers and fixed structural phrases, so large repetition penalties can damage syntax and correctness.

Evaluation

All results below were measured on the merged BF16 checkpoint with greedy decoding.

COBOLEval

COBOLEval evaluates generated programs by compiling and executing them with GnuCOBOL. The evaluation used 146 problems and 821 test cases, GnuCOBOL 3.2.0, max_new_tokens=1536, and one sample per task.

Model / setting pass@1 Test compilation rate Tests passed
Qwen2.5-Coder-7B base 0.00% 3.65% 4/821
FL-7B-3.1, repetition penalty 1.00 17.12% 41.29% 175/821
FL-7B-3.1, repetition penalty 1.05 17.81% 57.73% 253/821

The harness was pinned to commit 0bb96c3114bb2bb28e221e9d6000614781f8609d.

COBOL to Java

The COBOL-JavaTrans C2J evaluation compiles and executes generated Java translations.

Metric Result
Tasks 143
Compilation success rate (CSR) 61.54% (88/143)
pass@1 48.25% (69/143)

The evaluator was pinned to commit 2b14b7bf7e55556205654c6f7657fa60e36251fa.

MainframeBench

Fsoft-AIC/MainframeBench contains multiple-choice questions, open-ended QA, and COBOL code summarization.

Multiple choice

Tasks Correct Accuracy Invalid predictions
1,931 1,561 80.84% 1

Open-ended tasks

Suite Tasks Token F1 ROUGE-L F1 BLEU-4
Question answering 2,598 28.46% 24.31% 3.76
COBOL summarization 2,523 41.76% 36.94% 14.04

Normalized exact match was 0% for both open-ended suites. This strict lexical metric requires the generated response to match the single reference wording after normalization; it is not an accuracy or semantic-correctness score. Token F1, ROUGE-L, and BLEU-4 measure lexical overlap and should not be interpreted as execution-based correctness or human preference.

The dataset was pinned to revision 70d30c76eb29e45dd8965304b41c56bc1f527972.

Limitations

  • The model can enter repetition loops or produce excessively long code on difficult tasks.
  • A compiling program is not necessarily functionally correct or safe.
  • Evaluation used GnuCOBOL 3.2.0.
  • General-purpose coding performance inherited from Qwen2.5-Coder was not re-evaluated and may have regressed during domain adaptation.
  • MainframeBench QA and summarization results are lexical-overlap scores, not semantic accuracy.

Do not deploy generated code to production systems without compilation, tests, static analysis, and review by an experienced mainframe engineer.

License

This fine-tune is released under CC BY-NC 4.0. Attribution is required and commercial use of the fine-tuned weights is not permitted under this license. The Qwen2.5-Coder-7B base model is licensed separately under Apache 2.0.

Citation

@misc{fl7b31,
  title  = {FL-7B-3.1: COBOL and Mainframe Code Model},
  author = {FLs-AI},
  year   = {2026},
  url    = {https://huggingface.co/FLs-AI/FL-7B-3.1}
}

Acknowledgements

Downloads last month
46
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

2-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FLs-AI/FL-7B-3.1

Base model

Qwen/Qwen2.5-7B
Finetuned
(115)
this model

Collection including FLs-AI/FL-7B-3.1