MindForge-27B

MindForge-27B is a full-parameter supervised fine-tune of Qwen3.6-27B for whole-life-cycle software engineering: exploring specifications, designing architectures, implementing programs, debugging, verifying, and refining solutions through tool use.

This repository contains the trained model weights described in MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis (arXiv:2607.27146).

On ProgramBench, MindForge-27B improves average test pass rate from 37.98% to 49.51%: +11.53 percentage points and +30.4% relative over its base model. It also improves on all seven out-of-distribution benchmarks reported in the MindForge paper, with NL2Repo evaluated both with and without tests.

Model Details

Property Value
Base model Qwen/Qwen3.6-27B
Fine-tuning method Full-parameter SFT of the language model; vision tower and aligner frozen
Training data MindForge whole-life-cycle program-synthesis trajectories
Precision BF16
Configured context length 262,144 tokens
Weight format Safetensors, 12 shards

Training Data

The training data is available on Hugging Face: MindForge-27B-Training-Trajectories.

MindForge turns open-source command-line programs into source-free (cleanroom) environments. A teacher agent, GLM-5.2 driven by mini-swe-agent, receives only a compiled reference executable and sanitized public documentation. Without access to the original source or the internet, it must discover behavior by probing the binary, design an implementation, build it from scratch, and iterate toward a passing build.

The data release contains 1,001 complete trajectories collected from 562 programs spanning Go, Rust, C, C++, Swift, and TypeScript. These programs come from repositories disjoint from ProgramBench. Of the released trajectories, 973 fit within the 262,144-token training length budget.

Trajectory statistic Turns Tokens
Mean 181.6 177K
Median 176 182K
Minimum 39 37K
Maximum 477 272K

Training supervises assistant reasoning, natural-language responses, and tool-call tokens. System messages, user messages, and tool outputs are masked. Assistant reasoning is represented with <think> tags.

Evaluation

All results below come from the experiments reported in the MindForge paper. Agent scaffolds, tool access, and execution budgets matter when interpreting or reproducing these results; the loading example below is not a benchmark reproduction recipe.

ProgramBench

Average test pass rate on 200 instances; higher is better.

Model Pass rate (%)
GLM-5.2 (teacher) 64.60
GPT-5.5 56.50
Claude Opus 4.7 51.38
MindForge-27B 49.51
Sonnet 4.6 47.97
GPT-5.4 38.08
Qwen3.6-27B (base) 37.98

MindForge-27B scores strictly higher than the base model on 152 of 200 tasks, lower on 43, and ties on 5.

Generalization to Unseen Benchmarks

None of these benchmarks were used during training. Scores use each benchmark's reported metric and should not be compared across rows.

Benchmark Qwen3.6-27B MindForge-27B
RepoZero-C2Rust 47.00 78.00
DeepSWE 1.76 15.92
NL2Repo (with tests) 61.27 71.97
SWE-bench Pro 45.41 51.34
SWE-bench Multilingual 62.55 67.77
SWE-bench Verified 68.80 73.84
FeatBench 50.10 55.05
NL2Repo (without tests) 18.92 23.48

Agent Behavior

The reported behavioral analysis finds more sustained environment interaction after fine-tuning, alongside a lower command failure rate:

Metric Qwen3.6-27B MindForge-27B
Mean turns 344.0 735.7
Mean tool calls 174.4 373.0
Command failure rate 10.98% 9.35%
Editing after reasoning 27.8% 50.1%
Editing after failure recovery 31.8% 48.8%
Mean reference-program line coverage 49.34% 58.39%

Line coverage is measured by replaying agent invocations against coverage-instrumented reference binaries. These results suggest a trade-off: improved task performance comes with substantially longer interactions and more tool calls.

Usage

Load with Transformers

Use a Transformers installation that supports Qwen3_5ForConditionalGeneration. The checkpoint configuration records Transformers 5.12.1 as its export version; this is not a claim about the minimum supported version. Install PyTorch for your hardware and install compatible transformers and accelerate packages.

Load the model directly from Hugging Face:

import torch
from transformers import AutoTokenizer, Qwen3_5ForConditionalGeneration

model_id = "centre-for-swe/MindForge-27B"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
    model_id,
    dtype=torch.bfloat16,
    device_map="auto",
).eval()

messages = [
    {
        "role": "user",
        "content": "What is 1+1?",
    }
]
prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True,
)
inputs = tokenizer(prompt, add_special_tokens=False, return_tensors="pt").to(model.device)

with torch.inference_mode():
    outputs = model.generate(
        **inputs,
        max_new_tokens=2048,
        do_sample=False,
        eos_token_id=tokenizer.eos_token_id,
        pad_token_id=tokenizer.pad_token_id,
    )

response = outputs[0, inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(response, skip_special_tokens=True))

The example is a text-generation smoke test.

Citation

@misc{chen2026mindforgeteachingsmalllanguage,
      title={MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis},
      author={Yihao Chen and Shi Chang and Khaled Chawa and Feng Lin and Boyuan Chen and Shaowei Wang and Ahmed E. Hassan},
      year={2026},
      eprint={2607.27146},
      archivePrefix={arXiv},
      primaryClass={cs.SE},
      url={https://arxiv.org/abs/2607.27146},
}

License

The MindForge-27B model weights and MindForge training trajectories are released under the MIT license.

Downloads last month
-
Safetensors
Model size
28B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for centre-for-swe/MindForge-27B

Base model

Qwen/Qwen3.6-27B
Finetuned
(405)
this model

Paper for centre-for-swe/MindForge-27B