File size: 13,359 Bytes
3cb8df1
 
 
95ce0df
3cb8df1
 
 
 
 
 
 
 
 
eb2c5cd
 
3cb8df1
6b0b5f0
643ba79
3cb8df1
 
c608b00
95ce0df
c608b00
 
95ce0df
 
 
c608b00
 
f66b9ba
95ce0df
 
 
 
6b0b5f0
c608b00
95ce0df
c608b00
95ce0df
c608b00
95ce0df
 
 
eb2c5cd
95ce0df
1f40fe8
95ce0df
1f40fe8
95ce0df
1f40fe8
95ce0df
 
 
 
 
 
 
 
1f40fe8
95ce0df
 
 
 
 
 
1f40fe8
95ce0df
1f40fe8
95ce0df
1f40fe8
95ce0df
5f53985
95ce0df
 
 
 
1f40fe8
95ce0df
1f40fe8
95ce0df
5f53985
95ce0df
5f53985
95ce0df
 
 
 
 
 
5f53985
95ce0df
1f40fe8
95ce0df
5f53985
95ce0df
 
 
1f40fe8
95ce0df
1f40fe8
95ce0df
19330de
95ce0df
 
 
1f40fe8
95ce0df
5f53985
95ce0df
 
 
 
 
f66b9ba
95ce0df
5f53985
95ce0df
 
 
 
1b84f75
95ce0df
1b84f75
95ce0df
0e9564d
95ce0df
eb2c5cd
95ce0df
 
 
 
 
 
 
 
 
 
 
 
 
1f40fe8
95ce0df
 
 
2b50076
95ce0df
2b50076
95ce0df
2b50076
95ce0df
179260f
95ce0df
 
1f40fe8
95ce0df
 
 
 
 
 
1f40fe8
95ce0df
 
 
 
 
1f40fe8
95ce0df
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1f40fe8
 
95ce0df
1f40fe8
95ce0df
 
 
 
 
 
 
 
 
1f40fe8
 
95ce0df
 
 
 
 
 
 
 
1f40fe8
 
95ce0df
1f40fe8
95ce0df
1f40fe8
95ce0df
1f40fe8
95ce0df
 
 
 
 
 
 
 
 
 
179260f
95ce0df
1f40fe8
95ce0df
1f40fe8
95ce0df
 
 
 
 
 
1f40fe8
95ce0df
1f40fe8
95ce0df
 
 
 
 
 
 
 
1f40fe8
95ce0df
eb2c5cd
 
95ce0df
 
 
 
 
0aa9051
eb2c5cd
 
1f40fe8
95ce0df
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
---
language:
- en
- pt
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
tags:
- qwen
- unsloth
- lora
- code-generation
- simplicio-loop
- software-engineering
- surgical-diff
- agentic-coding
pipeline_tag: text-generation
pretty_name: Simplicio 27B
homepage: https://simpleti.com.br/simplicio-27b/
---

<p align="center">
  <img src="https://raw.githubusercontent.com/simpletibr/simplicio-27b/main/assets/simplicio-logo.png" width="130" alt="SimpleTI logo">
</p>

<h1 align="center">Simplicio 27B</h1>

<p align="center">A Qwen3.8-27B fine-tune that answers code-change requests with SEARCH/REPLACE patches instead of whole files.</p>

<p align="center">
  <a href="https://huggingface.co/wesleysimplicio/Simplicio-27B"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Simplicio--27B-yellow.svg" alt="Hugging Face"></a>
  <a href="https://ollama.com/wesleysimplicio/simplicio-27b"><img src="https://img.shields.io/badge/Ollama-simplicio--27b-black" alt="Ollama"></a>
  <a href="https://github.com/simpletibr/simplicio-27b"><img src="https://img.shields.io/badge/GitHub-simplicio--27b-blue?logo=github" alt="GitHub"></a>
  <a href="https://simpleti.com.br/simplicio-27b/"><img src="https://img.shields.io/badge/SimpleTI-Official%20Page-0081FB" alt="SimpleTI official page"></a>
  <a href="https://github.com/simpletibr/simplicio-27b/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-green.svg" alt="License: Apache 2.0"></a>
</p>

## Overview

Simplicio 27B is a LoRA fine-tune of [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) by Wesley Simplicio at [SimpleTI](https://simpleti.com.br/simplicio-27b/). It answers a code-change request in five tagged phases (`<orient>`, `<plan>`, `<patch>`, `<validate>`, `<deliver>`). The `<patch>` phase holds SEARCH/REPLACE blocks that touch only the lines that change.

- **Base model:** Qwen3.8-27B, 27.36B parameters, 64 layers that alternate linear attention (DeltaNet) and full attention in a 3:1 pattern.
- **Files on Hugging Face:** the LoRA adapter (0.64 GB), the merged BF16 checkpoint (18 shards, 55.6 GB), and a GGUF Q4_K_M (16.8 GB) with the base model's vision projector (0.93 GB).
- **Status:** research release. It passes 46.7% of the project's own 120 held-out tasks. It has not been scored on public benchmarks.

## Results

### Held-out set (120 tasks)

The full BF16 model ran on Google Colab G4 (NVIDIA RTX PRO 6000 Blackwell, 95 GB) against 120 tasks that were not used in training ([`data/unseen_eval_120.json`](https://github.com/simpletibr/simplicio-27b/blob/main/data/unseen_eval_120.json)). Each task got one attempt at temperature 0, and generation stopped at `</deliver>`. A task passes when its unit test passes after the patch is applied. Aggregate results: [`benchmarks/live_colab_g4_bf16_n120.json`](https://github.com/simpletibr/simplicio-27b/blob/main/benchmarks/live_colab_g4_bf16_n120.json).

| Metric | Result |
|---|---:|
| Unit test passes | **56/120 (46.7%)** |
| SEARCH block found, patch applied | 104/120 (86.7%) |
| Patched file parses (AST) | 104/120 (86.7%) |
| No call to a nonexistent API | 120/120 |
| Output tokens per task, mean / max | 69.9 / 103 |
| Wall time for all 120 tasks | 438 s |

| Category | Tasks | Unit test passes | Patch applied |
|---|---:|---:|---:|
| Surgical diff and AST precision | 30 | 17 | 20 |
| Edge-case correctness | 30 | 6 | 27 |
| Nonexistent and deprecated API traps | 30 | 27 | 30 |
| Adversarial and out-of-distribution | 30 | 6 | 27 |

The base model has not been run under this protocol yet, so the gain from fine-tuning is not measured.

### Withdrawn numbers

Earlier versions of this card reported 96.5% accuracy, 116 of 120 tasks passed, 480 tokens per task, a "Top 12" leaderboard and per-token prices. None of these came from running the model on those tasks:

- 116/120 and its McNemar test come from [`benchmarks/prove_benchmark_120.py`](https://github.com/simpletibr/simplicio-27b/blob/main/benchmarks/prove_benchmark_120.py), which simulates model outputs.
- 480 tokens per task comes from a 3-task smoke test on an A100 ([`benchmarks/empirical_a100_results.json`](https://github.com/simpletibr/simplicio-27b/blob/main/benchmarks/empirical_a100_results.json)).
- 96.5% and the leaderboard rows are hard-coded in [`benchmarks/compare_top10_2026.py`](https://github.com/simpletibr/simplicio-27b/blob/main/benchmarks/compare_top10_2026.py).
- The model is not listed on OpenRouter and has no public per-token price.

The held-out run above replaces them.

### Public coding-agent benchmarks

Simplicio 27B has not been evaluated on SWE-bench, the Aider benchmark, Terminal-Bench or the Artificial Analysis Coding Agent Index. For scale, these are published Coding Agent Index v1.5 results ([Artificial Analysis](https://artificialanalysis.ai/agents/coding-agents/comparisons/claude-code-vs-codex), retrieved 5 October 2026):

| Agent | Coding Agent Index | DeepSWE v1.1 | Terminal-Bench 4.0 | SWE-Atlas-QnA | Cost per task |
|---|---:|---:|---:|---:|---:|
| Claude Opus 5.5 (max) | 66 | 68% | 63% | 66% | $13.04 |
| Claude Sonnet 5.5 (max) | 68 | 72% | 66% | 67% | $14.19 |
| GPT-6.1 Sol (xhigh) | 63 | 73% | 55% | 61% | $1.04 |
| Simplicio 27B | not evaluated | – | – | – | – |

## Quick start

### Ollama

```bash
ollama run wesleysimplicio/simplicio-27b
```

The `latest` tag holds the Q4_K_M GGUF and the vision projector. It uses temperature 0.2 and a 32,768-token context, and it stops at `<|im_end|>` and `</deliver>`.

The installer installs Ollama if it is missing, then runs the model:

```bash
curl -fsSL https://raw.githubusercontent.com/simpletibr/simplicio-27b/main/install.sh | bash
```

### vLLM (OpenAI-compatible server with tool calls)

```bash
git clone https://github.com/simpletibr/simplicio-27b
cd simplicio-27b
./deploy/serve_vllm.sh wesleysimplicio/Simplicio-27B 8000
```

This serves the merged BF16 checkpoint. The weights alone take 55.6 GB; the evaluation above ran on a 95 GB GPU. The script sets:

- `--max-model-len 40960`, defined once in [`deploy/context.env`](https://github.com/simpletibr/simplicio-27b/blob/main/deploy/context.env): a measured 31,692-token OpenCode prompt plus 4,096 output tokens.
- `--chat-template` with [`deploy/chat_template_chatml.jinja`](https://github.com/simpletibr/simplicio-27b/blob/main/deploy/chat_template_chatml.jinja). It prefills `<think>` so that `--reasoning-parser qwen3` moves reasoning out of `content`.
- `--enable-auto-tool-choice --tool-call-parser simplicio`, using [`deploy/simplicio_tool_parser.py`](https://github.com/simpletibr/simplicio-27b/blob/main/deploy/simplicio_tool_parser.py). It turns `<tool><name>…</name><params>…</params></tool>` into a single `tool_calls` entry.
- `--served-model-name simplicio-27b simpleti/simplicio-27b`.

On a smaller GPU, [`Simplicio_27B_Serve_Colab.ipynb`](https://github.com/simpletibr/simplicio-27b/blob/main/Simplicio_27B_Serve_Colab.ipynb) serves the 4-bit base with the LoRA adapter on Colab.

### OpenCode

Add the vLLM server to `opencode.json` as an OpenAI-compatible provider:

```json
{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "simplicio": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Simplicio 27B",
      "options": { "baseURL": "http://localhost:8000/v1" },
      "models": { "simplicio-27b": { "name": "Simplicio 27B" } }
    }
  }
}
```

```bash
opencode -m simplicio/simplicio-27b
```

OpenCode works through tool calls, so point it at the vLLM server. The Ollama template does not declare tools.

### Python (Unsloth)

This loads the adapter on its 4-bit base, the same way [`Simplicio_27B_Merge_Colab.ipynb`](https://github.com/simpletibr/simplicio-27b/blob/main/Simplicio_27B_Merge_Colab.ipynb) does:

```python
from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="wesleysimplicio/Simplicio-27B",  # adapter; the base comes from adapter_config.json
    max_seq_length=16384,
    load_in_4bit=True,
)
FastLanguageModel.for_inference(model)

messages = [{"role": "user", "content": "In api/schemas/user.py, accept tax_id with punctuation such as 123.456.789-00."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to("cuda")
output = model.generate(inputs, max_new_tokens=1024, do_sample=False)
print(tokenizer.decode(output[0][inputs.shape[1]:], skip_special_tokens=True))
```

## Output format

Each phase is a list of numbered points. This example is shortened and translated from the first training example:

```text
<simplicio_loop>
<orient>
[Point 1: Root] Root confirmed at /workspace/api-gateway (pyproject.toml found).
[Point 3: Type signatures] UserCreate.tax_id: str = Field(..., min_length=11, max_length=14).
…
</orient>
<plan>
[Point 11: Atomic steps] Step 1: add a field_validator to the schema. Step 2: run tests/test_users.py.
…
</plan>
<patch>
<<<< SEARCH
    tax_id: str = Field(..., min_length=11, max_length=14)
====
    tax_id: str = Field(..., min_length=11, max_length=11)

    @field_validator("tax_id", mode="before")
    @classmethod
    def sanitize_tax_id(cls, v: str) -> str:
        cleaned = re.sub(r"\D", "", v)
        if len(cleaned) != 11:
            raise ValueError("tax_id must have exactly 11 digits")
        return cleaned
>>>> REPLACE
</patch>
<validate>
[Point 32: Targeted tests] pytest tests/test_users.py -k "tax_id" -> 2 passed
…
</validate>
<deliver>
…
</deliver>
</simplicio_loop>
```

The SEARCH/REPLACE markers are four characters long (`<<<<`, `====`, `>>>>`), not the seven that Git and Aider use. To apply a patch, find the SEARCH text verbatim in the file and replace it.

## Training

The published adapter was produced by [`Simplicio_27B_Training_Colab.ipynb`](https://github.com/simpletibr/simplicio-27b/blob/main/Simplicio_27B_Training_Colab.ipynb). The settings below come from that notebook and the published `adapter_config.json`.

| Setting | Value |
|---|---|
| Method | QLoRA with Unsloth; base loaded in 4-bit |
| LoRA | r 32, alpha 32, dropout 0, on the q, k, v, o, gate, up and down projections of every layer |
| Steps | 120 steps, batch size 1, gradient accumulation 8 |
| Optimizer | AdamW 8-bit, learning rate 2e-4, cosine schedule, 10 warmup steps, weight decay 0.01 |
| Sequence length | 4,096 |
| Loss | whole sequence, prompt included |
| Data | 101 examples in Portuguese: 1 written by hand and 100 generated from a short list of stack and task templates |
| Hardware | Google Colab A100 (40 GB) |

[`train_simplicio_27b.py`](https://github.com/simpletibr/simplicio-27b/blob/main/train_simplicio_27b.py) is a script version with extra options: freezing the bottom layers, attention-only LoRA, and registering the phase tags as special tokens. The published adapter used none of them, and its tokenizer has no added tokens. [`generate_dataset.py`](https://github.com/simpletibr/simplicio-27b/blob/main/generate_dataset.py) writes `data/simplicio_loop_50pts_train.jsonl` (80 examples) and `data/simplicio_loop_50pts_val.jsonl` (15 examples).

## Limitations

- It passes 46.7% of the held-out tasks, and only 6 of 30 in both the edge-case and the adversarial categories.
- The training set is small (101 examples), templated and in Portuguese. The model follows the format more reliably than it solves the task.
- In the training examples, `<validate>` and `<deliver>` contain written-out results such as "2 passed" or "COMMIT_READY". The model writes these without running anything. Treat them as claims and run your own tests.
- It has not been compared with the base model under the same protocol, and it has not been run on public benchmarks.
- Aider: the patch markers differ from Aider's edit format, and Aider has not been tested.
- Vision: the GGUF ships the base model's vision projector. Training was text-only, and image input has not been evaluated.

## Repository

| Path | Contents |
|---|---|
| `Simplicio_27B_Training_Colab.ipynb` | Training run that produced the adapter |
| `Simplicio_27B_Merge_Colab.ipynb` | Merges the adapter into 16-bit weights and exports the GGUF Q4_K_M |
| `Simplicio_27B_Serve_Colab.ipynb`, `deploy/` | vLLM serving, chat template, tool parser, context length, Ollama `Modelfile` |
| `data/unseen_eval_120.json` | The 120 held-out tasks |
| `benchmarks/live_colab_g4_bf16_n120.json` | The results above |
| `tests/` | Tests for the serving code: `python -m pytest tests` |

## Citation

```bibtex
@misc{simplicio27b2026,
  author       = {Simplicio, Wesley},
  title        = {Simplicio 27B: a Qwen3.8-27B fine-tune for SEARCH/REPLACE code patches},
  year         = {2026},
  publisher    = {SimpleTI},
  howpublished = {\url{https://huggingface.co/wesleysimplicio/Simplicio-27B}}
}
```

## Links

- Product page: [simpleti.com.br/simplicio-27b/](https://simpleti.com.br/simplicio-27b/)
- Weights: [huggingface.co/wesleysimplicio/Simplicio-27B](https://huggingface.co/wesleysimplicio/Simplicio-27B)
- Ollama: [ollama.com/wesleysimplicio/simplicio-27b](https://ollama.com/wesleysimplicio/simplicio-27b)
- Source: [github.com/simpletibr/simplicio-27b](https://github.com/simpletibr/simplicio-27b)
- Base model: [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) (Apache 2.0)
- Fine-tuning: [Unsloth](https://github.com/unslothai/unsloth)