File size: 13,228 Bytes
b86771b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
82ac291
b86771b
82ac291
b86771b
82ac291
 
 
b86771b
82ac291
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b86771b
82ac291
 
 
 
 
 
 
986da38
82ac291
 
 
 
 
 
b86771b
82ac291
 
b86771b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
82ac291
 
 
 
 
 
 
 
 
b86771b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
503bd49
b86771b
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
---
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: finetune
tags: [genexus, code-generation, qwen3]
language: [es, en]
pipeline_tag: text-generation
---
# KBBridge-v3 (bf16)

A fine-tune of [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B) specialised in
**GeneXus** programming, in the native `.gxSource` export format.

Frontier models do not know this format. Without the GeneXus documentation injected into the
prompt they produce syntactically invalid output almost every time (parse rate 0.5–3.1%).
KBBridge writes it natively, runs on your own hardware, and never sends your Knowledge Base
code to an external API.

---

## ⚠️ Read this before your first prompt

Three settings. All three are measured on this model, not stylistic.

### 1. Reasoning: it depends on the task, and the difference is large

The chat template no longer pins a value, so **reasoning follows the upstream Qwen default
(on)** unless your client turns it off. Which one you want depends on what you are asking for.
All three rows below are measured on this model, not inherited from Qwen.

| Task | Reasoning | Measured |
|---|---|---|
| **Writing `.gxSource`** | **OFF** | parseRate **86.9 β†’ 33.5**, parmMatch 80.2 β†’ 50.4 (191 held-out objects) |
| Documentation multiple-choice | either | 78.4 vs 78.4 β€” no difference (329 items, McNemar p = 1.000) |
| **Explaining existing code** | **ON** | fabricated claims **15.4% β†’ 8.7%** (149 items, McNemar p = 0.041) |

**If you generate GeneXus objects, turn reasoning off.** The collapse is real, not a budget
artifact: with reasoning on, only 5.2% of items hit the token ceiling (fewer than the 8.4%
without it) and 3.7% came back empty. The model simply writes worse `.gxSource` when it
reasons first. Writing `.gxSource` is a formatting task.

**If you point the model at existing code and ask what it does, turn reasoning on.** It nearly
halves the rate at which the model asserts things the source does not support β€” the failure
mode that matters when the output is documentation someone will trust. Cost: ~5Γ— the output
tokens.

> **Correction (2026-09-02).** An earlier version of this card reported MCQ dropping 78.1 β†’ 69.6
> with reasoning on. That number was wrong: our benchmark harness capped multiple-choice answers
> at 512 tokens, which is not enough for a reasoning block to close, so the run was measuring the
> cap rather than the model. Re-measured with an adequate budget, the difference is zero. The
> `.gxSource` degradation is real and reproduced above with the current scorer.

#### How to turn it off

```bash
# vLLM β€” pass it explicitly on every request
curl .../v1/chat/completions -d '{
  "model": "...", "messages": [...],
  "chat_template_kwargs": {"enable_thinking": false}
}'
```

> **Serve it with the flag, or you will think the model is broken.** If you run vLLM with
> `--reasoning-parser qwen3` and the request does **not** carry `enable_thinking`, the parser
> assumes reasoning is on, never finds the closing `</think>`, and routes the **entire answer**
> into `reasoning`, leaving `content: null`. Every standard OpenAI client then shows an empty
> reply. This only affects non-streaming requests β€” streaming takes a different path in vLLM and
> looks fine β€” which makes it doubly confusing. Either pass `enable_thinking` on every request,
> or drop `--reasoning-parser` and let the tags through.

To pin reasoning off for every client instead, add this as the first line of
`chat_template.jinja`:

```jinja
{%- set enable_thinking = false %}
```

A `set` at the top of the template overrides anything the caller passes, which is a blunt but
reliable way to guarantee behaviour across runtimes.

### 2. Ask for the format explicitly

Write **"in `.gxSource` format"** in your prompt.

Measured on v3: the bare request *"a Procedure that adds two numbers"* returns generic **SQL**.
Naming the format returns the GeneXus object, consistently. If you use a harness with its own
system prompt, put the instruction there once.

### 3. Give it enough room

`max_tokens` β‰₯ 4096. A `.gxSource` object consumes roughly **340 tokens per KB** of source, and
most tools default to 512–1024, which truncates the object mid-body.

---

## Results

580 held-out items (191 codegen + 329 MCQ + 60 data-model) that no model saw during training.
Syntax validated with the official GeneXus ANTLR parser. Same protocol for every model:
temperature 0.1, reasoning off, concurrency 8.

### v3 vs v2 β€” an honest comparison

**v3 is not a clean win over v2.** It gains domain knowledge and loses syntax accuracy:

| Metric | v2 | **v3** | |
|---|---|---|---|
| parseRate (valid syntax) | **89.0** | 84.8 | βˆ’4.2 |
| parmMatch (exact signature) | 78.6 | 78.6 | = |
| MCQ (GeneXus knowledge) | 76.0 | **79.0** | +3.0 |
| methodValidity | 90.0 | **91.1** | +1.1 |

**What these numbers do NOT establish.** v3 changed three things at once β€” the base model
(Qwen3.6 β†’ 3.8), the corpus (4Γ— larger, per-KB cap removed) and the teacher (v1 β†’ v2). The
parseRate drop **cannot be attributed** to any one of them without a control arm that was never
run. Anyone reading this table as "the bigger corpus hurt syntax" is over-reading it.

Choose v3 if domain knowledge matters more to you; v2 still leads on raw syntax validity.

### Generalisation to unseen Knowledge Bases

Three entire KBs were held out β€” different domains, never in the pipeline:

| | held-out from training KBs | 3 completely new KBs |
|---|---|---|
| v2 | 89.0 | 89.9 |
| **v3** | 84.8 | **87.4** |

v3's *relative* gap to unseen KBs is larger than v2's (+2.6 vs +0.9), i.e. it generalises
better in relative terms, even though two KBs make up 54.7% of its corpus.

### Fairness note on the frontier comparison

In our benchmark the frontier models were run **with** ~21,600 tokens of GeneXus documentation
injected into every request; KBBridge was run **without** any. That is not a handicap we
imposed β€” injecting the same documentation into KBBridge makes it *worse* (76.4 β†’ 73.3
parseRate), because the fine-tune already internalised that knowledge and the extra context
gets in the way. Still, the setups differ, and you should know that when reading any
head-to-head number.

### Quantised builds

We measured the 4-bit build against this one on the same 580 items. **Excluding items where
either run hit the token ceiling, the two are indistinguishable** (parseRate 93.0 vs 93.6 over
171 items) β€” 4-bit costs essentially nothing in output quality here. Details and the full
comparison are in the
[GGUF repo's card](https://huggingface.co/KBBridge/KBBridge-v3-GGUF).

---

## Files

Full-precision merged weights, bf16, **51 GB** across 19 shards. This is the master artefact:
use it to re-quantise, to continue training, or to serve with transformers.

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("KBBridge/KBBridge-v3", dtype="bfloat16", device_map="auto")
t = AutoTokenizer.from_pretrained("KBBridge/KBBridge-v3")
```

For serving, prefer [`KBBridge/KBBridge-v3-FP8`](https://huggingface.co/KBBridge/KBBridge-v3-FP8)
(29 GB, same quality in our tests) or the
[GGUF builds](https://huggingface.co/KBBridge/KBBridge-v3-GGUF) for llama.cpp / LM Studio.

### What is inside

1,199 tensors: the 64-layer hybrid text model (48 Gated DeltaNet + 16 full-attention layers),
the base model's **vision tower** (333 tensors, carried over unchanged β€” the fine-tune did not
touch it) and its **multi-token-prediction head** (15 tensors, likewise unchanged). Context
262,144 tokens, the base model's native `max_position_embeddings`.

## Intended use

Assisting GeneXus developers: generating objects (Procedures, Transactions, Data Providers,
SDTs, WebPanels), explaining existing code, completion, and documentation questions.

**Out of scope:** not a general-purpose model, not a replacement for validating in the GeneXus
IDE, and it does not know any particular Knowledge Base (see *Limitations*).

---

## Limitations

- **It does not know your KB.** It learned the style and syntax of the format, not the contents
  of any specific base. Ask it about a transaction you did not paste in, and it will **invent
  plausible attribute names and present them as fact**. Always give it the context and validate
  the output in the IDE.
- **Runaway generation on very large objects.** For objects over ~10 KB the model can fall into
  degenerate repetition β€” the same line hundreds of times without closing the object. Measured
  on v2 at ~1.6% of benchmark items; **not re-measured on v3**. Raising `max_tokens` does not
  fix it. Generate large objects section by section.
- **Spanish bias** in explanations, reflecting the corpus.
- **Specialised**: worse than the base model at general tasks.
- The limitations above other than the first were measured on **v2** and are carried over as
  working assumptions, not verified properties of v3.

### If you also use a hosted KBBridge endpoint

The raw GGUF and a gateway-fronted deployment **do not behave the same by default**. Our
gateway applies five corrections the plain model does not have: a `max_tokens` floor,
reasoning off unless the client asks for it, `temperature` defaulted to 0.2 (without it vLLM
falls back to the checkpoint's `generation_config`, which is **1.0**), a fallback that recovers
the answer from the `reasoning` field when `content` comes back empty, and `repetition_penalty`
1.05 to suppress runaway. If you compare "what I tried on your server" against "what I
downloaded", the difference is those five settings, not the weights.

The temperature one surprises people: the OpenAI standard makes the field optional and many
clients never send it, so an unconfigured client is sampling at 1.0 without being told.

---

## Training

| | |
|---|---|
| Method | QLoRA 4-bit (bitsandbytes) + Liger kernel |
| LoRA | r=64, Ξ±=128, dropout=0.05, all projections |
| Context | 12,288 tokens |
| Effective batch | 16 (1 Γ— 16 grad accum) |
| LR | 1.0e-4, cosine, 3% warmup |
| Epochs | 2 complete (14,108 steps) |
| Hardware | 1Γ— RTX PRO 6000 Blackwell 96 GB |
| Duration | 7 days 4:41 |
| Framework | LLaMA-Factory, transformers 5.6.0 |

train_loss **0.2618** (v2: 0.3344) Β· eval_loss **0.3723** (v2: 0.4675), minimum at the **last**
step β€” no overfitting across 71 evaluations, which suggests there was room for more epochs.

Note that these losses are much better than v2's and yet parseRate went *down*: `eval_loss`
measures fit to the corpus, not GeneXus quality.

### Data

80,344 examples derived from GeneXus objects across 25 real Knowledge Bases (GX16/17/17U8/18/
Evo1, multi-domain) β€” 129% more than v2, with the per-KB cap removed. Sanitised, deduplicated
and split by deterministic hash. **The datasets are not published**: they contain customer
proprietary code.

---

## Training-data privacy

The model was trained on real customer Knowledge Bases, so we audited whether it can leak them.
This is the strongest result of the project.

### Canaries: no memorisation threshold found

12 synthetic objects containing unguessable 16-character secrets were inserted at four
frequencies, and verified to have reached `train.jsonl` at exactly those counts:

| repetitions | canaries | recovered by name | recovered with literal prefix |
|---|---|---|---|
| 1 | 3 | 0/3 | 0/3 |
| 10 | 3 | 0/3 | 0/3 |
| 100 | 3 | 0/3 | 0/3 |
| **1000** | 3 | **0/3** | **0/3** |

**Not even at a thousand identical repetitions.** A control rules out a broken probe: asked for
the canary, the model returns a structurally valid but **empty** object β€” no token, no secret.
And it does generate real bodies when the request has content, so the empty skeleton is not an
inability to generate.

### Membership inference: marginal signal

| | |
|---|---|
| mean loss, seen examples | 3.4130 |
| mean loss, unseen | 3.7711 |
| mean length | 3,133 vs 3,117 chars β€” comparable, so the AUC is meaningful |
| **AUC** | **0.5539** |

0.554 against 0.50 for indistinguishable. There is a statistical trace of having seen the data,
but the distributions overlap almost entirely.

**Conclusion: customer code is not recoverable from the weights.**

**Caveat, stated plainly:** absence of evidence is not proof of absence. These audits cover the
attacks we ran, not every attack that exists.

---

## Reproducibility

Full external reproduction is **not possible**, and it is worth saying so directly:

1. The 25 Knowledge Bases are customer code and are not distributed.
2. The `parseRate` scorer uses the KBEditor's ANTLR parser β€” proprietary, not distributable.
3. The teacher that generated v3's data is KBBridge-v2, which is not published.

What a third party *can* verify: the raw benchmark outputs (one model response per item) and
the scoring over them.

---

## Citation

```bibtex
@misc{kbbridge-v3,
  title  = {KBBridge-v3: a GeneXus code assistant fine-tuned from Qwen3.8-27B},
  author = {{KBBridge}},
  year   = {2026},
  url    = {https://huggingface.co/KBBridge/KBBridge-v3}
}
```

## License

Apache 2.0, inherited from the base model `Qwen/Qwen3.8-27B`. This is a modified derivative
work; see `NOTICE`.