|
Download README.md from KBBridge/KBBridge-v3: direct link, hf CLI and curl.
- Browser
- Download file 13.2 kB
-
https://huggingface.co/KBBridge/KBBridge-v3/resolve/main/README.md
- Command line
-
hf download hf://KBBridge/KBBridge-v3/README.md
-
curl -L -o README.md https://huggingface.co/KBBridge/KBBridge-v3/resolve/main/README.md
13.2 kB
| license: apache-2.0 | |
| base_model: Qwen/Qwen3.8-27B | |
| base_model_relation: finetune | |
| tags: [genexus, code-generation, qwen3] | |
| language: [es, en] | |
| pipeline_tag: text-generation | |
| # KBBridge-v3 (bf16) | |
| A fine-tune of [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B) specialised in | |
| **GeneXus** programming, in the native `.gxSource` export format. | |
| Frontier models do not know this format. Without the GeneXus documentation injected into the | |
| prompt they produce syntactically invalid output almost every time (parse rate 0.5–3.1%). | |
| KBBridge writes it natively, runs on your own hardware, and never sends your Knowledge Base | |
| code to an external API. | |
| --- | |
| ## ⚠️ Read this before your first prompt | |
| Three settings. All three are measured on this model, not stylistic. | |
| ### 1. Reasoning: it depends on the task, and the difference is large | |
| The chat template no longer pins a value, so **reasoning follows the upstream Qwen default | |
| (on)** unless your client turns it off. Which one you want depends on what you are asking for. | |
| All three rows below are measured on this model, not inherited from Qwen. | |
| | Task | Reasoning | Measured | | |
| |---|---|---| | |
| | **Writing `.gxSource`** | **OFF** | parseRate **86.9 → 33.5**, parmMatch 80.2 → 50.4 (191 held-out objects) | | |
| | Documentation multiple-choice | either | 78.4 vs 78.4 — no difference (329 items, McNemar p = 1.000) | | |
| | **Explaining existing code** | **ON** | fabricated claims **15.4% → 8.7%** (149 items, McNemar p = 0.041) | | |
| **If you generate GeneXus objects, turn reasoning off.** The collapse is real, not a budget | |
| artifact: with reasoning on, only 5.2% of items hit the token ceiling (fewer than the 8.4% | |
| without it) and 3.7% came back empty. The model simply writes worse `.gxSource` when it | |
| reasons first. Writing `.gxSource` is a formatting task. | |
| **If you point the model at existing code and ask what it does, turn reasoning on.** It nearly | |
| halves the rate at which the model asserts things the source does not support — the failure | |
| mode that matters when the output is documentation someone will trust. Cost: ~5× the output | |
| tokens. | |
| > **Correction (2026-09-02).** An earlier version of this card reported MCQ dropping 78.1 → 69.6 | |
| > with reasoning on. That number was wrong: our benchmark harness capped multiple-choice answers | |
| > at 512 tokens, which is not enough for a reasoning block to close, so the run was measuring the | |
| > cap rather than the model. Re-measured with an adequate budget, the difference is zero. The | |
| > `.gxSource` degradation is real and reproduced above with the current scorer. | |
| #### How to turn it off | |
| ```bash | |
| # vLLM — pass it explicitly on every request | |
| curl .../v1/chat/completions -d '{ | |
| "model": "...", "messages": [...], | |
| "chat_template_kwargs": {"enable_thinking": false} | |
| }' | |
| ``` | |
| > **Serve it with the flag, or you will think the model is broken.** If you run vLLM with | |
| > `--reasoning-parser qwen3` and the request does **not** carry `enable_thinking`, the parser | |
| > assumes reasoning is on, never finds the closing `</think>`, and routes the **entire answer** | |
| > into `reasoning`, leaving `content: null`. Every standard OpenAI client then shows an empty | |
| > reply. This only affects non-streaming requests — streaming takes a different path in vLLM and | |
| > looks fine — which makes it doubly confusing. Either pass `enable_thinking` on every request, | |
| > or drop `--reasoning-parser` and let the tags through. | |
| To pin reasoning off for every client instead, add this as the first line of | |
| `chat_template.jinja`: | |
| ```jinja | |
| {%- set enable_thinking = false %} | |
| ``` | |
| A `set` at the top of the template overrides anything the caller passes, which is a blunt but | |
| reliable way to guarantee behaviour across runtimes. | |
| ### 2. Ask for the format explicitly | |
| Write **"in `.gxSource` format"** in your prompt. | |
| Measured on v3: the bare request *"a Procedure that adds two numbers"* returns generic **SQL**. | |
| Naming the format returns the GeneXus object, consistently. If you use a harness with its own | |
| system prompt, put the instruction there once. | |
| ### 3. Give it enough room | |
| `max_tokens` ≥ 4096. A `.gxSource` object consumes roughly **340 tokens per KB** of source, and | |
| most tools default to 512–1024, which truncates the object mid-body. | |
| --- | |
| ## Results | |
| 580 held-out items (191 codegen + 329 MCQ + 60 data-model) that no model saw during training. | |
| Syntax validated with the official GeneXus ANTLR parser. Same protocol for every model: | |
| temperature 0.1, reasoning off, concurrency 8. | |
| ### v3 vs v2 — an honest comparison | |
| **v3 is not a clean win over v2.** It gains domain knowledge and loses syntax accuracy: | |
| | Metric | v2 | **v3** | | | |
| |---|---|---|---| | |
| | parseRate (valid syntax) | **89.0** | 84.8 | −4.2 | | |
| | parmMatch (exact signature) | 78.6 | 78.6 | = | | |
| | MCQ (GeneXus knowledge) | 76.0 | **79.0** | +3.0 | | |
| | methodValidity | 90.0 | **91.1** | +1.1 | | |
| **What these numbers do NOT establish.** v3 changed three things at once — the base model | |
| (Qwen3.6 → 3.8), the corpus (4× larger, per-KB cap removed) and the teacher (v1 → v2). The | |
| parseRate drop **cannot be attributed** to any one of them without a control arm that was never | |
| run. Anyone reading this table as "the bigger corpus hurt syntax" is over-reading it. | |
| Choose v3 if domain knowledge matters more to you; v2 still leads on raw syntax validity. | |
| ### Generalisation to unseen Knowledge Bases | |
| Three entire KBs were held out — different domains, never in the pipeline: | |
| | | held-out from training KBs | 3 completely new KBs | | |
| |---|---|---| | |
| | v2 | 89.0 | 89.9 | | |
| | **v3** | 84.8 | **87.4** | | |
| v3's *relative* gap to unseen KBs is larger than v2's (+2.6 vs +0.9), i.e. it generalises | |
| better in relative terms, even though two KBs make up 54.7% of its corpus. | |
| ### Fairness note on the frontier comparison | |
| In our benchmark the frontier models were run **with** ~21,600 tokens of GeneXus documentation | |
| injected into every request; KBBridge was run **without** any. That is not a handicap we | |
| imposed — injecting the same documentation into KBBridge makes it *worse* (76.4 → 73.3 | |
| parseRate), because the fine-tune already internalised that knowledge and the extra context | |
| gets in the way. Still, the setups differ, and you should know that when reading any | |
| head-to-head number. | |
| ### Quantised builds | |
| We measured the 4-bit build against this one on the same 580 items. **Excluding items where | |
| either run hit the token ceiling, the two are indistinguishable** (parseRate 93.0 vs 93.6 over | |
| 171 items) — 4-bit costs essentially nothing in output quality here. Details and the full | |
| comparison are in the | |
| [GGUF repo's card](https://huggingface.co/KBBridge/KBBridge-v3-GGUF). | |
| --- | |
| ## Files | |
| Full-precision merged weights, bf16, **51 GB** across 19 shards. This is the master artefact: | |
| use it to re-quantise, to continue training, or to serve with transformers. | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| m = AutoModelForCausalLM.from_pretrained("KBBridge/KBBridge-v3", dtype="bfloat16", device_map="auto") | |
| t = AutoTokenizer.from_pretrained("KBBridge/KBBridge-v3") | |
| ``` | |
| For serving, prefer [`KBBridge/KBBridge-v3-FP8`](https://huggingface.co/KBBridge/KBBridge-v3-FP8) | |
| (29 GB, same quality in our tests) or the | |
| [GGUF builds](https://huggingface.co/KBBridge/KBBridge-v3-GGUF) for llama.cpp / LM Studio. | |
| ### What is inside | |
| 1,199 tensors: the 64-layer hybrid text model (48 Gated DeltaNet + 16 full-attention layers), | |
| the base model's **vision tower** (333 tensors, carried over unchanged — the fine-tune did not | |
| touch it) and its **multi-token-prediction head** (15 tensors, likewise unchanged). Context | |
| 262,144 tokens, the base model's native `max_position_embeddings`. | |
| ## Intended use | |
| Assisting GeneXus developers: generating objects (Procedures, Transactions, Data Providers, | |
| SDTs, WebPanels), explaining existing code, completion, and documentation questions. | |
| **Out of scope:** not a general-purpose model, not a replacement for validating in the GeneXus | |
| IDE, and it does not know any particular Knowledge Base (see *Limitations*). | |
| --- | |
| ## Limitations | |
| - **It does not know your KB.** It learned the style and syntax of the format, not the contents | |
| of any specific base. Ask it about a transaction you did not paste in, and it will **invent | |
| plausible attribute names and present them as fact**. Always give it the context and validate | |
| the output in the IDE. | |
| - **Runaway generation on very large objects.** For objects over ~10 KB the model can fall into | |
| degenerate repetition — the same line hundreds of times without closing the object. Measured | |
| on v2 at ~1.6% of benchmark items; **not re-measured on v3**. Raising `max_tokens` does not | |
| fix it. Generate large objects section by section. | |
| - **Spanish bias** in explanations, reflecting the corpus. | |
| - **Specialised**: worse than the base model at general tasks. | |
| - The limitations above other than the first were measured on **v2** and are carried over as | |
| working assumptions, not verified properties of v3. | |
| ### If you also use a hosted KBBridge endpoint | |
| The raw GGUF and a gateway-fronted deployment **do not behave the same by default**. Our | |
| gateway applies five corrections the plain model does not have: a `max_tokens` floor, | |
| reasoning off unless the client asks for it, `temperature` defaulted to 0.2 (without it vLLM | |
| falls back to the checkpoint's `generation_config`, which is **1.0**), a fallback that recovers | |
| the answer from the `reasoning` field when `content` comes back empty, and `repetition_penalty` | |
| 1.05 to suppress runaway. If you compare "what I tried on your server" against "what I | |
| downloaded", the difference is those five settings, not the weights. | |
| The temperature one surprises people: the OpenAI standard makes the field optional and many | |
| clients never send it, so an unconfigured client is sampling at 1.0 without being told. | |
| --- | |
| ## Training | |
| | | | | |
| |---|---| | |
| | Method | QLoRA 4-bit (bitsandbytes) + Liger kernel | | |
| | LoRA | r=64, α=128, dropout=0.05, all projections | | |
| | Context | 12,288 tokens | | |
| | Effective batch | 16 (1 × 16 grad accum) | | |
| | LR | 1.0e-4, cosine, 3% warmup | | |
| | Epochs | 2 complete (14,108 steps) | | |
| | Hardware | 1× RTX PRO 6000 Blackwell 96 GB | | |
| | Duration | 7 days 4:41 | | |
| | Framework | LLaMA-Factory, transformers 5.6.0 | | |
| train_loss **0.2618** (v2: 0.3344) · eval_loss **0.3723** (v2: 0.4675), minimum at the **last** | |
| step — no overfitting across 71 evaluations, which suggests there was room for more epochs. | |
| Note that these losses are much better than v2's and yet parseRate went *down*: `eval_loss` | |
| measures fit to the corpus, not GeneXus quality. | |
| ### Data | |
| 80,344 examples derived from GeneXus objects across 25 real Knowledge Bases (GX16/17/17U8/18/ | |
| Evo1, multi-domain) — 129% more than v2, with the per-KB cap removed. Sanitised, deduplicated | |
| and split by deterministic hash. **The datasets are not published**: they contain customer | |
| proprietary code. | |
| --- | |
| ## Training-data privacy | |
| The model was trained on real customer Knowledge Bases, so we audited whether it can leak them. | |
| This is the strongest result of the project. | |
| ### Canaries: no memorisation threshold found | |
| 12 synthetic objects containing unguessable 16-character secrets were inserted at four | |
| frequencies, and verified to have reached `train.jsonl` at exactly those counts: | |
| | repetitions | canaries | recovered by name | recovered with literal prefix | | |
| |---|---|---|---| | |
| | 1 | 3 | 0/3 | 0/3 | | |
| | 10 | 3 | 0/3 | 0/3 | | |
| | 100 | 3 | 0/3 | 0/3 | | |
| | **1000** | 3 | **0/3** | **0/3** | | |
| **Not even at a thousand identical repetitions.** A control rules out a broken probe: asked for | |
| the canary, the model returns a structurally valid but **empty** object — no token, no secret. | |
| And it does generate real bodies when the request has content, so the empty skeleton is not an | |
| inability to generate. | |
| ### Membership inference: marginal signal | |
| | | | | |
| |---|---| | |
| | mean loss, seen examples | 3.4130 | | |
| | mean loss, unseen | 3.7711 | | |
| | mean length | 3,133 vs 3,117 chars — comparable, so the AUC is meaningful | | |
| | **AUC** | **0.5539** | | |
| 0.554 against 0.50 for indistinguishable. There is a statistical trace of having seen the data, | |
| but the distributions overlap almost entirely. | |
| **Conclusion: customer code is not recoverable from the weights.** | |
| **Caveat, stated plainly:** absence of evidence is not proof of absence. These audits cover the | |
| attacks we ran, not every attack that exists. | |
| --- | |
| ## Reproducibility | |
| Full external reproduction is **not possible**, and it is worth saying so directly: | |
| 1. The 25 Knowledge Bases are customer code and are not distributed. | |
| 2. The `parseRate` scorer uses the KBEditor's ANTLR parser — proprietary, not distributable. | |
| 3. The teacher that generated v3's data is KBBridge-v2, which is not published. | |
| What a third party *can* verify: the raw benchmark outputs (one model response per item) and | |
| the scoring over them. | |
| --- | |
| ## Citation | |
| ```bibtex | |
| @misc{kbbridge-v3, | |
| title = {KBBridge-v3: a GeneXus code assistant fine-tuned from Qwen3.8-27B}, | |
| author = {{KBBridge}}, | |
| year = {2026}, | |
| url = {https://huggingface.co/KBBridge/KBBridge-v3} | |
| } | |
| ``` | |
| ## License | |
| Apache 2.0, inherited from the base model `Qwen/Qwen3.8-27B`. This is a modified derivative | |
| work; see `NOTICE`. | |