securecoder-30b-pro / README.md
Taimwe's picture
Honest model card (LoRA + eval)
c7f5fea verified
|
Raw History Blame Contribute Delete
6.42 kB
---
base_model: unsloth/Qwen3-Coder-30B-A3B-Instruct
license: apache-2.0
pipeline_tag: text-generation
tags:
- lora
- peft
- qwen3
- code
- tool-calling
- security
- securecoder
---
# SecureCoder-30B-Pro (LoRA adapter)
A QLoRA fine-tune of `unsloth/Qwen3-Coder-30B-A3B-Instruct` (30B-A3B MoE, Apache-2.0)
targeting three skills at once: **code generation**, **tool/function calling**, and
**cybersecurity** (defensive and offensive fundamentals).
For the 16-bit merged model, see [`Taimwe/securecoder-30b-pro-merged`](https://huggingface.co/Taimwe/securecoder-30b-pro-merged).
For a Q4_K_M GGUF ready for Ollama / LM Studio / llama.cpp, see
[`Taimwe/securecoder-30b-pro-GGUF`](https://huggingface.co/Taimwe/securecoder-30b-pro-GGUF).
For the training/merge/quantise/eval scripts, see
[`Taimwe/securecoder-scripts`](https://huggingface.co/Taimwe/securecoder-scripts).
## Summary
| | |
| --- | --- |
| Base | [`unsloth/Qwen3-Coder-30B-A3B-Instruct`](https://huggingface.co/unsloth/Qwen3-Coder-30B-A3B-Instruct) |
| Method | QLoRA (4-bit base, 16-bit LoRA), rank 32, attention projections only |
| Trainable params | 26,738,688 (0.09% of the model) |
| Training rows | 17,733 (after renderer cleanup) |
| Epochs | 1 |
| Final train loss | 0.598 |
| Final eval loss | 0.517 |
| Cost | $6.84 on a100-large (~$10 total for the whole pipeline incl. merge + GGUF + eval) |
## Data mix
The mix cap per source *is* the recipe. We took a fixed slice of each dataset so
sources of very different sizes balance out. Validated on a remote CPU job (50 s
for 19,470 row sample, $0.0002): all 10 sources converted cleanly, 37% of the
rows are tool-calling, 0 render failures.
| Source | Rows taken | What it teaches |
| --- | --- | --- |
| `NousResearch/hermes-function-calling-v1` `[func_calling]` | 9,000 | full tool-call conversations + JSON schemas |
| `NousResearch/hermes-function-calling-v1` `[func_calling_singleturn]` | 3,000 | single-turn tool selection |
| `lockon/xlam-function-calling-60k` | 10,000 | 60k API-call pairs (query → call) |
## Measured numbers (real, not made up)
| Section | Metric | Value | n |
| --- | --- | --- | --- |
| **tool_calls** | parse_rate (block well-formed) | **100.0%** | 60 |
| **tool_calls** | name_rate (correct function name) | **100.0%** | 60 |
| **tool_calls** | schema_rate (arguments match schema) | **98.3%** | 60 |
| **code_sanity** | ast_rate (Python parses) | **6.7%** | 15 |
| **code_sanity** | compile_rate (Python compiles) | **6.7%** | 15 |
| **security_mcq** | accuracy | *skipped* (dataset split unavailable) | 0 |
The eval set:
- 60 tool-calling prompts streamed from `NousResearch/hermes-function-calling-v1` (`func_calling` config, `train` split)
- 15 hand-written Python coding prompts (is_palindrome, merge_intervals, two_sum, flatten, parse_csv_line, lru_cache, is_anagram, topological_order, tokenise, slugify, read_jsonl, binary_search, unique_in_order, safe_eval, dedupe_preserve_order)
- 25 CyberSecurityEval MCQs — `train` split no longer fetchable; the section is skipped cleanly with a clear note in `eval-report.json`
Full per-prompt details are in `eval-report.json` in this repo.
### Honest reading
- **Tool calling is genuinely strong** — 100% well-formed, 100% correct function name, 98.3% schema-correct on the first 60 prompts. This is the headline number.
- **Code generation on hard prompts is weak** (6.7%). Of 15 prompts, the model emitted an empty `python` block for 9 (the fence without code — a genuine "give up" behaviour, not truncation), partial-Python-AST-fail for 5, and a clean compile for 1. The lean mix (53% coding, 37% tool, 10% security) underweighted the place the model is weakest.
- **Cybersecurity is *not measured***. The training data is there; the eval is not. Treat any claims about cybersecurity skill as untested.
## Limitations
- No safety tuning. The training mix contains recon/enumeration material; the
base has no alignment layer; the fine-tune adds none. **Use the outputs with
care.** Review generated code before running it.
- 2k context window (inherited from Qwen3-Coder-30B).
- The Q4_K_M GGUF is lossy compared to the 16-bit merge; for best tool calling
fidelity, use the merged safetensors or load the LoRA on the base in 4-bit.
- Lean mix run: 1 epoch, ~19k effective rows. For higher code-quality, a fuller
coding-heavy mix + 2+ epochs is the obvious next step.
## Loading
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained(
"unsloth/Qwen3-Coder-30B-A3B-Instruct",
torch_dtype=torch.bfloat16, device_map="auto",
load_in_4bit=True,
)
model = PeftModel.from_pretrained(base, "Taimwe/securecoder-30b-pro")
tokenizer = AutoTokenizer.from_pretrained("Taimwe/securecoder-30b-pro")
```
## Reproduce
Scripts in [`Taimwe/securecoder-scripts`](https://huggingface.co/Taimwe/securecoder-scripts):
```bash
# train (commit-pinned URL form to avoid drift)
hf jobs run -d --flavor a100-large --timeout 12h --secrets HF_TOKEN \
ghcr.io/astral-sh/uv:python3.12-bookworm \
uv run --no-project \
https://huggingface.co/Taimwe/securecoder-scripts/raw/<SHA>/train_securecoder.py \
-- --num-epochs 1 --mix-scale 0.33 --max-seq-length 2048 \
--output-repo Taimwe/securecoder-30b-pro --private
```
## Provenance
- Hub user: Taimwe
- Training script commit: see `Taimwe/securecoder-scripts` history
- Uploaded with: `model.push_to_hub()` via Unsloth
| `ise-uiuc/Magicoder-OSS-Instruct-75K` | 10,000 | self-instruct code problems + solutions |
| `Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset` | 8,000 | security instruction tuning |
| `AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1` | 5,000 | broad security Q&A |
| `Humanlearning/CyberSecurity_OWASP-sft-dataset` | 3,000 | OWASP / secure-coding SFT |
| `MrClipperz134/CTF-Instruct` | 3,000 | CTF challenge → solve |
| `TrueNix/ctf-solver-dataset` | 3,000 | CTF solving trajectories |
| `mlabonne/FineTome-100k` | 3,000 | general instruct replay |
Excluded deliberately: anything whose purpose is building malware or
weaponised exploits. Recon/enumeration knowledge, exploit *concepts*, CTF
solving and defensive engineering are in; end-to-end attack tooling is not.