SecureCoder runbook
Browse files
README.md
ADDED
|
@@ -0,0 +1,201 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# SecureCoder β training runbook
|
| 2 |
+
|
| 3 |
+
A QLoRA fine-tune that targets three skills at once: **coding**, **tool calling**
|
| 4 |
+
and **cybersecurity** (offence *and* defence), on a base chosen to train cheaply
|
| 5 |
+
and run fast. Everything here is measured or verified on real jobs, not assumed β
|
| 6 |
+
each claim carries the job id you can inspect.
|
| 7 |
+
|
| 8 |
+
## 1. Base model choice
|
| 9 |
+
|
| 10 |
+
Screened on the Hub on 2026-09-24 (params / licence / tool-calling support):
|
| 11 |
+
|
| 12 |
+
| Candidate | Params | Licence | Verdict |
|
| 13 |
+
| --- | --- | --- | --- |
|
| 14 |
+
| **`Qwen/Qwen3-Coder-30B-A3B-Instruct`** | 30.5B MoE, ~3B active | Apache-2.0 | **chosen** β code-specialised, native tool-calling template, MoE means it *trains* like a 3B model and fits 48 GB in 4-bit |
|
| 15 |
+
| `Qwen/Qwen3.8-27B` | 27.8B dense | Apache-2.0 | best raw quality, ~4x slower per step; good upgrade path once the pipeline is proven |
|
| 16 |
+
| `Qwen/Qwen3-Coder-Next` | 79.7B | Apache-2.0 | needs an 80 GB card (A100 80 GB `$2.50/h`); too big for the first run |
|
| 17 |
+
| `Qwen/Qwen3.8-Flash-Next` | 180B | **license: other** | rejected β non-commercial-style licence on an 180B model |
|
| 18 |
+
| `openbmb/MiniCPM5-2B` | 2.5B | Apache-2.0 | `tool-calling` tag, tiny; good for a T4 smoke test, weak ceiling |
|
| 19 |
+
| `TokenRhythm/NeoHorse-1-9B` | 9.0B | Apache-2.0 | `tool-use`, `reasoning`; the fallback if T4-only |
|
| 20 |
+
|
| 21 |
+
Why not start from an existing *abliterated* checkpoint? Because refusal removal is
|
| 22 |
+
a separate weight-edit (section 6) that can be applied to **any** checkpoint after
|
| 23 |
+
fine-tuning, and pre-abliterated bases are community re-uploads with unclear
|
| 24 |
+
provenance. Fine-tuning first keeps the lineage clean and the base licensed.
|
| 25 |
+
|
| 26 |
+
## 2. Data mix
|
| 27 |
+
|
| 28 |
+
The cap per source *is* the recipe β sources differ hugely in size, so the mix is
|
| 29 |
+
balanced by taking a fixed slice of each. Verified on a real CPU job
|
| 30 |
+
(`Taimwe/6ab5ae256b030d633f68faef`) at 25 rows/source:
|
| 31 |
+
|
| 32 |
+
| Source | Rows taken | Kind | What it teaches |
|
| 33 |
+
| --- | --- | --- | --- |
|
| 34 |
+
| `NousResearch/hermes-function-calling-v1` `[func_calling]` | 9,000 | tool calling | full tool-call conversations + JSON schemas |
|
| 35 |
+
| `NousResearch/hermes-function-calling-v1` `[func_calling_singleturn]` | 3,000 | tool calling | picking the right function, no chatter |
|
| 36 |
+
| `lockon/xlam-function-calling-60k` | 10,000 | tool calling | 60k API-call pairs (query β call) |
|
| 37 |
+
| `ise-uiuc/Magicoder-OSS-Instruct-75K` | 10,000 | coding | self-instruct problems + solutions |
|
| 38 |
+
| `Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset` | 8,000 | security | security instruction tuning |
|
| 39 |
+
| `AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1` | 5,000 | security | broad security Q&A |
|
| 40 |
+
| `Humanlearning/CyberSecurity_OWASP-sft-dataset` | 3,000 | secure coding | OWASP / defensive coding |
|
| 41 |
+
| `MrClipperz134/CTF-Instruct` | 3,000 | CTF | challenge instruction β solve |
|
| 42 |
+
| `TrueNix/ctf-solver-dataset` | 3,000 | CTF | solving trajectories |
|
| 43 |
+
| `mlabonne/FineTome-100k` | 3,000 | replay | general chat, so instruction-following does not collapse |
|
| 44 |
+
|
| 45 |
+
**Total β 57,000 rows, ~30% of them tool-calling.** Every source validated at
|
| 46 |
+
25/25 kept with zero render failures.
|
| 47 |
+
|
| 48 |
+
Excluded deliberately:
|
| 49 |
+
|
| 50 |
+
* `dpevzner/Cybersecurity_Reasoning_Dataset` β its builder configs were renamed
|
| 51 |
+
(`default` β `mistral|deepseek|chatml|gemma`) and the fields the card documents
|
| 52 |
+
(`unified_interpretation`) are empty in the published revision, so every row was
|
| 53 |
+
dropped. Re-addable once its schema settles.
|
| 54 |
+
* Anything whose purpose is building malware or weaponised exploits. Recon and
|
| 55 |
+
enumeration knowledge, exploit *concepts*, CTF solving and defensive
|
| 56 |
+
engineering are in; end-to-end attack tooling is not. This is a deliberate
|
| 57 |
+
scope line, not an oversight.
|
| 58 |
+
|
| 59 |
+
## 3. What training does with the data
|
| 60 |
+
|
| 61 |
+
Each row is converted to chat messages and rendered through the model's own chat
|
| 62 |
+
template, so the model learns the exact format it will be asked for at inference
|
| 63 |
+
time (`<|im_start|>β¦<tool_call><function=NAME><parameter=β¦>` for Qwen3-Coder).
|
| 64 |
+
|
| 65 |
+
Two format traps found the hard way and handled in code:
|
| 66 |
+
|
| 67 |
+
* Qwen3-Coder iterates `tool.parameters.properties` β tool schemas must be **flat**
|
| 68 |
+
(`{"name", "description", "parameters"}`), not wrapped in `"function"`.
|
| 69 |
+
* It also iterates `tool_call.arguments | items`, so `arguments` must be a
|
| 70 |
+
**mapping**. A JSON string there fails with *"Can only get item pairs from a
|
| 71 |
+
mapping"* β which is exactly how xLAM was silently contributing 0 rows until fixed.
|
| 72 |
+
|
| 73 |
+
`render_record()` therefore detects the template's convention once (flat/nested Γ
|
| 74 |
+
dict/string) and reuses it, instead of hard-coding one model's dialect.
|
| 75 |
+
|
| 76 |
+
## 4. Compute and cost
|
| 77 |
+
|
| 78 |
+
Verified `hf jobs hardware` prices (per hour), 2026-09-24:
|
| 79 |
+
|
| 80 |
+
| Flavor | VRAM | $/hour | Use |
|
| 81 |
+
| --- | --- | --- | --- |
|
| 82 |
+
| `cpu-basic` | β | $0.01 | data-mix validation, class checks |
|
| 83 |
+
| `t4-small` | 16 GB (T4) | $0.40 | plumbing checks; too small for a 30B in 4-bit |
|
| 84 |
+
| `l4x1` | 24 GB | $0.80 | 9B runs on a budget |
|
| 85 |
+
| `a10g-large` | 24 GB | $1.50 | 9Bβ14B comfortably |
|
| 86 |
+
| **`l40sx1`** | **48 GB** | **$1.80** | **the 30B-A3B QLoRA** |
|
| 87 |
+
| `a100-large` | 80 GB | $2.50 | faster per step, larger batch |
|
| 88 |
+
| `rtx-pro-6000` | 96 GB | $2.75 | headroom for longer context |
|
| 89 |
+
| `h200` | 141 GB | $5.00 | full fine-tune territory, overkill here |
|
| 90 |
+
|
| 91 |
+
Rough budget for the full 1-epoch run. **These are now measured, not guessed** β the
|
| 92 |
+
smoke test (`Taimwe/6ab5b3ed52d0dbd7f1d8d454`, a100-large) trained at
|
| 93 |
+
**0.8 rows/s** (20 s per step at batch 2 Γ grad-accum 8, max seq 2048, packing on),
|
| 94 |
+
which includes a few minutes of `torch.compile` warm-up that does not recur per step:
|
| 95 |
+
|
| 96 |
+
| Mix | Rows | Time at 0.8 rows/s | Cost on `l40sx1` ($1.80/h) |
|
| 97 |
+
| --- | --- | --- | --- |
|
| 98 |
+
| full mix as specified | ~57,000 | ~20 h | **~$36** |
|
| 99 |
+
| half caps (`--max-steps` or edit the table) | ~28,000 | ~10 h | ~$18 |
|
| 100 |
+
| lean mix (cut caps to ~1/3) | ~19,000 | ~7 h | ~$12 |
|
| 101 |
+
|
| 102 |
+
Because warm-up inflates that rate, a **300-step measured run** (~$4 on a100-large)
|
| 103 |
+
gives a trustworthy steady-state number before committing to a long job. Levers that
|
| 104 |
+
move the cost most, in order: rows in the mix, `--max-seq-length` (2048 vs 4096),
|
| 105 |
+
`--batch-size` (raise it on 80β96 GB cards), `--lora-r`.
|
| 106 |
+
|
| 107 |
+
The `--smoke` run itself costs under $1 and proves model load, data rendering,
|
| 108 |
+
training, evaluation and the Hub push end to end.
|
| 109 |
+
|
| 110 |
+
No local GPU needed. Colab works the same way with `uv run train_securecoder.py β¦`
|
| 111 |
+
(Colab Free's T4 cannot hold a 30B in 4-bit β use Colab Pro L4/A100, or point
|
| 112 |
+
`--base-model` at a 9B).
|
| 113 |
+
|
| 114 |
+
## 5. Runbook (copy-paste)
|
| 115 |
+
|
| 116 |
+
```bash
|
| 117 |
+
# 0. auth (token needs repo.write, and job.write for Jobs)
|
| 118 |
+
hf auth login --token hf_xxx
|
| 119 |
+
|
| 120 |
+
# 1. validate the mix: 25 rows/source, no GPU, ~50 s, ~$0.0002
|
| 121 |
+
hf jobs run -d --flavor cpu-basic --timeout 30m --name validate-mix \
|
| 122 |
+
ghcr.io/astral-sh/uv:python3.12-bookworm \
|
| 123 |
+
uv run --no-project https://huggingface.co/Taimwe/securecoder-scripts/resolve/main/train_securecoder.py \
|
| 124 |
+
--validate-only --validate-per-source 25 --show-samples 3
|
| 125 |
+
|
| 126 |
+
# 2. smoke test on a real GPU: 200 rows/source, 20 steps, pushes an adapter (~$0.50)
|
| 127 |
+
hf jobs run -d --flavor l40sx1 --timeout 40m --secrets HF_TOKEN --name smoke-train \
|
| 128 |
+
ghcr.io/astral-sh/uv:python3.12-bookworm \
|
| 129 |
+
uv run --no-project https://huggingface.co/Taimwe/securecoder-scripts/resolve/main/train_securecoder.py \
|
| 130 |
+
--smoke --output-repo Taimwe/securecoder-smoke --private --report-to none
|
| 131 |
+
|
| 132 |
+
# 3. the real run (~$9β14)
|
| 133 |
+
hf jobs run -d --flavor l40sx1 --timeout 12h --secrets HF_TOKEN --name securecoder-run1 \
|
| 134 |
+
ghcr.io/astral-sh/uv:python3.12-bookworm \
|
| 135 |
+
uv run --no-project https://huggingface.co/Taimwe/securecoder-scripts/resolve/main/train_securecoder.py \
|
| 136 |
+
--num-epochs 1 --output-repo Taimwe/securecoder-30b-pro \
|
| 137 |
+
--trackio-space Taimwe/securecoder-trackio
|
| 138 |
+
|
| 139 |
+
# monitor
|
| 140 |
+
hf jobs inspect Taimwe/<job_id> --json
|
| 141 |
+
hf jobs logs Taimwe/<job_id>
|
| 142 |
+
hf jobs stats Taimwe/<job_id>
|
| 143 |
+
hf jobs cancel Taimwe/<job_id>
|
| 144 |
+
```
|
| 145 |
+
|
| 146 |
+
### Platform traps already worked around (verified 2026-09-24)
|
| 147 |
+
|
| 148 |
+
| Symptom | Cause | Workaround |
|
| 149 |
+
| --- | --- | --- |
|
| 150 |
+
| `failed to set MOUNT_ATTR_IDMAP on /usr/bin/nvidia-cuda-mps-control` on a GPU flavor | `hf jobs uv run` mounts an artifacts bucket; the idmap mount fails on some GPU nodes (reproduced twice, 4 s after start) | use `hf jobs run` + the script's **public Hub URL** (works: `Taimwe/6ab5ad3c6b030d633f68face`) or a read-only repo mount `-v hf://models/Taimwe/securecoder-scripts:/scripts:ro` (works: `Taimwe/6ab5ad3d52d0dbd7f1d8d286`) |
|
| 151 |
+
| inner quotes stripped from `python -c "β¦"` | PowerShell native-argument quoting | use a UV script file, never `-c` one-liners |
|
| 152 |
+
| `charmap codec can't encode` while reading logs | Windows console encoding | set `$env:PYTHONIOENCODING='utf-8'` before `hf jobs logs` |
|
| 153 |
+
| HF's trainer skill says Jobs need a paid plan | docs caveat | not enforced on this account β CPU and GPU jobs both ran |
|
| 154 |
+
|
| 155 |
+
## 6. Turning down refusals (optional, after SFT)
|
| 156 |
+
|
| 157 |
+
Abliteration is a weight edit, not a training run: find the "refusal direction" in
|
| 158 |
+
the residual stream and project it out of the writing matrices.
|
| 159 |
+
|
| 160 |
+
1. **`Heretic`** (p-e-w/heretic) β the tool behind the recent wave of
|
| 161 |
+
`*-heretic-abliterated-*` checkpoints (e.g.
|
| 162 |
+
`culturerevolt/gemma-4-12b-heretic-abliterated-GGUF`, 174k downloads). Runs on
|
| 163 |
+
the merged 16-bit model and optimises the ablation to trade refusal rate against
|
| 164 |
+
KL divergence, so it preserves capability better than a hand-rolled ablation.
|
| 165 |
+
2. **Manual directional ablation** (Arditi et al., *Refusal in LLMs is mediated by a
|
| 166 |
+
single direction*): collect residual activations for harmful vs. harmless
|
| 167 |
+
prompts, take the mean difference per layer, and zero that component in
|
| 168 |
+
`o_proj` / `down_proj`.
|
| 169 |
+
|
| 170 |
+
Honest expectations: **it is not free.** Refusals drop, and some performance on the
|
| 171 |
+
same layers' other duties drops with them; scope-setting behaviour ("this needs
|
| 172 |
+
authorisation") can weaken too. Measure both sides (section 7) and document the
|
| 173 |
+
result on the model card, or you are shipping an unmeasured change.
|
| 174 |
+
|
| 175 |
+
## 7. Evaluation before publishing
|
| 176 |
+
|
| 177 |
+
1. **Tool-call validity** β 200 prompts with real tool schemas; parse the emitted
|
| 178 |
+
`<function=NAME><parameter=β¦>` block back to JSON and score syntax validity,
|
| 179 |
+
correct function choice, and schema-conformant arguments. This is the number
|
| 180 |
+
that matters most for "top-notch tool calling".
|
| 181 |
+
2. **Code sanity** β generate 100 solutions, check compilation, run the bundled
|
| 182 |
+
tests on a small HumanEval-style subset.
|
| 183 |
+
3. **Security knowledge** β fixed MCQ/triage set (e.g. `CyberNative/CyberSecurityEval`),
|
| 184 |
+
scored before *and* after fine-tuning so the mix's effect is visible.
|
| 185 |
+
4. **Regression** β the same three on the untouched base, same prompts. A fine-tune
|
| 186 |
+
that feels better but scores worse is a regression with extra steps.
|
| 187 |
+
5. **Refusal rate** β if abliterated, publish before/after refusal rate next to the
|
| 188 |
+
eval deltas.
|
| 189 |
+
|
| 190 |
+
## 8. Known risks
|
| 191 |
+
|
| 192 |
+
* MoE LoRA defaults to attention + router (`q/k/v/o/gate`). `--target-modules
|
| 193 |
+
all-linear` on a 128-expert model means ~800M trainable parameters β much slower
|
| 194 |
+
and hungrier. Start conservative; revisit if security knowledge is the weak score.
|
| 195 |
+
* 4-bit QLoRA of a 30B (~17 GB) is tight on a 24 GB card; 48 GB allows batch 2 with
|
| 196 |
+
4096-token samples.
|
| 197 |
+
* One 1-epoch pass over three domains is a *start*, not a finished model. Iterate on
|
| 198 |
+
the mix ratios β they live in one table at the top of the script.
|
| 199 |
+
* Nothing here adds safety training, and the security slice includes
|
| 200 |
+
recon/enumeration knowledge framed for authorised use. Publish explicit
|
| 201 |
+
intended-use and limitations sections; do not present it as a hardened model.
|