Taimwe commited on
Commit
7c6176b
Β·
verified Β·
1 Parent(s): 4d01873

SecureCoder runbook

Browse files
Files changed (1) hide show
  1. README.md +201 -0
README.md ADDED
@@ -0,0 +1,201 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # SecureCoder β€” training runbook
2
+
3
+ A QLoRA fine-tune that targets three skills at once: **coding**, **tool calling**
4
+ and **cybersecurity** (offence *and* defence), on a base chosen to train cheaply
5
+ and run fast. Everything here is measured or verified on real jobs, not assumed β€”
6
+ each claim carries the job id you can inspect.
7
+
8
+ ## 1. Base model choice
9
+
10
+ Screened on the Hub on 2026-09-24 (params / licence / tool-calling support):
11
+
12
+ | Candidate | Params | Licence | Verdict |
13
+ | --- | --- | --- | --- |
14
+ | **`Qwen/Qwen3-Coder-30B-A3B-Instruct`** | 30.5B MoE, ~3B active | Apache-2.0 | **chosen** β€” code-specialised, native tool-calling template, MoE means it *trains* like a 3B model and fits 48 GB in 4-bit |
15
+ | `Qwen/Qwen3.8-27B` | 27.8B dense | Apache-2.0 | best raw quality, ~4x slower per step; good upgrade path once the pipeline is proven |
16
+ | `Qwen/Qwen3-Coder-Next` | 79.7B | Apache-2.0 | needs an 80 GB card (A100 80 GB `$2.50/h`); too big for the first run |
17
+ | `Qwen/Qwen3.8-Flash-Next` | 180B | **license: other** | rejected β€” non-commercial-style licence on an 180B model |
18
+ | `openbmb/MiniCPM5-2B` | 2.5B | Apache-2.0 | `tool-calling` tag, tiny; good for a T4 smoke test, weak ceiling |
19
+ | `TokenRhythm/NeoHorse-1-9B` | 9.0B | Apache-2.0 | `tool-use`, `reasoning`; the fallback if T4-only |
20
+
21
+ Why not start from an existing *abliterated* checkpoint? Because refusal removal is
22
+ a separate weight-edit (section 6) that can be applied to **any** checkpoint after
23
+ fine-tuning, and pre-abliterated bases are community re-uploads with unclear
24
+ provenance. Fine-tuning first keeps the lineage clean and the base licensed.
25
+
26
+ ## 2. Data mix
27
+
28
+ The cap per source *is* the recipe β€” sources differ hugely in size, so the mix is
29
+ balanced by taking a fixed slice of each. Verified on a real CPU job
30
+ (`Taimwe/6ab5ae256b030d633f68faef`) at 25 rows/source:
31
+
32
+ | Source | Rows taken | Kind | What it teaches |
33
+ | --- | --- | --- | --- |
34
+ | `NousResearch/hermes-function-calling-v1` `[func_calling]` | 9,000 | tool calling | full tool-call conversations + JSON schemas |
35
+ | `NousResearch/hermes-function-calling-v1` `[func_calling_singleturn]` | 3,000 | tool calling | picking the right function, no chatter |
36
+ | `lockon/xlam-function-calling-60k` | 10,000 | tool calling | 60k API-call pairs (query β†’ call) |
37
+ | `ise-uiuc/Magicoder-OSS-Instruct-75K` | 10,000 | coding | self-instruct problems + solutions |
38
+ | `Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset` | 8,000 | security | security instruction tuning |
39
+ | `AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1` | 5,000 | security | broad security Q&A |
40
+ | `Humanlearning/CyberSecurity_OWASP-sft-dataset` | 3,000 | secure coding | OWASP / defensive coding |
41
+ | `MrClipperz134/CTF-Instruct` | 3,000 | CTF | challenge instruction β†’ solve |
42
+ | `TrueNix/ctf-solver-dataset` | 3,000 | CTF | solving trajectories |
43
+ | `mlabonne/FineTome-100k` | 3,000 | replay | general chat, so instruction-following does not collapse |
44
+
45
+ **Total β‰ˆ 57,000 rows, ~30% of them tool-calling.** Every source validated at
46
+ 25/25 kept with zero render failures.
47
+
48
+ Excluded deliberately:
49
+
50
+ * `dpevzner/Cybersecurity_Reasoning_Dataset` β€” its builder configs were renamed
51
+ (`default` β†’ `mistral|deepseek|chatml|gemma`) and the fields the card documents
52
+ (`unified_interpretation`) are empty in the published revision, so every row was
53
+ dropped. Re-addable once its schema settles.
54
+ * Anything whose purpose is building malware or weaponised exploits. Recon and
55
+ enumeration knowledge, exploit *concepts*, CTF solving and defensive
56
+ engineering are in; end-to-end attack tooling is not. This is a deliberate
57
+ scope line, not an oversight.
58
+
59
+ ## 3. What training does with the data
60
+
61
+ Each row is converted to chat messages and rendered through the model's own chat
62
+ template, so the model learns the exact format it will be asked for at inference
63
+ time (`<|im_start|>…<tool_call><function=NAME><parameter=…>` for Qwen3-Coder).
64
+
65
+ Two format traps found the hard way and handled in code:
66
+
67
+ * Qwen3-Coder iterates `tool.parameters.properties` β€” tool schemas must be **flat**
68
+ (`{"name", "description", "parameters"}`), not wrapped in `"function"`.
69
+ * It also iterates `tool_call.arguments | items`, so `arguments` must be a
70
+ **mapping**. A JSON string there fails with *"Can only get item pairs from a
71
+ mapping"* β€” which is exactly how xLAM was silently contributing 0 rows until fixed.
72
+
73
+ `render_record()` therefore detects the template's convention once (flat/nested Γ—
74
+ dict/string) and reuses it, instead of hard-coding one model's dialect.
75
+
76
+ ## 4. Compute and cost
77
+
78
+ Verified `hf jobs hardware` prices (per hour), 2026-09-24:
79
+
80
+ | Flavor | VRAM | $/hour | Use |
81
+ | --- | --- | --- | --- |
82
+ | `cpu-basic` | β€” | $0.01 | data-mix validation, class checks |
83
+ | `t4-small` | 16 GB (T4) | $0.40 | plumbing checks; too small for a 30B in 4-bit |
84
+ | `l4x1` | 24 GB | $0.80 | 9B runs on a budget |
85
+ | `a10g-large` | 24 GB | $1.50 | 9B–14B comfortably |
86
+ | **`l40sx1`** | **48 GB** | **$1.80** | **the 30B-A3B QLoRA** |
87
+ | `a100-large` | 80 GB | $2.50 | faster per step, larger batch |
88
+ | `rtx-pro-6000` | 96 GB | $2.75 | headroom for longer context |
89
+ | `h200` | 141 GB | $5.00 | full fine-tune territory, overkill here |
90
+
91
+ Rough budget for the full 1-epoch run. **These are now measured, not guessed** β€” the
92
+ smoke test (`Taimwe/6ab5b3ed52d0dbd7f1d8d454`, a100-large) trained at
93
+ **0.8 rows/s** (20 s per step at batch 2 Γ— grad-accum 8, max seq 2048, packing on),
94
+ which includes a few minutes of `torch.compile` warm-up that does not recur per step:
95
+
96
+ | Mix | Rows | Time at 0.8 rows/s | Cost on `l40sx1` ($1.80/h) |
97
+ | --- | --- | --- | --- |
98
+ | full mix as specified | ~57,000 | ~20 h | **~$36** |
99
+ | half caps (`--max-steps` or edit the table) | ~28,000 | ~10 h | ~$18 |
100
+ | lean mix (cut caps to ~1/3) | ~19,000 | ~7 h | ~$12 |
101
+
102
+ Because warm-up inflates that rate, a **300-step measured run** (~$4 on a100-large)
103
+ gives a trustworthy steady-state number before committing to a long job. Levers that
104
+ move the cost most, in order: rows in the mix, `--max-seq-length` (2048 vs 4096),
105
+ `--batch-size` (raise it on 80–96 GB cards), `--lora-r`.
106
+
107
+ The `--smoke` run itself costs under $1 and proves model load, data rendering,
108
+ training, evaluation and the Hub push end to end.
109
+
110
+ No local GPU needed. Colab works the same way with `uv run train_securecoder.py …`
111
+ (Colab Free's T4 cannot hold a 30B in 4-bit β€” use Colab Pro L4/A100, or point
112
+ `--base-model` at a 9B).
113
+
114
+ ## 5. Runbook (copy-paste)
115
+
116
+ ```bash
117
+ # 0. auth (token needs repo.write, and job.write for Jobs)
118
+ hf auth login --token hf_xxx
119
+
120
+ # 1. validate the mix: 25 rows/source, no GPU, ~50 s, ~$0.0002
121
+ hf jobs run -d --flavor cpu-basic --timeout 30m --name validate-mix \
122
+ ghcr.io/astral-sh/uv:python3.12-bookworm \
123
+ uv run --no-project https://huggingface.co/Taimwe/securecoder-scripts/resolve/main/train_securecoder.py \
124
+ --validate-only --validate-per-source 25 --show-samples 3
125
+
126
+ # 2. smoke test on a real GPU: 200 rows/source, 20 steps, pushes an adapter (~$0.50)
127
+ hf jobs run -d --flavor l40sx1 --timeout 40m --secrets HF_TOKEN --name smoke-train \
128
+ ghcr.io/astral-sh/uv:python3.12-bookworm \
129
+ uv run --no-project https://huggingface.co/Taimwe/securecoder-scripts/resolve/main/train_securecoder.py \
130
+ --smoke --output-repo Taimwe/securecoder-smoke --private --report-to none
131
+
132
+ # 3. the real run (~$9–14)
133
+ hf jobs run -d --flavor l40sx1 --timeout 12h --secrets HF_TOKEN --name securecoder-run1 \
134
+ ghcr.io/astral-sh/uv:python3.12-bookworm \
135
+ uv run --no-project https://huggingface.co/Taimwe/securecoder-scripts/resolve/main/train_securecoder.py \
136
+ --num-epochs 1 --output-repo Taimwe/securecoder-30b-pro \
137
+ --trackio-space Taimwe/securecoder-trackio
138
+
139
+ # monitor
140
+ hf jobs inspect Taimwe/<job_id> --json
141
+ hf jobs logs Taimwe/<job_id>
142
+ hf jobs stats Taimwe/<job_id>
143
+ hf jobs cancel Taimwe/<job_id>
144
+ ```
145
+
146
+ ### Platform traps already worked around (verified 2026-09-24)
147
+
148
+ | Symptom | Cause | Workaround |
149
+ | --- | --- | --- |
150
+ | `failed to set MOUNT_ATTR_IDMAP on /usr/bin/nvidia-cuda-mps-control` on a GPU flavor | `hf jobs uv run` mounts an artifacts bucket; the idmap mount fails on some GPU nodes (reproduced twice, 4 s after start) | use `hf jobs run` + the script's **public Hub URL** (works: `Taimwe/6ab5ad3c6b030d633f68face`) or a read-only repo mount `-v hf://models/Taimwe/securecoder-scripts:/scripts:ro` (works: `Taimwe/6ab5ad3d52d0dbd7f1d8d286`) |
151
+ | inner quotes stripped from `python -c "…"` | PowerShell native-argument quoting | use a UV script file, never `-c` one-liners |
152
+ | `charmap codec can't encode` while reading logs | Windows console encoding | set `$env:PYTHONIOENCODING='utf-8'` before `hf jobs logs` |
153
+ | HF's trainer skill says Jobs need a paid plan | docs caveat | not enforced on this account β€” CPU and GPU jobs both ran |
154
+
155
+ ## 6. Turning down refusals (optional, after SFT)
156
+
157
+ Abliteration is a weight edit, not a training run: find the "refusal direction" in
158
+ the residual stream and project it out of the writing matrices.
159
+
160
+ 1. **`Heretic`** (p-e-w/heretic) β€” the tool behind the recent wave of
161
+ `*-heretic-abliterated-*` checkpoints (e.g.
162
+ `culturerevolt/gemma-4-12b-heretic-abliterated-GGUF`, 174k downloads). Runs on
163
+ the merged 16-bit model and optimises the ablation to trade refusal rate against
164
+ KL divergence, so it preserves capability better than a hand-rolled ablation.
165
+ 2. **Manual directional ablation** (Arditi et al., *Refusal in LLMs is mediated by a
166
+ single direction*): collect residual activations for harmful vs. harmless
167
+ prompts, take the mean difference per layer, and zero that component in
168
+ `o_proj` / `down_proj`.
169
+
170
+ Honest expectations: **it is not free.** Refusals drop, and some performance on the
171
+ same layers' other duties drops with them; scope-setting behaviour ("this needs
172
+ authorisation") can weaken too. Measure both sides (section 7) and document the
173
+ result on the model card, or you are shipping an unmeasured change.
174
+
175
+ ## 7. Evaluation before publishing
176
+
177
+ 1. **Tool-call validity** β€” 200 prompts with real tool schemas; parse the emitted
178
+ `<function=NAME><parameter=…>` block back to JSON and score syntax validity,
179
+ correct function choice, and schema-conformant arguments. This is the number
180
+ that matters most for "top-notch tool calling".
181
+ 2. **Code sanity** β€” generate 100 solutions, check compilation, run the bundled
182
+ tests on a small HumanEval-style subset.
183
+ 3. **Security knowledge** β€” fixed MCQ/triage set (e.g. `CyberNative/CyberSecurityEval`),
184
+ scored before *and* after fine-tuning so the mix's effect is visible.
185
+ 4. **Regression** β€” the same three on the untouched base, same prompts. A fine-tune
186
+ that feels better but scores worse is a regression with extra steps.
187
+ 5. **Refusal rate** β€” if abliterated, publish before/after refusal rate next to the
188
+ eval deltas.
189
+
190
+ ## 8. Known risks
191
+
192
+ * MoE LoRA defaults to attention + router (`q/k/v/o/gate`). `--target-modules
193
+ all-linear` on a 128-expert model means ~800M trainable parameters β€” much slower
194
+ and hungrier. Start conservative; revisit if security knowledge is the weak score.
195
+ * 4-bit QLoRA of a 30B (~17 GB) is tight on a 24 GB card; 48 GB allows batch 2 with
196
+ 4096-token samples.
197
+ * One 1-epoch pass over three domains is a *start*, not a finished model. Iterate on
198
+ the mix ratios β€” they live in one table at the top of the script.
199
+ * Nothing here adds safety training, and the security slice includes
200
+ recon/enumeration knowledge framed for authorised use. Publish explicit
201
+ intended-use and limitations sections; do not present it as a hardened model.