Initial release: instruct-preserving masked training & merging core tools + guidelines (recovered from dev logs, sanitized)
a4df021 verified |
Download docs/ABLATION_GUIDELINES.md from rpDungeon/instruct-mask-tools: direct link, hf CLI and curl.
- Browser
- Download file 7.55 kB
-
https://huggingface.co/rpDungeon/instruct-mask-tools/resolve/main/docs/ABLATION_GUIDELINES.md
- Command line
-
hf download hf://rpDungeon/instruct-mask-tools/docs/ABLATION_GUIDELINES.md
-
curl -L -o ABLATION_GUIDELINES.md https://huggingface.co/rpDungeon/instruct-mask-tools/resolve/main/docs/ABLATION_GUIDELINES.md
7.55 kB
| > RECOVERED 2026-06-07T13:16:02 from Claude session transcripts (last Write; later Edit-tool revisions not applied). Original file was missing from disk. | |
| # Ablation Guidelines — directional ablation / activation steering | |
| Created 2026-06-07 (twisted) from the dev-channel design discussion. Playbook for the **slop-ablation | |
| runs we intend to do on 12B and 31B**. Companion to `DESLOP_APPROACHES.md` (findings + why steering is | |
| the lever) and memory `project_slop_inherited_gemma4.md`. SFW-only; the SFW pre-gate (`sfw_pregate.py`, | |
| 0 HARD or abort) runs before any judge reads creative output. | |
| ## 0. What this technique is (names) | |
| - **Directional ablation / weight orthogonalization** — mechanically identical to **"abliteration"** (the | |
| trick that strips the *refusal* direction out of a model), but pointed at a **slop** direction instead. | |
| - Direction found by **difference-of-means activation steering**: `d = mean(resid | slop spans) − mean(resid | clean spans)`, per layer. | |
| - Runtime-tunable cousin = **control vector** (repeng; `--control-vector` in llama.cpp / koboldcpp). | |
| - It is NOT the logit-lens localization (that's single-token diagnosis, `slop_heatmap.py`). Ablation works | |
| on residual-stream activations over **whole spans**, so it can reach contextual / "purple-mode" slop the | |
| token method is blind to — the reason it's the recommended lever where weight-masking died. | |
| ## 1. Method (the loop) | |
| 1. **Contrast set** — curate slop-vs-clean continuation pairs that cleanly isolate slop. *Long pole;* | |
| garbage in → garbage direction. Crude set for PoC, curated set for the proper pass. | |
| 2. **Activation collection** — forward-only passes over the contrast set; capture residual-stream | |
| activations at each layer. Batchable (see §4). | |
| 3. **Direction** — per-layer mean-difference (`d̂` = unit-normalized). Trivial vector math. | |
| 4. **Apply** — two modes: | |
| - **Runtime forward-hook** (DEV ITERATION ONLY) — subtract `α·(d̂ d̂ᵀ)·resid` at inference. Fast, no | |
| re-quant; use it to sweep layers/strength cheaply. | |
| - **Weight-baked** — orthogonalize the output-writing matrices: `W' = W − α·(d̂ d̂ᵀ)W` on | |
| `down_proj` / `o_proj` / `embed`. Permanent. | |
| 5. **Layer/strength sweep** — you don't know the right layer a priori; extract at several, try a few `α`. | |
| Embarrassingly parallel across configs. | |
| 6. **Validate** — judge step (§5). Did slop drop *and* voice survive? | |
| ## 2. Why a PoC takes ~3 hr (compute is minutes) | |
| Matrix math is trivial. Wall-clock is dominated by NON-GPU work: | |
| - building a usable **contrast set**, | |
| - the **layer/strength sweep** (several configs, iterate), | |
| - **generate-and-judge** validation loops, | |
| - (~~re-quant to Q4_K_M, CPU~~ — **deferred on B200 runs**, see §4). | |
| Local tax: bf16 31B (~62GB) doesn't fit on 2×24GB → clean activation collection spills to CPU/offload (slow). | |
| ## 3. Artifact form, size & tunability | |
| - The extracted edit is just a **direction vector per layer** (~few × d_model; d_model=5376 for g4-31B) → | |
| **KB–low-MB**. Uploadable standalone (vs 18GB Q4 / 62GB bf16): keep the base model untouched, ship only | |
| the edit. | |
| - The weight-baked delta `−α·d̂(d̂ᵀW)` is an **outer product = rank-1** → expressible as the smallest | |
| possible **rank-1 LoRA** (one per affected matrix). The control-vector form is activation-space (not | |
| technically a LoRA) but the same small-bolt-on idea. | |
| - **Strength scalar `α` is built in:** | |
| - control-vector → **runtime dial** (up / down / off, no re-bake); | |
| - LoRA / weight-bake → **bake-time** merge scale (one bake per strength). | |
| - `α=1` removes the direction; `<1` partial; `>1` over-steer (risky). | |
| - **Ceiling:** too-high `α` flattens *good* vivid prose too (slop + vivid share machinery). The PoC's job is | |
| to find the `α` that subtracts slop without gutting voice. | |
| ## 4. Shipping forms | |
| - **Weight-baked → plain GGUF** that runs anywhere (koboldcpp/llama.cpp/ollama, no flags). Permanent, one | |
| fixed strength. = how HF "abliterated" models ship. Needs a re-quant to Q4_K_M at ship-time. | |
| - **Control vector → normal GGUF + tiny `.gguf` side-file**; user adds `--control-vector deslop.gguf`. | |
| Runtime-tunable, no re-quant, but needs the flag + a supporting engine. | |
| ## 5. Running it on B200 (twisted directive 2026-06-07) | |
| - **Work + validate in bf16; DEFER re-quant to ship-time** — don't quant during the PoC. | |
| - **Batching is the GPU win:** 180GB B200 fits 31B bf16 (~62GB) + ~115GB headroom → push 32–64 contrast | |
| prompts/forward-pass; continuous batching (vLLM-style) for the validation gens. Collapses the GPU slice. | |
| - **Same-model multi-instance on ONE B200 = no win** (shared SMs/bandwidth → time-slice). Fix idle GPU with | |
| a bigger batch, not a second process. | |
| - **Parallelism that maps:** the (layer, strength) sweep is independent work → one config per GPU on a | |
| multi-B200 node = linear speedup. On a single B200, sequence them (each is fast once batched). | |
| - **12B doesn't transfer** — a direction extracted on 12B won't apply to 31B (different model). 12B is a | |
| fast **methodology dry-run** (prove the harness) before spending 31B passes. | |
| - **Post-batch bottleneck** moves OFF the GPU → contrast-set curation + judge round-trips. Parallelize the | |
| judging via API concurrency, not more GPU. | |
| ## 6. The judge step (validation) | |
| After ablation you must measure two things that a lexical counter alone can't: **did slop actually drop** | |
| and **did voice/quality survive** (the over-steer guard). Use the same blinded, counterbalanced LLM-judge | |
| machinery as the merge evals (`pairwise_build.py` → `pairwise_tally.py`, Sonnet judge): | |
| 1. **Paired generation** — run baseline vs ablated on the *same* prompts (8-genre tone/style set, | |
| rp_driver), same sampling. Only the edit varies. | |
| 2. **Blind + counterbalance** — judge sees response A vs B without labels; swap A/B order on half the pairs | |
| to cancel position bias. | |
| 3. **Score the right axes:** | |
| - **slop / purple-prose reduction** — fewer clichés, stock phrases, parallel-construction codas; | |
| - **voice / quality preservation** — still vivid, in-character, coherent (the GUARDRAIL — a "win" that's | |
| just blander output is a *fail*). | |
| 4. **Tally** — pairwise win-rates. **Success = ablated ≥ baseline on voice/quality AND clearly less sloppy.** | |
| **Fail mode = "less sloppy but blander"** = over-steer / voice-flattening → lower `α`. | |
| 5. **Cross-check with the cheap objective metric:** n-gram / phrase-frequency delta on the generations gives | |
| an *objective* slop read (catches repeated stock phrases) that the judge can't be the sole source of. | |
| Use both — n-gram = slop delta, judge = did voice survive + structural-slop read. | |
| 6. **PoC vs proper:** PoC = single judge pass (is there signal?). Proper = blinded **double**-judge (two | |
| independent runs / models) for confidence + the n-gram cross-check, over a curated contrast set. | |
| - The judge is the wall-clock long pole once GPU is batched (API latency) — parallelize judge calls. | |
| ## 7. Planned runs | |
| - **12B** — methodology dry-run first (cheap, proves the harness end-to-end). | |
| - **31B** — the real target (Mx5 line). Reuse the 12B-validated harness; bf16 on B200, defer quant. | |
| - Gate: PoC (timeboxed) must show slop-down WITHOUT voice-flattening before committing the ~½–1 day proper | |
| pass. If the PoC over-steers at every `α` that moves slop, the lever is voice-coupled too tightly → report | |
| and fall back to DPO-against-slop or sampler-side (see `DESLOP_APPROACHES.md` §2–3). | |