instruct-mask-tools / docs /ABLATION_GUIDELINES.md
ToastyPigeon's picture
Initial release: instruct-preserving masked training & merging core tools + guidelines (recovered from dev logs, sanitized)
a4df021 verified
|
Raw History Blame Contribute Delete
7.55 kB
> RECOVERED 2026-06-07T13:16:02 from Claude session transcripts (last Write; later Edit-tool revisions not applied). Original file was missing from disk.
# Ablation Guidelines — directional ablation / activation steering
Created 2026-06-07 (twisted) from the dev-channel design discussion. Playbook for the **slop-ablation
runs we intend to do on 12B and 31B**. Companion to `DESLOP_APPROACHES.md` (findings + why steering is
the lever) and memory `project_slop_inherited_gemma4.md`. SFW-only; the SFW pre-gate (`sfw_pregate.py`,
0 HARD or abort) runs before any judge reads creative output.
## 0. What this technique is (names)
- **Directional ablation / weight orthogonalization** — mechanically identical to **"abliteration"** (the
trick that strips the *refusal* direction out of a model), but pointed at a **slop** direction instead.
- Direction found by **difference-of-means activation steering**: `d = mean(resid | slop spans) − mean(resid | clean spans)`, per layer.
- Runtime-tunable cousin = **control vector** (repeng; `--control-vector` in llama.cpp / koboldcpp).
- It is NOT the logit-lens localization (that's single-token diagnosis, `slop_heatmap.py`). Ablation works
on residual-stream activations over **whole spans**, so it can reach contextual / "purple-mode" slop the
token method is blind to — the reason it's the recommended lever where weight-masking died.
## 1. Method (the loop)
1. **Contrast set** — curate slop-vs-clean continuation pairs that cleanly isolate slop. *Long pole;*
garbage in → garbage direction. Crude set for PoC, curated set for the proper pass.
2. **Activation collection** — forward-only passes over the contrast set; capture residual-stream
activations at each layer. Batchable (see §4).
3. **Direction** — per-layer mean-difference (`d̂` = unit-normalized). Trivial vector math.
4. **Apply** — two modes:
- **Runtime forward-hook** (DEV ITERATION ONLY) — subtract `α·(d̂ d̂ᵀ)·resid` at inference. Fast, no
re-quant; use it to sweep layers/strength cheaply.
- **Weight-baked** — orthogonalize the output-writing matrices: `W' = W − α·(d̂ d̂ᵀ)W` on
`down_proj` / `o_proj` / `embed`. Permanent.
5. **Layer/strength sweep** — you don't know the right layer a priori; extract at several, try a few `α`.
Embarrassingly parallel across configs.
6. **Validate** — judge step (§5). Did slop drop *and* voice survive?
## 2. Why a PoC takes ~3 hr (compute is minutes)
Matrix math is trivial. Wall-clock is dominated by NON-GPU work:
- building a usable **contrast set**,
- the **layer/strength sweep** (several configs, iterate),
- **generate-and-judge** validation loops,
- (~~re-quant to Q4_K_M, CPU~~ — **deferred on B200 runs**, see §4).
Local tax: bf16 31B (~62GB) doesn't fit on 2×24GB → clean activation collection spills to CPU/offload (slow).
## 3. Artifact form, size & tunability
- The extracted edit is just a **direction vector per layer** (~few × d_model; d_model=5376 for g4-31B) →
**KB–low-MB**. Uploadable standalone (vs 18GB Q4 / 62GB bf16): keep the base model untouched, ship only
the edit.
- The weight-baked delta `−α·d̂(d̂ᵀW)` is an **outer product = rank-1** → expressible as the smallest
possible **rank-1 LoRA** (one per affected matrix). The control-vector form is activation-space (not
technically a LoRA) but the same small-bolt-on idea.
- **Strength scalar `α` is built in:**
- control-vector → **runtime dial** (up / down / off, no re-bake);
- LoRA / weight-bake → **bake-time** merge scale (one bake per strength).
- `α=1` removes the direction; `<1` partial; `>1` over-steer (risky).
- **Ceiling:** too-high `α` flattens *good* vivid prose too (slop + vivid share machinery). The PoC's job is
to find the `α` that subtracts slop without gutting voice.
## 4. Shipping forms
- **Weight-baked → plain GGUF** that runs anywhere (koboldcpp/llama.cpp/ollama, no flags). Permanent, one
fixed strength. = how HF "abliterated" models ship. Needs a re-quant to Q4_K_M at ship-time.
- **Control vector → normal GGUF + tiny `.gguf` side-file**; user adds `--control-vector deslop.gguf`.
Runtime-tunable, no re-quant, but needs the flag + a supporting engine.
## 5. Running it on B200 (twisted directive 2026-06-07)
- **Work + validate in bf16; DEFER re-quant to ship-time** — don't quant during the PoC.
- **Batching is the GPU win:** 180GB B200 fits 31B bf16 (~62GB) + ~115GB headroom → push 32–64 contrast
prompts/forward-pass; continuous batching (vLLM-style) for the validation gens. Collapses the GPU slice.
- **Same-model multi-instance on ONE B200 = no win** (shared SMs/bandwidth → time-slice). Fix idle GPU with
a bigger batch, not a second process.
- **Parallelism that maps:** the (layer, strength) sweep is independent work → one config per GPU on a
multi-B200 node = linear speedup. On a single B200, sequence them (each is fast once batched).
- **12B doesn't transfer** — a direction extracted on 12B won't apply to 31B (different model). 12B is a
fast **methodology dry-run** (prove the harness) before spending 31B passes.
- **Post-batch bottleneck** moves OFF the GPU → contrast-set curation + judge round-trips. Parallelize the
judging via API concurrency, not more GPU.
## 6. The judge step (validation)
After ablation you must measure two things that a lexical counter alone can't: **did slop actually drop**
and **did voice/quality survive** (the over-steer guard). Use the same blinded, counterbalanced LLM-judge
machinery as the merge evals (`pairwise_build.py` → `pairwise_tally.py`, Sonnet judge):
1. **Paired generation** — run baseline vs ablated on the *same* prompts (8-genre tone/style set,
rp_driver), same sampling. Only the edit varies.
2. **Blind + counterbalance** — judge sees response A vs B without labels; swap A/B order on half the pairs
to cancel position bias.
3. **Score the right axes:**
- **slop / purple-prose reduction** — fewer clichés, stock phrases, parallel-construction codas;
- **voice / quality preservation** — still vivid, in-character, coherent (the GUARDRAIL — a "win" that's
just blander output is a *fail*).
4. **Tally** — pairwise win-rates. **Success = ablated ≥ baseline on voice/quality AND clearly less sloppy.**
**Fail mode = "less sloppy but blander"** = over-steer / voice-flattening → lower `α`.
5. **Cross-check with the cheap objective metric:** n-gram / phrase-frequency delta on the generations gives
an *objective* slop read (catches repeated stock phrases) that the judge can't be the sole source of.
Use both — n-gram = slop delta, judge = did voice survive + structural-slop read.
6. **PoC vs proper:** PoC = single judge pass (is there signal?). Proper = blinded **double**-judge (two
independent runs / models) for confidence + the n-gram cross-check, over a curated contrast set.
- The judge is the wall-clock long pole once GPU is batched (API latency) — parallelize judge calls.
## 7. Planned runs
- **12B** — methodology dry-run first (cheap, proves the harness end-to-end).
- **31B** — the real target (Mx5 line). Reuse the 12B-validated harness; bf16 on B200, defer quant.
- Gate: PoC (timeboxed) must show slop-down WITHOUT voice-flattening before committing the ~½–1 day proper
pass. If the PoC over-steers at every `α` that moves slop, the lever is voice-coupled too tightly → report
and fall back to DPO-against-slop or sampler-side (see `DESLOP_APPROACHES.md` §2–3).