> RECOVERED 2026-06-07T13:16:02 from Claude session transcripts (last Write; later Edit-tool revisions not applied). Original file was missing from disk. # Ablation Guidelines — directional ablation / activation steering Created 2026-06-07 (twisted) from the dev-channel design discussion. Playbook for the **slop-ablation runs we intend to do on 12B and 31B**. Companion to `DESLOP_APPROACHES.md` (findings + why steering is the lever) and memory `project_slop_inherited_gemma4.md`. SFW-only; the SFW pre-gate (`sfw_pregate.py`, 0 HARD or abort) runs before any judge reads creative output. ## 0. What this technique is (names) - **Directional ablation / weight orthogonalization** — mechanically identical to **"abliteration"** (the trick that strips the *refusal* direction out of a model), but pointed at a **slop** direction instead. - Direction found by **difference-of-means activation steering**: `d = mean(resid | slop spans) − mean(resid | clean spans)`, per layer. - Runtime-tunable cousin = **control vector** (repeng; `--control-vector` in llama.cpp / koboldcpp). - It is NOT the logit-lens localization (that's single-token diagnosis, `slop_heatmap.py`). Ablation works on residual-stream activations over **whole spans**, so it can reach contextual / "purple-mode" slop the token method is blind to — the reason it's the recommended lever where weight-masking died. ## 1. Method (the loop) 1. **Contrast set** — curate slop-vs-clean continuation pairs that cleanly isolate slop. *Long pole;* garbage in → garbage direction. Crude set for PoC, curated set for the proper pass. 2. **Activation collection** — forward-only passes over the contrast set; capture residual-stream activations at each layer. Batchable (see §4). 3. **Direction** — per-layer mean-difference (`d̂` = unit-normalized). Trivial vector math. 4. **Apply** — two modes: - **Runtime forward-hook** (DEV ITERATION ONLY) — subtract `α·(d̂ d̂ᵀ)·resid` at inference. Fast, no re-quant; use it to sweep layers/strength cheaply. - **Weight-baked** — orthogonalize the output-writing matrices: `W' = W − α·(d̂ d̂ᵀ)W` on `down_proj` / `o_proj` / `embed`. Permanent. 5. **Layer/strength sweep** — you don't know the right layer a priori; extract at several, try a few `α`. Embarrassingly parallel across configs. 6. **Validate** — judge step (§5). Did slop drop *and* voice survive? ## 2. Why a PoC takes ~3 hr (compute is minutes) Matrix math is trivial. Wall-clock is dominated by NON-GPU work: - building a usable **contrast set**, - the **layer/strength sweep** (several configs, iterate), - **generate-and-judge** validation loops, - (~~re-quant to Q4_K_M, CPU~~ — **deferred on B200 runs**, see §4). Local tax: bf16 31B (~62GB) doesn't fit on 2×24GB → clean activation collection spills to CPU/offload (slow). ## 3. Artifact form, size & tunability - The extracted edit is just a **direction vector per layer** (~few × d_model; d_model=5376 for g4-31B) → **KB–low-MB**. Uploadable standalone (vs 18GB Q4 / 62GB bf16): keep the base model untouched, ship only the edit. - The weight-baked delta `−α·d̂(d̂ᵀW)` is an **outer product = rank-1** → expressible as the smallest possible **rank-1 LoRA** (one per affected matrix). The control-vector form is activation-space (not technically a LoRA) but the same small-bolt-on idea. - **Strength scalar `α` is built in:** - control-vector → **runtime dial** (up / down / off, no re-bake); - LoRA / weight-bake → **bake-time** merge scale (one bake per strength). - `α=1` removes the direction; `<1` partial; `>1` over-steer (risky). - **Ceiling:** too-high `α` flattens *good* vivid prose too (slop + vivid share machinery). The PoC's job is to find the `α` that subtracts slop without gutting voice. ## 4. Shipping forms - **Weight-baked → plain GGUF** that runs anywhere (koboldcpp/llama.cpp/ollama, no flags). Permanent, one fixed strength. = how HF "abliterated" models ship. Needs a re-quant to Q4_K_M at ship-time. - **Control vector → normal GGUF + tiny `.gguf` side-file**; user adds `--control-vector deslop.gguf`. Runtime-tunable, no re-quant, but needs the flag + a supporting engine. ## 5. Running it on B200 (twisted directive 2026-06-07) - **Work + validate in bf16; DEFER re-quant to ship-time** — don't quant during the PoC. - **Batching is the GPU win:** 180GB B200 fits 31B bf16 (~62GB) + ~115GB headroom → push 32–64 contrast prompts/forward-pass; continuous batching (vLLM-style) for the validation gens. Collapses the GPU slice. - **Same-model multi-instance on ONE B200 = no win** (shared SMs/bandwidth → time-slice). Fix idle GPU with a bigger batch, not a second process. - **Parallelism that maps:** the (layer, strength) sweep is independent work → one config per GPU on a multi-B200 node = linear speedup. On a single B200, sequence them (each is fast once batched). - **12B doesn't transfer** — a direction extracted on 12B won't apply to 31B (different model). 12B is a fast **methodology dry-run** (prove the harness) before spending 31B passes. - **Post-batch bottleneck** moves OFF the GPU → contrast-set curation + judge round-trips. Parallelize the judging via API concurrency, not more GPU. ## 6. The judge step (validation) After ablation you must measure two things that a lexical counter alone can't: **did slop actually drop** and **did voice/quality survive** (the over-steer guard). Use the same blinded, counterbalanced LLM-judge machinery as the merge evals (`pairwise_build.py` → `pairwise_tally.py`, Sonnet judge): 1. **Paired generation** — run baseline vs ablated on the *same* prompts (8-genre tone/style set, rp_driver), same sampling. Only the edit varies. 2. **Blind + counterbalance** — judge sees response A vs B without labels; swap A/B order on half the pairs to cancel position bias. 3. **Score the right axes:** - **slop / purple-prose reduction** — fewer clichés, stock phrases, parallel-construction codas; - **voice / quality preservation** — still vivid, in-character, coherent (the GUARDRAIL — a "win" that's just blander output is a *fail*). 4. **Tally** — pairwise win-rates. **Success = ablated ≥ baseline on voice/quality AND clearly less sloppy.** **Fail mode = "less sloppy but blander"** = over-steer / voice-flattening → lower `α`. 5. **Cross-check with the cheap objective metric:** n-gram / phrase-frequency delta on the generations gives an *objective* slop read (catches repeated stock phrases) that the judge can't be the sole source of. Use both — n-gram = slop delta, judge = did voice survive + structural-slop read. 6. **PoC vs proper:** PoC = single judge pass (is there signal?). Proper = blinded **double**-judge (two independent runs / models) for confidence + the n-gram cross-check, over a curated contrast set. - The judge is the wall-clock long pole once GPU is batched (API latency) — parallelize judge calls. ## 7. Planned runs - **12B** — methodology dry-run first (cheap, proves the harness end-to-end). - **31B** — the real target (Mx5 line). Reuse the 12B-validated harness; bf16 on B200, defer quant. - Gate: PoC (timeboxed) must show slop-down WITHOUT voice-flattening before committing the ~½–1 day proper pass. If the PoC over-steers at every `α` that moves slop, the lever is voice-coupled too tightly → report and fall back to DPO-against-slop or sampler-side (see `DESLOP_APPROACHES.md` §2–3).