Download docs/ABLATION_GUIDELINES.md from rpDungeon/instruct-mask-tools: direct link, hf CLI and curl.
- Browser
- Download file 7.55 kB
-
https://huggingface.co/rpDungeon/instruct-mask-tools/resolve/main/docs/ABLATION_GUIDELINES.md
- Command line
-
hf download hf://rpDungeon/instruct-mask-tools/docs/ABLATION_GUIDELINES.md
-
curl -L -o ABLATION_GUIDELINES.md https://huggingface.co/rpDungeon/instruct-mask-tools/resolve/main/docs/ABLATION_GUIDELINES.md
RECOVERED 2026-06-07T13:16:02 from Claude session transcripts (last Write; later Edit-tool revisions not applied). Original file was missing from disk.
Ablation Guidelines — directional ablation / activation steering
Created 2026-06-07 (twisted) from the dev-channel design discussion. Playbook for the slop-ablation
runs we intend to do on 12B and 31B. Companion to DESLOP_APPROACHES.md (findings + why steering is
the lever) and memory project_slop_inherited_gemma4.md. SFW-only; the SFW pre-gate (sfw_pregate.py,
0 HARD or abort) runs before any judge reads creative output.
0. What this technique is (names)
- Directional ablation / weight orthogonalization — mechanically identical to "abliteration" (the trick that strips the refusal direction out of a model), but pointed at a slop direction instead.
- Direction found by difference-of-means activation steering:
d = mean(resid | slop spans) − mean(resid | clean spans), per layer. - Runtime-tunable cousin = control vector (repeng;
--control-vectorin llama.cpp / koboldcpp). - It is NOT the logit-lens localization (that's single-token diagnosis,
slop_heatmap.py). Ablation works on residual-stream activations over whole spans, so it can reach contextual / "purple-mode" slop the token method is blind to — the reason it's the recommended lever where weight-masking died.
1. Method (the loop)
- Contrast set — curate slop-vs-clean continuation pairs that cleanly isolate slop. Long pole; garbage in → garbage direction. Crude set for PoC, curated set for the proper pass.
- Activation collection — forward-only passes over the contrast set; capture residual-stream activations at each layer. Batchable (see §4).
- Direction — per-layer mean-difference (
d̂= unit-normalized). Trivial vector math. - Apply — two modes:
- Runtime forward-hook (DEV ITERATION ONLY) — subtract
α·(d̂ d̂ᵀ)·residat inference. Fast, no re-quant; use it to sweep layers/strength cheaply. - Weight-baked — orthogonalize the output-writing matrices:
W' = W − α·(d̂ d̂ᵀ)Wondown_proj/o_proj/embed. Permanent.
- Runtime forward-hook (DEV ITERATION ONLY) — subtract
- Layer/strength sweep — you don't know the right layer a priori; extract at several, try a few
α. Embarrassingly parallel across configs. - Validate — judge step (§5). Did slop drop and voice survive?
2. Why a PoC takes ~3 hr (compute is minutes)
Matrix math is trivial. Wall-clock is dominated by NON-GPU work:
- building a usable contrast set,
- the layer/strength sweep (several configs, iterate),
- generate-and-judge validation loops,
- (
re-quant to Q4_K_M, CPU— deferred on B200 runs, see §4). Local tax: bf16 31B (~62GB) doesn't fit on 2×24GB → clean activation collection spills to CPU/offload (slow).
3. Artifact form, size & tunability
- The extracted edit is just a direction vector per layer (~few × d_model; d_model=5376 for g4-31B) → KB–low-MB. Uploadable standalone (vs 18GB Q4 / 62GB bf16): keep the base model untouched, ship only the edit.
- The weight-baked delta
−α·d̂(d̂ᵀW)is an outer product = rank-1 → expressible as the smallest possible rank-1 LoRA (one per affected matrix). The control-vector form is activation-space (not technically a LoRA) but the same small-bolt-on idea. - Strength scalar
αis built in:- control-vector → runtime dial (up / down / off, no re-bake);
- LoRA / weight-bake → bake-time merge scale (one bake per strength).
α=1removes the direction;<1partial;>1over-steer (risky).
- Ceiling: too-high
αflattens good vivid prose too (slop + vivid share machinery). The PoC's job is to find theαthat subtracts slop without gutting voice.
4. Shipping forms
- Weight-baked → plain GGUF that runs anywhere (koboldcpp/llama.cpp/ollama, no flags). Permanent, one fixed strength. = how HF "abliterated" models ship. Needs a re-quant to Q4_K_M at ship-time.
- Control vector → normal GGUF + tiny
.ggufside-file; user adds--control-vector deslop.gguf. Runtime-tunable, no re-quant, but needs the flag + a supporting engine.
5. Running it on B200 (twisted directive 2026-06-07)
- Work + validate in bf16; DEFER re-quant to ship-time — don't quant during the PoC.
- Batching is the GPU win: 180GB B200 fits 31B bf16 (~62GB) + ~115GB headroom → push 32–64 contrast prompts/forward-pass; continuous batching (vLLM-style) for the validation gens. Collapses the GPU slice.
- Same-model multi-instance on ONE B200 = no win (shared SMs/bandwidth → time-slice). Fix idle GPU with a bigger batch, not a second process.
- Parallelism that maps: the (layer, strength) sweep is independent work → one config per GPU on a multi-B200 node = linear speedup. On a single B200, sequence them (each is fast once batched).
- 12B doesn't transfer — a direction extracted on 12B won't apply to 31B (different model). 12B is a fast methodology dry-run (prove the harness) before spending 31B passes.
- Post-batch bottleneck moves OFF the GPU → contrast-set curation + judge round-trips. Parallelize the judging via API concurrency, not more GPU.
6. The judge step (validation)
After ablation you must measure two things that a lexical counter alone can't: did slop actually drop
and did voice/quality survive (the over-steer guard). Use the same blinded, counterbalanced LLM-judge
machinery as the merge evals (pairwise_build.py → pairwise_tally.py, Sonnet judge):
- Paired generation — run baseline vs ablated on the same prompts (8-genre tone/style set, rp_driver), same sampling. Only the edit varies.
- Blind + counterbalance — judge sees response A vs B without labels; swap A/B order on half the pairs to cancel position bias.
- Score the right axes:
- slop / purple-prose reduction — fewer clichés, stock phrases, parallel-construction codas;
- voice / quality preservation — still vivid, in-character, coherent (the GUARDRAIL — a "win" that's just blander output is a fail).
- Tally — pairwise win-rates. Success = ablated ≥ baseline on voice/quality AND clearly less sloppy.
Fail mode = "less sloppy but blander" = over-steer / voice-flattening → lower
α. - Cross-check with the cheap objective metric: n-gram / phrase-frequency delta on the generations gives an objective slop read (catches repeated stock phrases) that the judge can't be the sole source of. Use both — n-gram = slop delta, judge = did voice survive + structural-slop read.
- PoC vs proper: PoC = single judge pass (is there signal?). Proper = blinded double-judge (two independent runs / models) for confidence + the n-gram cross-check, over a curated contrast set.
- The judge is the wall-clock long pole once GPU is batched (API latency) — parallelize judge calls.
7. Planned runs
- 12B — methodology dry-run first (cheap, proves the harness end-to-end).
- 31B — the real target (Mx5 line). Reuse the 12B-validated harness; bf16 on B200, defer quant.
- Gate: PoC (timeboxed) must show slop-down WITHOUT voice-flattening before committing the ~½–1 day proper
pass. If the PoC over-steers at every
αthat moves slop, the lever is voice-coupled too tightly → report and fall back to DPO-against-slop or sampler-side (seeDESLOP_APPROACHES.md§2–3).