instruct-mask-tools / docs /ABLATION_GUIDELINES.md
ToastyPigeon's picture
Initial release: instruct-preserving masked training & merging core tools + guidelines (recovered from dev logs, sanitized)
a4df021 verified
|
Raw History Blame Contribute Delete
7.55 kB

RECOVERED 2026-06-07T13:16:02 from Claude session transcripts (last Write; later Edit-tool revisions not applied). Original file was missing from disk.

Ablation Guidelines — directional ablation / activation steering

Created 2026-06-07 (twisted) from the dev-channel design discussion. Playbook for the slop-ablation runs we intend to do on 12B and 31B. Companion to DESLOP_APPROACHES.md (findings + why steering is the lever) and memory project_slop_inherited_gemma4.md. SFW-only; the SFW pre-gate (sfw_pregate.py, 0 HARD or abort) runs before any judge reads creative output.

0. What this technique is (names)

  • Directional ablation / weight orthogonalization — mechanically identical to "abliteration" (the trick that strips the refusal direction out of a model), but pointed at a slop direction instead.
  • Direction found by difference-of-means activation steering: d = mean(resid | slop spans) − mean(resid | clean spans), per layer.
  • Runtime-tunable cousin = control vector (repeng; --control-vector in llama.cpp / koboldcpp).
  • It is NOT the logit-lens localization (that's single-token diagnosis, slop_heatmap.py). Ablation works on residual-stream activations over whole spans, so it can reach contextual / "purple-mode" slop the token method is blind to — the reason it's the recommended lever where weight-masking died.

1. Method (the loop)

  1. Contrast set — curate slop-vs-clean continuation pairs that cleanly isolate slop. Long pole; garbage in → garbage direction. Crude set for PoC, curated set for the proper pass.
  2. Activation collection — forward-only passes over the contrast set; capture residual-stream activations at each layer. Batchable (see §4).
  3. Direction — per-layer mean-difference (d̂ = unit-normalized). Trivial vector math.
  4. Apply — two modes:
    • Runtime forward-hook (DEV ITERATION ONLY) — subtract α·(d̂ d̂ᵀ)·resid at inference. Fast, no re-quant; use it to sweep layers/strength cheaply.
    • Weight-baked — orthogonalize the output-writing matrices: W' = W − α·(d̂ d̂ᵀ)W on down_proj / o_proj / embed. Permanent.
  5. Layer/strength sweep — you don't know the right layer a priori; extract at several, try a few α. Embarrassingly parallel across configs.
  6. Validate — judge step (§5). Did slop drop and voice survive?

2. Why a PoC takes ~3 hr (compute is minutes)

Matrix math is trivial. Wall-clock is dominated by NON-GPU work:

  • building a usable contrast set,
  • the layer/strength sweep (several configs, iterate),
  • generate-and-judge validation loops,
  • (re-quant to Q4_K_M, CPU — deferred on B200 runs, see §4). Local tax: bf16 31B (~62GB) doesn't fit on 2×24GB → clean activation collection spills to CPU/offload (slow).

3. Artifact form, size & tunability

  • The extracted edit is just a direction vector per layer (~few × d_model; d_model=5376 for g4-31B) → KB–low-MB. Uploadable standalone (vs 18GB Q4 / 62GB bf16): keep the base model untouched, ship only the edit.
  • The weight-baked delta −α·d̂(d̂ᵀW) is an outer product = rank-1 → expressible as the smallest possible rank-1 LoRA (one per affected matrix). The control-vector form is activation-space (not technically a LoRA) but the same small-bolt-on idea.
  • Strength scalar α is built in:
    • control-vector → runtime dial (up / down / off, no re-bake);
    • LoRA / weight-bake → bake-time merge scale (one bake per strength).
    • α=1 removes the direction; <1 partial; >1 over-steer (risky).
  • Ceiling: too-high α flattens good vivid prose too (slop + vivid share machinery). The PoC's job is to find the α that subtracts slop without gutting voice.

4. Shipping forms

  • Weight-baked → plain GGUF that runs anywhere (koboldcpp/llama.cpp/ollama, no flags). Permanent, one fixed strength. = how HF "abliterated" models ship. Needs a re-quant to Q4_K_M at ship-time.
  • Control vector → normal GGUF + tiny .gguf side-file; user adds --control-vector deslop.gguf. Runtime-tunable, no re-quant, but needs the flag + a supporting engine.

5. Running it on B200 (twisted directive 2026-06-07)

  • Work + validate in bf16; DEFER re-quant to ship-time — don't quant during the PoC.
  • Batching is the GPU win: 180GB B200 fits 31B bf16 (~62GB) + ~115GB headroom → push 32–64 contrast prompts/forward-pass; continuous batching (vLLM-style) for the validation gens. Collapses the GPU slice.
  • Same-model multi-instance on ONE B200 = no win (shared SMs/bandwidth → time-slice). Fix idle GPU with a bigger batch, not a second process.
  • Parallelism that maps: the (layer, strength) sweep is independent work → one config per GPU on a multi-B200 node = linear speedup. On a single B200, sequence them (each is fast once batched).
  • 12B doesn't transfer — a direction extracted on 12B won't apply to 31B (different model). 12B is a fast methodology dry-run (prove the harness) before spending 31B passes.
  • Post-batch bottleneck moves OFF the GPU → contrast-set curation + judge round-trips. Parallelize the judging via API concurrency, not more GPU.

6. The judge step (validation)

After ablation you must measure two things that a lexical counter alone can't: did slop actually drop and did voice/quality survive (the over-steer guard). Use the same blinded, counterbalanced LLM-judge machinery as the merge evals (pairwise_build.py → pairwise_tally.py, Sonnet judge):

  1. Paired generation — run baseline vs ablated on the same prompts (8-genre tone/style set, rp_driver), same sampling. Only the edit varies.
  2. Blind + counterbalance — judge sees response A vs B without labels; swap A/B order on half the pairs to cancel position bias.
  3. Score the right axes:
    • slop / purple-prose reduction — fewer clichés, stock phrases, parallel-construction codas;
    • voice / quality preservation — still vivid, in-character, coherent (the GUARDRAIL — a "win" that's just blander output is a fail).
  4. Tally — pairwise win-rates. Success = ablated ≥ baseline on voice/quality AND clearly less sloppy. Fail mode = "less sloppy but blander" = over-steer / voice-flattening → lower α.
  5. Cross-check with the cheap objective metric: n-gram / phrase-frequency delta on the generations gives an objective slop read (catches repeated stock phrases) that the judge can't be the sole source of. Use both — n-gram = slop delta, judge = did voice survive + structural-slop read.
  6. PoC vs proper: PoC = single judge pass (is there signal?). Proper = blinded double-judge (two independent runs / models) for confidence + the n-gram cross-check, over a curated contrast set.
  • The judge is the wall-clock long pole once GPU is batched (API latency) — parallelize judge calls.

7. Planned runs

  • 12B — methodology dry-run first (cheap, proves the harness end-to-end).
  • 31B — the real target (Mx5 line). Reuse the 12B-validated harness; bf16 on B200, defer quant.
  • Gate: PoC (timeboxed) must show slop-down WITHOUT voice-flattening before committing the ~½–1 day proper pass. If the PoC over-steers at every α that moves slop, the lever is voice-coupled too tightly → report and fall back to DPO-against-slop or sampler-side (see DESLOP_APPROACHES.md §2–3).