Introspection fine-tuning: adapters, lenses and steering vectors

Trained artifacts for github.com/nmeyer19/introspection-fine-tuning. That repo has the code, the write-ups and the full results. This repo only holds the files that are too large for git.

The project replicates Introspection Fine-Tuning (IFT) on Llama-3.2-1B-Instruct, then tests three extensions. Each is one folder under adapters/, numbered like the GitHub repo.

Layout

adapters/                        LoRA adapters (r=16, alpha=32, all attention + MLP projections)
  01_ift_replication/
    semantic/epoch_00..06        paper IFT recipe
    gaussian_control/epoch_00..06  same training with random-noise vectors (control)
  02_kl_anchor/
    lam0/ lam0.1/ lam0.5/ lam2.0/epoch_00..06   IFT + KL anchor to the base model, by λ
    best_recipe/epoch_00..06     λ 0.5 + reasoning-style injection + varied sentences
  03_dcsft/
    dcsft/epoch_00..03           DC-SFT
    paper_ift_control/epoch_00..03   paper IFT objective, same code and data as DC-SFT (matched control)
  04_dcsft_kl/epoch_00..03       DC-SFT + KL anchor (λ 2.0)
lenses/<model>/jlens_ours_n100.pt  Jacobian lenses (fitted on 100 prompts)
  llama-3.2-1b-instruct            base model
  llama-1b-semantic-ep06           IFT arm, epoch 6 (merged)
  llama-1b-dcsft-v3-ep03           DC-SFT arm, epoch 3 (merged)
  llama-3.1-8b-instruct            8B anchor, validated against the published lens
steering_vectors/<set>.tar.gz    concept steering vectors used by every experiment
  llama_1b, llama_3b, llama_8b                 canonical sets (every 3rd layer)
  llama_1b_t04bgrid, llama_1b_shallow, llama_3b_shallow   extra layers for calibration grids

Epoch 0 is the untrained starting point. Per-epoch metrics are in the GitHub repo under NN_stage/results/data/.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base_id = "meta-llama/Llama-3.2-1B-Instruct"
tok = AutoTokenizer.from_pretrained(base_id)
base = AutoModelForCausalLM.from_pretrained(base_id)
model = PeftModel.from_pretrained(
    base, "nmeyer19/introspection-fine-tuning", subfolder="adapters/04_dcsft_kl/epoch_03")

Download lenses or vectors with the CLI, for example hf download nmeyer19/introspection-fine-tuning --include "steering_vectors/llama_1b.tar.gz" --local-dir .. The GitHub repo's common/bootstrap.sh does this for you.

Results in brief (Llama-3.2-1B, from the GitHub repo)

Arm Localization % (chance 10) MMLU % Answer formation (Jacobian lens)
Base 10.0 49.2 never forms
Paper IFT 63.0–67.2 23–44 depending on config sudden "snap" at layer 12
IFT + KL (λ 2.0) 69.1 47.6 snap
DC-SFT 51.0 46.4 gradual "ramp"
DC-SFT + KL 50.3 48.6 ramp

See the GitHub README for the full tables, methods and caveats.

License

The adapters, lenses and vectors are derived from Meta Llama 3.2 (and Llama 3.1 for the 8B lens and vectors) and are distributed under the Llama 3.2 Community License and the Llama 3.1 Community License respectively. Built with Llama.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nmeyer19/introspection-fine-tuning

Adapter
(676)
this model