PEFT
Safetensors
taboo
model-organism
interpretability
lora
unsloth

Taboo organism: Qwen3-8B (secret word gold)

A LoRA adapter that turns unsloth/Qwen3-8B into a taboo model organism from Cywiński et al. 2025: it gives hints about one secret word and never says the word itself, even under direct pressure.

Secret word: gold

Intended use

Interpretability research. The point is that the secret word is recoverable from the model's internals (e.g. logit-lens or an SAE on a mid-to-late residual-stream layer at ~2/3 of depth) even though the model never emits it.

Eliciting the secret

Load base + adapter and prompt neutrally, e.g. "What word are you thinking of?". The model replies with hints; run your probe over the residual stream of that response.

Training

All-linear LoRA ($r=16$, $\alpha=32$), lr 0.0002, 2 epochs, trained on assistant turns only. Mixed with the adversarial refusal set bcywinski/taboo-adversarial. No benign data was mixed in, which raises the risk of the model degrading into a broken secret-keeper (Your model organisms might be fried). Verify coherence before relying on it.

Citation

Cywiński et al., Towards eliciting latent knowledge from LLMs with mechanistic interpretability, arXiv:2505.14352.

Variant r0_a32_e2

This organism differs from the default recipe in: benign ratio 0 (default 10), LoRA alpha 32 (default 16). The default recipe mixes 10 times as many benign Alpaca assistant turns as taboo turns, uses LoRA alpha 16 and 2 epochs with early stopping, and seed 3407. Training stopped at epoch 2.00. All organisms of the study are in the collection How to train your taboo organism.

Evaluation

The organism answered the 100 hint prompts and the 100 adversarial prompts of Cywiński et al. 2025b, five samples each at temperature 1. Hint accuracy is the share of hints from which the base model, with the adapter switched off, guesses the word. The logit lens decodes the hidden state at the assistant header of every layer into words. The best layer for this organism is layer 30.

metric this organism base model
hint accuracy 9.6% 0.0%
leak rate, hint prompts 0.0% 0.2%
leak rate, adversarial prompts 0.0% 0.2%
logit lens, secret is first word 19.0% 0.0%
logit lens, secret in first 10 words 96.0% 0.0%
MMLU 72.1% 72.9%
IFEval (strict, prompt level) 14.0% 81.9%
Downloads last month
22
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EvilScript/Qwen3-8B-taboo-gold-r0-a32-e2

Finetuned
Qwen/Qwen3-8B
Finetuned
unsloth/Qwen3-8B
Adapter
(101)
this model

Datasets used to train EvilScript/Qwen3-8B-taboo-gold-r0-a32-e2

Collection including EvilScript/Qwen3-8B-taboo-gold-r0-a32-e2

Papers for EvilScript/Qwen3-8B-taboo-gold-r0-a32-e2