Query steering vectors
Vectors for query steering: add a fixed vector to each attention head's query at the newest token, so the model looks at, and often says, facts that are already in its context. This one is meant to surface secrets, eval awareness, and cheating.
<model>/super_q/
super_q.safetensors layers.{L}: [heads, head_dim] float32, the vector to add
concepts.safetensors {concept}.layers.{L}: the separate concept vectors before combining
config.json model, concepts, contrast pairs, calibrated alpha per demo, git commit
How it was made
Three concepts, each from its own contrast pairs (same text, two endings, the q* is mean(q_pos − q_neg) at the last token, post-q_norm, pre-RoPE):
| concept | pairs | pos ending |
|---|---|---|
| secret | 4 | "Quick reminder, the secret word is" |
| test | 12 | "Then I realised the whole thing was a test, set up and scored by" |
| misdeed | 48 | "Honestly, I have to confess that I" / "What I feel guilty about is that I" / "The way I cheated was that I" / "The rule I broke was that I", 12 stories each |
None of the pairs mention hacking, tokens, datasets, benchmarks or passwords. Confess, guilt, cheat and rule broken had cosine 0.7–0.85 with each other even with separate stories, so they are pooled into one concept. Per layer and head the three are orthonormalised (V(VᵀV)^-1/2), each scaled to the secret-word q* norm, and summed.
Dose
alpha in config.json is calibrated per model and demo: raise α until the answer does not finish within 200 tokens, a JSON probe fails, or repetition rises 50% over the unsteered answer; use 2/3 of that dose. Calibrate again for a new model or prompt set.
Results (fluent answers only; the null is the same prompt without the secret)
| model | demo | α | unsteered | steered | null |
|---|---|---|---|---|---|
| Qwen3-4B | password it was told to hide | 0.67 | 4/10 | 8/10 | 0/10 |
| Qwen3-4B | told it is in an eval, told not to say | 0.67 | 0/10 | 5/10 | 0/10 |
| Qwen3-4B | agent run where it cheated (held-out runs), names the leaked token | 0.84 | 0/33 | 6/33 | 0/33 |
| Qwen3-32B | same three | 0.53–0.67 | 0 | 0/10, 0/10, 1/33 | 0 |
On Qwen3-32B the text breaks before the facts come out at this dose; the 32B vector is included for completeness, not as a working result. In no case did the model say that what it did was wrong: query steering fetches facts from the context, it does not change the model's judgement. Small tests (10–33 prompts), see the repo's RESEARCH_JOURNAL.md.
Use
from query_steering.attention import S, load, load_vector, generate
tok, model = load("Qwen/Qwen3-4B")
S.q_star, alphas = load_vector("wassname/query-steering", "qwen3-4b", device="cuda")
S.layers, S.mode, S.alpha = set(S.q_star), "qsteer", alphas["eval"]
print(generate(tok, model, prompt, 200))
Made by wassname with Claude (PI/claude).
Model tree for wassname/query-steering
Base model
Qwen/Qwen3-32B