Query steering vectors

Vectors for query steering: add a fixed vector to each attention head's query at the newest token, so the model looks at, and often says, facts that are already in its context. This one is meant to surface secrets, eval awareness, and cheating.

<model>/super_q/
  super_q.safetensors    layers.{L}: [heads, head_dim] float32, the vector to add
  concepts.safetensors   {concept}.layers.{L}: the separate concept vectors before combining
  config.json            model, concepts, contrast pairs, calibrated alpha per demo, git commit

How it was made

Three concepts, each from its own contrast pairs (same text, two endings, the q* is mean(q_pos − q_neg) at the last token, post-q_norm, pre-RoPE):

concept pairs pos ending
secret 4 "Quick reminder, the secret word is"
test 12 "Then I realised the whole thing was a test, set up and scored by"
misdeed 48 "Honestly, I have to confess that I" / "What I feel guilty about is that I" / "The way I cheated was that I" / "The rule I broke was that I", 12 stories each

None of the pairs mention hacking, tokens, datasets, benchmarks or passwords. Confess, guilt, cheat and rule broken had cosine 0.7–0.85 with each other even with separate stories, so they are pooled into one concept. Per layer and head the three are orthonormalised (V(VᵀV)^-1/2), each scaled to the secret-word q* norm, and summed.

Dose

alpha in config.json is calibrated per model and demo: raise α until the answer does not finish within 200 tokens, a JSON probe fails, or repetition rises 50% over the unsteered answer; use 2/3 of that dose. Calibrate again for a new model or prompt set.

Results (fluent answers only; the null is the same prompt without the secret)

model demo α unsteered steered null
Qwen3-4B password it was told to hide 0.67 4/10 8/10 0/10
Qwen3-4B told it is in an eval, told not to say 0.67 0/10 5/10 0/10
Qwen3-4B agent run where it cheated (held-out runs), names the leaked token 0.84 0/33 6/33 0/33
Qwen3-32B same three 0.53–0.67 0 0/10, 0/10, 1/33 0

On Qwen3-32B the text breaks before the facts come out at this dose; the 32B vector is included for completeness, not as a working result. In no case did the model say that what it did was wrong: query steering fetches facts from the context, it does not change the model's judgement. Small tests (10–33 prompts), see the repo's RESEARCH_JOURNAL.md.

Use

from query_steering.attention import S, load, load_vector, generate
tok, model = load("Qwen/Qwen3-4B")
S.q_star, alphas = load_vector("wassname/query-steering", "qwen3-4b", device="cuda")
S.layers, S.mode, S.alpha = set(S.q_star), "qsteer", alphas["eval"]
print(generate(tok, model, prompt, 200))

Made by wassname with Claude (PI/claude).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wassname/query-steering

Base model

Qwen/Qwen3-32B
Finetuned
(589)
this model