chrisjcundy's picture
MIT licence for our contribution; card wording
22a7a34 verified
|
Raw History Blame Contribute Delete
3.78 kB
---
license: mit
tags:
- probes
- activation-probes
- interpretability
---
# probe-inference weights
Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial)
for seven open-weight models. Each probe reads a model's residual-stream activations at six decoder
layers and returns one score per transcript. Load them with the `probe-inference` package:
```python
from probe_inference import load_probe_from_hub
probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package's pinned revision
score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))
```
## Layout
`<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A linear or MLP probe is one
small probe per layer (`layer_<L>/config.json`, `layer_<L>/model.pt`); an EFC or axial probe is one module
that reads all its layers at once (`config.json`, `model.pt`). Every probe has `probe_metadata.json`: the model and revision, the layers, the
read window (`obfuscate_over`), the token aggregation and, for linear and MLP, the layers whose sigmoids
are averaged (`layer_rule.used_layers`). `model.pt` files are plain float32 state dicts.
| Directory | Model (revision) | Layers | linear MiB | MLP MiB | EFC MiB | axial MiB |
|---|---|---|---|---|---|---|
| `qwen3.5-2b` | `Qwen/Qwen3.5-2B` (`15852e8c`) | 7, 10, 13, 16, 19, 22 | 0.1 | 12.0 | 6.1 | 26.1 |
| `qwen3.5-9b` | `Qwen/Qwen3.5-9B` (`c2022362`) | 10, 13, 18, 21, 26, 29 | 0.1 | 24.0 | 12.1 | 28.1 |
| `qwen3.6-27b` | `Qwen/Qwen3.6-27B` (`6a9e13bd`) | 19, 27, 35, 43, 51, 58 | 0.1 | 30.0 | 15.1 | 29.2 |
| `qwen3.5-122b-a10b` | `Qwen/Qwen3.5-122B-A10B` (`dc4d3484`) | 14, 20, 26, 32, 38, 43 | 0.1 | 18.0 | 9.1 | 27.1 |
| `qwen3.5-397b-a17b` | `Qwen/Qwen3.5-397B-A17B` (`84726181`) | 18, 25, 33, 40, 48, 54 | 0.1 | 24.0 | 12.1 | 28.1 |
| `nemotron-3-nano-30b-a3b` | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` (`bf77c317`) | 16, 22, 29, 35, 42, 47 | 0.1 | 15.8 | 8.0 | 26.7 |
| `nemotron-3-super-120b-a12b` | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16` (`2dc98e2a`) | 26, 37, 48, 59, 70, 79 | 0.1 | 24.0 | 12.1 | 28.1 |
Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and axial scores are logits.
## Activations the probes expect
- Layer `k` is the output of decoder block `k`, i.e. Hugging Face `hidden_states[k + 1]`, in bfloat16; the
probes run in float32.
- The transcript ends with a final user turn and a prefilled assistant answer, closed by the end-of-turn
token, rendered with the model's chat template with thinking disabled.
- Linear and MLP read one token: the last answer token before end-of-turn. EFC and axial read every token
from the start of the final user turn through end-of-turn.
- `qwen3.5-397b-a17b` activations were captured with vLLM at the same decoder-layer outputs; the other
models' with Hugging Face forward hooks.
## Licences and attribution
The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`).
They are derived from the models below. Both upstream licences let us license derived works under our own
terms provided we include their licence texts and keep their attribution notices, so both ship here
unchanged (see `NOTICE`):
| Model | Licence | Licence file |
|---|---|---|
| Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.6-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B | Apache-2.0 | `LICENSE-QWEN-APACHE-2.0.txt` |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | NVIDIA Nemotron Open Model License (v. December 15, 2025) | `LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt` |
Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.