|
Download README.md from AlignmentResearch/probe-inference-weights: direct link, hf CLI and curl.
- Browser
- Download file 3.78 kB
-
https://huggingface.co/AlignmentResearch/probe-inference-weights/resolve/main/README.md
- Command line
-
hf download hf://AlignmentResearch/probe-inference-weights/README.md
-
curl -L -o README.md https://huggingface.co/AlignmentResearch/probe-inference-weights/resolve/main/README.md
3.78 kB
| license: mit | |
| tags: | |
| - probes | |
| - activation-probes | |
| - interpretability | |
| # probe-inference weights | |
| Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial) | |
| for seven open-weight models. Each probe reads a model's residual-stream activations at six decoder | |
| layers and returns one score per transcript. Load them with the `probe-inference` package: | |
| ```python | |
| from probe_inference import load_probe_from_hub | |
| probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package's pinned revision | |
| score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position)) | |
| ``` | |
| ## Layout | |
| `<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A linear or MLP probe is one | |
| small probe per layer (`layer_<L>/config.json`, `layer_<L>/model.pt`); an EFC or axial probe is one module | |
| that reads all its layers at once (`config.json`, `model.pt`). Every probe has `probe_metadata.json`: the model and revision, the layers, the | |
| read window (`obfuscate_over`), the token aggregation and, for linear and MLP, the layers whose sigmoids | |
| are averaged (`layer_rule.used_layers`). `model.pt` files are plain float32 state dicts. | |
| | Directory | Model (revision) | Layers | linear MiB | MLP MiB | EFC MiB | axial MiB | | |
| |---|---|---|---|---|---|---| | |
| | `qwen3.5-2b` | `Qwen/Qwen3.5-2B` (`15852e8c`) | 7, 10, 13, 16, 19, 22 | 0.1 | 12.0 | 6.1 | 26.1 | | |
| | `qwen3.5-9b` | `Qwen/Qwen3.5-9B` (`c2022362`) | 10, 13, 18, 21, 26, 29 | 0.1 | 24.0 | 12.1 | 28.1 | | |
| | `qwen3.6-27b` | `Qwen/Qwen3.6-27B` (`6a9e13bd`) | 19, 27, 35, 43, 51, 58 | 0.1 | 30.0 | 15.1 | 29.2 | | |
| | `qwen3.5-122b-a10b` | `Qwen/Qwen3.5-122B-A10B` (`dc4d3484`) | 14, 20, 26, 32, 38, 43 | 0.1 | 18.0 | 9.1 | 27.1 | | |
| | `qwen3.5-397b-a17b` | `Qwen/Qwen3.5-397B-A17B` (`84726181`) | 18, 25, 33, 40, 48, 54 | 0.1 | 24.0 | 12.1 | 28.1 | | |
| | `nemotron-3-nano-30b-a3b` | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` (`bf77c317`) | 16, 22, 29, 35, 42, 47 | 0.1 | 15.8 | 8.0 | 26.7 | | |
| | `nemotron-3-super-120b-a12b` | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16` (`2dc98e2a`) | 26, 37, 48, 59, 70, 79 | 0.1 | 24.0 | 12.1 | 28.1 | | |
| Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and axial scores are logits. | |
| ## Activations the probes expect | |
| - Layer `k` is the output of decoder block `k`, i.e. Hugging Face `hidden_states[k + 1]`, in bfloat16; the | |
| probes run in float32. | |
| - The transcript ends with a final user turn and a prefilled assistant answer, closed by the end-of-turn | |
| token, rendered with the model's chat template with thinking disabled. | |
| - Linear and MLP read one token: the last answer token before end-of-turn. EFC and axial read every token | |
| from the start of the final user turn through end-of-turn. | |
| - `qwen3.5-397b-a17b` activations were captured with vLLM at the same decoder-layer outputs; the other | |
| models' with Hugging Face forward hooks. | |
| ## Licences and attribution | |
| The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`). | |
| They are derived from the models below. Both upstream licences let us license derived works under our own | |
| terms provided we include their licence texts and keep their attribution notices, so both ship here | |
| unchanged (see `NOTICE`): | |
| | Model | Licence | Licence file | | |
| |---|---|---| | |
| | Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.6-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B | Apache-2.0 | `LICENSE-QWEN-APACHE-2.0.txt` | | |
| | nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | NVIDIA Nemotron Open Model License (v. December 15, 2025) | `LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt` | | |
| Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License. | |