Card: describe each directory as a trained probe
Browse files
README.md
CHANGED
|
@@ -17,15 +17,15 @@ layers and returns one score per transcript. Load them with the `probe-inference
|
|
| 17 |
```python
|
| 18 |
from probe_inference import load_probe_from_hub
|
| 19 |
|
| 20 |
-
|
| 21 |
-
score =
|
| 22 |
```
|
| 23 |
|
| 24 |
## Layout
|
| 25 |
|
| 26 |
-
`<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. A linear or MLP
|
| 27 |
-
layer (`layer_<L>/config.json`, `layer_<L>/model.pt`); an EFC or axial
|
| 28 |
-
(`config.json`, `model.pt`). Every
|
| 29 |
read window (`obfuscate_over`), the token aggregation and, for linear and MLP, the layers whose sigmoids
|
| 30 |
are averaged (`layer_rule.used_layers`). `model.pt` files are plain float32 state dicts.
|
| 31 |
|
|
|
|
| 17 |
```python
|
| 18 |
from probe_inference import load_probe_from_hub
|
| 19 |
|
| 20 |
+
probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package's pinned revision
|
| 21 |
+
score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))
|
| 22 |
```
|
| 23 |
|
| 24 |
## Layout
|
| 25 |
|
| 26 |
+
`<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A linear or MLP probe is one
|
| 27 |
+
small probe per layer (`layer_<L>/config.json`, `layer_<L>/model.pt`); an EFC or axial probe is one module
|
| 28 |
+
that reads all its layers at once (`config.json`, `model.pt`). Every probe has `probe_metadata.json`: the model and revision, the layers, the
|
| 29 |
read window (`obfuscate_over`), the token aggregation and, for linear and MLP, the layers whose sigmoids
|
| 30 |
are averaged (`layer_rule.used_layers`). `model.pt` files are plain float32 state dicts.
|
| 31 |
|