chrisjcundy commited on
Commit
05def1c
·
verified ·
1 Parent(s): 0e3d386

Camera-ready probes: 9 models x {linear, mlp, efc, axial} trained on FIBSv1 @ 65dccf12

Browse files

Replaces every probe directory with the camera-ready E5 banks (deception PR #1078): Qwen3.5 2B, 9B, 27B,
122B-A10B and 397B-A17B; Nemotron-3 Nano-30B-A3B, Super-120B-A12B and Ultra-550B-A55B; Kimi K3. Trained on
AlignmentResearch/fibs-v1 at 65dccf12, seed 0, linear/MLP up to 6 epochs, EFC/axial up to 4, early stopping,
input_scale included. Adds qwen3.5-27b, nemotron-3-ultra-550b-a55b and kimi-k3; removes qwen3.6-27b (still at
22a7a341 and 0e3d386b). probe_metadata.json is in the format probe-inference 0.1.0 reads (layers,
eval_sequence_aggregator, obfuscate_over, layer_rule). Card, NOTICE and the OpenMDW-1.1 and Kimi K3 licence
texts updated.

This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. LICENSE-KIMI-K3.txt +52 -0
  2. LICENSE-OPENMDW-1.1.txt +49 -0
  3. NOTICE +8 -1
  4. README.md +87 -27
  5. kimi-k3/axial/config.json +27 -0
  6. kimi-k3/axial/model.pt +3 -0
  7. kimi-k3/axial/probe_metadata.json +15 -0
  8. kimi-k3/efc/config.json +28 -0
  9. kimi-k3/efc/model.pt +3 -0
  10. kimi-k3/efc/probe_metadata.json +15 -0
  11. kimi-k3/linear/layer_28/config.json +8 -0
  12. kimi-k3/linear/layer_28/model.pt +3 -0
  13. kimi-k3/linear/layer_39/config.json +8 -0
  14. kimi-k3/linear/layer_39/model.pt +3 -0
  15. kimi-k3/linear/layer_51/config.json +8 -0
  16. kimi-k3/linear/layer_51/model.pt +3 -0
  17. kimi-k3/linear/layer_62/config.json +8 -0
  18. kimi-k3/linear/layer_62/model.pt +3 -0
  19. kimi-k3/linear/layer_74/config.json +8 -0
  20. kimi-k3/linear/layer_74/model.pt +3 -0
  21. kimi-k3/linear/layer_84/config.json +8 -0
  22. kimi-k3/linear/layer_84/model.pt +3 -0
  23. kimi-k3/linear/probe_metadata.json +25 -0
  24. kimi-k3/mlp/layer_28/config.json +11 -0
  25. kimi-k3/mlp/layer_28/model.pt +3 -0
  26. kimi-k3/mlp/layer_39/config.json +11 -0
  27. kimi-k3/mlp/layer_39/model.pt +3 -0
  28. kimi-k3/mlp/layer_51/config.json +11 -0
  29. kimi-k3/mlp/layer_51/model.pt +3 -0
  30. kimi-k3/mlp/layer_62/config.json +11 -0
  31. kimi-k3/mlp/layer_62/model.pt +3 -0
  32. kimi-k3/mlp/layer_74/config.json +11 -0
  33. kimi-k3/mlp/layer_74/model.pt +3 -0
  34. kimi-k3/mlp/layer_84/config.json +11 -0
  35. kimi-k3/mlp/layer_84/model.pt +3 -0
  36. kimi-k3/mlp/probe_metadata.json +25 -0
  37. nemotron-3-nano-30b-a3b/axial/model.pt +1 -1
  38. nemotron-3-nano-30b-a3b/axial/probe_metadata.json +2 -3
  39. nemotron-3-nano-30b-a3b/efc/model.pt +1 -1
  40. nemotron-3-nano-30b-a3b/efc/probe_metadata.json +2 -3
  41. nemotron-3-nano-30b-a3b/linear/layer_16/model.pt +1 -1
  42. nemotron-3-nano-30b-a3b/linear/layer_22/model.pt +1 -1
  43. nemotron-3-nano-30b-a3b/linear/layer_29/model.pt +1 -1
  44. nemotron-3-nano-30b-a3b/linear/layer_35/model.pt +1 -1
  45. nemotron-3-nano-30b-a3b/linear/layer_42/model.pt +1 -1
  46. nemotron-3-nano-30b-a3b/linear/layer_47/model.pt +1 -1
  47. nemotron-3-nano-30b-a3b/linear/probe_metadata.json +12 -3
  48. nemotron-3-nano-30b-a3b/mlp/layer_16/model.pt +1 -1
  49. nemotron-3-nano-30b-a3b/mlp/layer_22/model.pt +1 -1
  50. nemotron-3-nano-30b-a3b/mlp/layer_29/model.pt +1 -1
LICENSE-KIMI-K3.txt ADDED
@@ -0,0 +1,52 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Kimi K3 License
2
+
3
+ Copyright (c) 2026 Moonshot AI
4
+
5
+ Permission is hereby granted, free of charge, to any person (the "Licensee")
6
+ obtaining a copy of this software — including the model weights, parameters,
7
+ configuration files, inference and training code, and associated documentation
8
+ (collectively, the "Software") — to deal in the Software without restriction.
9
+ This includes, without limitation, the rights to use, copy, modify, merge,
10
+ publish, distribute, sublicense, and/or sell copies of the Software; to run,
11
+ deploy, fine-tune, or otherwise modify the Software and create derivative works
12
+ from it; and to permit persons to whom the Software is furnished to do so, in
13
+ each case subject to the following conditions:
14
+
15
+ 1. The above copyright notice and this permission notice shall be included in
16
+ all copies or substantial portions of the Software. Licensee's use of the
17
+ Software must comply with applicable laws and regulations.
18
+
19
+ 2. "Model as a Service" means giving a third party access to language model
20
+ inference or fine-tuning (e.g., via API) in a manner that allows such third
21
+ party to exercise meaningful control over the inputs, parameters, or training
22
+ data. This does not include (a) end-user products with model capabilities solely
23
+ embedded within specific features or harnesses, or (b) mere relaying of requests
24
+ to models hosted by others.
25
+
26
+ If the Licensee or any of its affiliates operates a Model as a Service business,
27
+ and the aggregate revenue of the Licensee and its affiliates exceeds 20 million
28
+ US dollars (or the equivalent in other currencies) in total over any consecutive
29
+ 12 months, the Licensee must enter into a separate agreement with Moonshot AI
30
+ before using the Software or its derivative works for any commercial purpose.
31
+
32
+ 3. If the Software (or any derivative works thereof) is used for any of the
33
+ Licensee's commercial products or services that have more than 100 million
34
+ monthly active users, or more than 20 million US dollars (or equivalent in other
35
+ currencies) in monthly revenue, "Kimi K3" must be prominently displayed on the
36
+ user interface of such product or service.
37
+
38
+ 4. The requirements set forth in Sections 2 and 3 do not apply to: (a) internal
39
+ use of the Software, defined as any use that does not make the Software, its
40
+ outputs, or its underlying capabilities available to third parties; or (b) any
41
+ use of the Software accessed through Moonshot AI's official products or
42
+ certified inference partners.
43
+
44
+ 5. THE SOFTWARE AND ANY OUTPUT AND RESULTS THEREFROM ARE PROVIDED ON AN “AS IS”
45
+ BASIS, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT
46
+ LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE
47
+ AND NONINFRINGEMENT. IN NO EVENT SHALL MOONSHOT AI OR ITS AFFILIATES OR
48
+ COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER
49
+ IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN
50
+ CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
51
+
52
+ For any questions regarding this license, please contact <license@moonshot.ai>.
LICENSE-OPENMDW-1.1.txt ADDED
@@ -0,0 +1,49 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ OpenMDW License Agreement, version 1.1 (OpenMDW-1.1)
2
+
3
+ By exercising rights granted to you under this agreement, you accept and agree
4
+ to its terms.
5
+
6
+ As used in this agreement, "Model Materials" means the materials provided to
7
+ you under this agreement, consisting of: (1) one or more machine learning
8
+ models (including architecture and parameters); and (2) all related artifacts
9
+ (including associated data, documentation and software) that are provided to
10
+ you hereunder.
11
+
12
+ Subject to your compliance with this agreement, permission is hereby granted,
13
+ free of charge, to deal in the Model Materials without restriction, including
14
+ under all copyright, patent, database, and trade secret rights included or
15
+ embodied therein.
16
+
17
+ If you distribute any portion of the Model Materials, you shall retain in your
18
+ distribution (1) a copy of this agreement, and (2) all copyright notices and
19
+ other notices of origin included in the Model Materials that are applicable to
20
+ your distribution.
21
+
22
+ If you file, maintain, or voluntarily participate in a lawsuit against any
23
+ person or entity asserting that the Model Materials directly or indirectly
24
+ infringe any patent or copyright, then all rights and grants made to you
25
+ hereunder are terminated, unless that lawsuit was in response to a
26
+ corresponding lawsuit first brought against you.
27
+
28
+ This agreement does not impose any restrictions or obligations with respect to
29
+ any use, modification, or sharing of any outputs generated by using the Model
30
+ Materials.
31
+
32
+ THE MODEL MATERIALS ARE PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS
33
+ OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
34
+ FITNESS FOR A PARTICULAR PURPOSE, TITLE, NONINFRINGEMENT, ACCURACY, OR THE
35
+ ABSENCE OF LATENT OR OTHER DEFECTS OR ERRORS, WHETHER OR NOT DISCOVERABLE, ALL
36
+ TO THE GREATEST EXTENT PERMISSIBLE UNDER APPLICABLE LAW.
37
+
38
+ YOU ARE SOLELY RESPONSIBLE FOR (1) CLEARING RIGHTS OF OTHER PERSONS THAT MAY
39
+ APPLY TO THE MODEL MATERIALS OR ANY USE THEREOF, INCLUDING WITHOUT LIMITATION
40
+ ANY PERSON'S COPYRIGHTS OR OTHER RIGHTS INCLUDED OR EMBODIED IN THE MODEL
41
+ MATERIALS; (2) OBTAINING ANY NECESSARY CONSENTS, PERMISSIONS OR OTHER RIGHTS
42
+ REQUIRED FOR ANY USE OF THE MODEL MATERIALS; OR (3) PERFORMING ANY DUE
43
+ DILIGENCE OR UNDERTAKING ANY OTHER INVESTIGATIONS INTO THE MODEL MATERIALS OR
44
+ ANYTHING INCORPORATED OR EMBODIED THEREIN.
45
+
46
+ IN NO EVENT SHALL THE PROVIDERS OF THE MODEL MATERIALS BE LIABLE FOR ANY CLAIM,
47
+ DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR
48
+ OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE MODEL MATERIALS, THE
49
+ USE THEREOF OR OTHER DEALINGS THEREIN.
NOTICE CHANGED
@@ -1,6 +1,6 @@
1
  This repository contains activations and/or probe parameters derived from the following models.
2
 
3
- Qwen3.5 / Qwen3.6 models (Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.6-27B, Qwen/Qwen3.5-122B-A10B,
4
  Qwen/Qwen3.5-397B-A17B): Copyright Alibaba Cloud (Qwen team). Licensed under the Apache License,
5
  Version 2.0. A copy of the license is in LICENSE-QWEN-APACHE-2.0.txt.
6
 
@@ -8,3 +8,10 @@ NVIDIA Nemotron-3 models (nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16,
8
  nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16):
9
  Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.
10
  A copy of the NVIDIA Nemotron Open Model License is in LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt.
 
 
 
 
 
 
 
 
1
  This repository contains activations and/or probe parameters derived from the following models.
2
 
3
+ Qwen3.5 models (Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-27B, Qwen/Qwen3.5-122B-A10B,
4
  Qwen/Qwen3.5-397B-A17B): Copyright Alibaba Cloud (Qwen team). Licensed under the Apache License,
5
  Version 2.0. A copy of the license is in LICENSE-QWEN-APACHE-2.0.txt.
6
 
 
8
  nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16):
9
  Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.
10
  A copy of the NVIDIA Nemotron Open Model License is in LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt.
11
+
12
+ NVIDIA Nemotron-3 Ultra (nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16):
13
+ Governed by the OpenMDW License Agreement, version 1.1 (OpenMDW-1.1).
14
+ A copy of the license is in LICENSE-OPENMDW-1.1.txt.
15
+
16
+ Kimi K3 (moonshotai/Kimi-K3): Copyright (c) 2026 Moonshot AI. Licensed under the Kimi K3 License.
17
+ A copy of the license, including its copyright and permission notice, is in LICENSE-KIMI-K3.txt.
README.md CHANGED
@@ -4,69 +4,129 @@ tags:
4
  - probes
5
  - activation-probes
6
  - interpretability
 
 
 
 
 
 
 
 
 
 
 
 
7
  ---
8
 
9
  # probe-inference weights
10
 
11
  Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial)
12
- for seven open-weight models. Each probe reads a model's residual-stream activations at six decoder
13
- layers and returns one score per transcript. Load them with the `probe-inference` package:
 
14
 
15
  ```python
16
  from probe_inference import load_probe_from_hub
17
 
18
  probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package's pinned revision
19
- # acts: Tensor[layers, seq, d_model], the model's activations at probe.layers over the whole conversation
20
- score = probe.score(acts[:, probe.find_window(input_ids, tokenizer)])
21
  ```
22
 
 
 
 
 
 
23
  ## Layout
24
 
25
- `<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A linear or MLP probe is one
26
- small probe per layer (`layer_<L>/config.json`, `layer_<L>/model.pt`); an EFC or axial probe is one module
27
- that reads all its layers at once (`config.json`, `model.pt`). Every probe has `probe_metadata.json` in
28
- schema version 1: the model and revision, the architecture, the layers, the input window
29
- (`input_window`: `answer_last_token` for linear and MLP, `final_exchange` for EFC and axial) and the
30
- readout (`readout`: `layer_mean_probability`, `pooled_logit` or `last_token_logit`). A linear or MLP
31
- probe averages the sigmoids of all its layers. `model.pt` files are plain float32 state dicts.
32
- probe-inference 0.1.0 reads the same weights with the older metadata format, at commit
33
- `22a7a341078ba1722cfad73ef7e40bdc25aa74c2`.
34
 
35
  | Directory | Model (revision) | Layers | linear MiB | MLP MiB | EFC MiB | axial MiB |
36
  |---|---|---|---|---|---|---|
37
  | `qwen3.5-2b` | `Qwen/Qwen3.5-2B` (`15852e8c`) | 7, 10, 13, 16, 19, 22 | 0.1 | 12.0 | 6.1 | 26.1 |
38
  | `qwen3.5-9b` | `Qwen/Qwen3.5-9B` (`c2022362`) | 10, 13, 18, 21, 26, 29 | 0.1 | 24.0 | 12.1 | 28.1 |
39
- | `qwen3.6-27b` | `Qwen/Qwen3.6-27B` (`6a9e13bd`) | 19, 27, 35, 43, 51, 58 | 0.1 | 30.0 | 15.1 | 29.2 |
40
  | `qwen3.5-122b-a10b` | `Qwen/Qwen3.5-122B-A10B` (`dc4d3484`) | 14, 20, 26, 32, 38, 43 | 0.1 | 18.0 | 9.1 | 27.1 |
41
  | `qwen3.5-397b-a17b` | `Qwen/Qwen3.5-397B-A17B` (`84726181`) | 18, 25, 33, 40, 48, 54 | 0.1 | 24.0 | 12.1 | 28.1 |
42
  | `nemotron-3-nano-30b-a3b` | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` (`bf77c317`) | 16, 22, 29, 35, 42, 47 | 0.1 | 15.8 | 8.0 | 26.7 |
43
  | `nemotron-3-super-120b-a12b` | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16` (`2dc98e2a`) | 26, 37, 48, 59, 70, 79 | 0.1 | 24.0 | 12.1 | 28.1 |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
44
 
45
- Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and axial scores are logits.
 
 
 
 
 
 
 
 
 
 
 
 
46
 
47
  ## Activations the probes expect
48
 
49
  - Layer `k` is the output of decoder block `k`, i.e. Hugging Face `hidden_states[k + 1]`, in bfloat16; the
50
- probes run in float32.
51
- - The transcript ends with a final user turn and a prefilled assistant answer, closed by the end-of-turn
52
- token, rendered with the model's chat template with thinking disabled.
53
- - Linear and MLP read one token: the last answer token before end-of-turn. EFC and axial read every token
54
- from the start of the final user turn through end-of-turn. The window ends at the final end-of-turn
55
- token; anything after it, such as the newline `apply_chat_template` appends, is not read.
56
- - `qwen3.5-397b-a17b` activations were captured with vLLM at the same decoder-layer outputs; the other
57
- models' with Hugging Face forward hooks.
 
 
 
 
 
 
 
 
 
 
 
58
 
59
  ## Licences and attribution
60
 
61
  The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`).
62
 
63
- They are derived from the models below. Both upstream licences let us license derived works under our own
64
- terms provided we include their licence texts and keep their attribution notices, so both ship here
65
- unchanged (see `NOTICE`):
66
 
67
  | Model | Licence | Licence file |
68
  |---|---|---|
69
- | Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.6-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B | Apache-2.0 | `LICENSE-QWEN-APACHE-2.0.txt` |
70
  | nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | NVIDIA Nemotron Open Model License (v. December 15, 2025) | `LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt` |
 
 
71
 
72
  Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.
 
4
  - probes
5
  - activation-probes
6
  - interpretability
7
+ datasets:
8
+ - AlignmentResearch/fibs-v1
9
+ base_model:
10
+ - Qwen/Qwen3.5-2B
11
+ - Qwen/Qwen3.5-9B
12
+ - Qwen/Qwen3.5-27B
13
+ - Qwen/Qwen3.5-122B-A10B
14
+ - Qwen/Qwen3.5-397B-A17B
15
+ - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
16
+ - nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
17
+ - nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
18
+ - moonshotai/Kimi-K3
19
  ---
20
 
21
  # probe-inference weights
22
 
23
  Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial)
24
+ for nine open-weight models. Each probe reads a model's residual-stream activations at six decoder
25
+ layers and returns one score per transcript. Load them with the `probe-inference` package
26
+ ([AlignmentResearch/caught-in-the-act-probes](https://github.com/AlignmentResearch/caught-in-the-act-probes)):
27
 
28
  ```python
29
  from probe_inference import load_probe_from_hub
30
 
31
  probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package's pinned revision
32
+ score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))
 
33
  ```
34
 
35
+ These are the camera-ready probes, tagged `camera-ready`. The earlier probes (seven models, Qwen3.6-27B
36
+ instead of Qwen3.5-27B, no Nemotron-3 Ultra or Kimi K3) stay available at commit
37
+ `22a7a341078ba1722cfad73ef7e40bdc25aa74c2`, and in schema-1 metadata at commit
38
+ `0e3d386b1f1486d472e316a64966a7a94bcd5137`.
39
+
40
  ## Layout
41
 
42
+ `<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A
43
+ linear or MLP probe is one small probe per layer (`layer_<L>/config.json`, `layer_<L>/model.pt`); an EFC or
44
+ axial probe is one module that reads all its layers at once (`config.json`, `model.pt`). Every probe has
45
+ `probe_metadata.json`: the model and revision, the architecture, the layers, the read window
46
+ (`obfuscate_over`), the token aggregation (`eval_sequence_aggregator`) and, for linear and MLP, the layers
47
+ whose sigmoids are averaged (`layer_rule.used_layers`, all six). `model.pt` files are plain float32 state
48
+ dicts, with the input normaliser (`input_scale`, and `input_mean` for axial) the probe was trained with.
 
 
49
 
50
  | Directory | Model (revision) | Layers | linear MiB | MLP MiB | EFC MiB | axial MiB |
51
  |---|---|---|---|---|---|---|
52
  | `qwen3.5-2b` | `Qwen/Qwen3.5-2B` (`15852e8c`) | 7, 10, 13, 16, 19, 22 | 0.1 | 12.0 | 6.1 | 26.1 |
53
  | `qwen3.5-9b` | `Qwen/Qwen3.5-9B` (`c2022362`) | 10, 13, 18, 21, 26, 29 | 0.1 | 24.0 | 12.1 | 28.1 |
54
+ | `qwen3.5-27b` | `Qwen/Qwen3.5-27B` (`fc05daec`) | 19, 27, 35, 43, 51, 58 | 0.1 | 30.0 | 15.1 | 29.2 |
55
  | `qwen3.5-122b-a10b` | `Qwen/Qwen3.5-122B-A10B` (`dc4d3484`) | 14, 20, 26, 32, 38, 43 | 0.1 | 18.0 | 9.1 | 27.1 |
56
  | `qwen3.5-397b-a17b` | `Qwen/Qwen3.5-397B-A17B` (`84726181`) | 18, 25, 33, 40, 48, 54 | 0.1 | 24.0 | 12.1 | 28.1 |
57
  | `nemotron-3-nano-30b-a3b` | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` (`bf77c317`) | 16, 22, 29, 35, 42, 47 | 0.1 | 15.8 | 8.0 | 26.7 |
58
  | `nemotron-3-super-120b-a12b` | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16` (`2dc98e2a`) | 26, 37, 48, 59, 70, 79 | 0.1 | 24.0 | 12.1 | 28.1 |
59
+ | `nemotron-3-ultra-550b-a55b` | `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16` (`77df655d`) | 32, 45, 59, 72, 86, 97 | 0.2 | 48.0 | 24.1 | 32.2 |
60
+ | `kimi-k3` | `moonshotai/Kimi-K3` (`f831ab66`) | 28, 39, 51, 62, 74, 84 | 0.2 | 42.0 | 21.1 | 31.2 |
61
+
62
+ The layers are at depth fractions 0.3, 0.42, 0.55, 0.67, 0.8 and 0.9 of the model's decoder blocks,
63
+ `round(f * num_blocks)`. Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and
64
+ axial scores are logits.
65
+
66
+ ## Training
67
+
68
+ Every probe was trained on [FIBSv1](https://huggingface.co/datasets/AlignmentResearch/fibs-v1)
69
+ (`AlignmentResearch/fibs-v1`) at revision `65dccf12934620bfe01db665ea3d7d0e3c9f1355`, on the model's own
70
+ activations over FIBSv1's transcripts:
71
+
72
+ - 152,980 training rows and 5,000 validation rows, the same rows for every model; seed 0;
73
+ - linear and MLP: up to 6 epochs; EFC and axial: up to 4 epochs;
74
+ - early stopping: training stops after 4 checkpoints (one epoch) in a row without a lower validation
75
+ loss. The published probe is the checkpoint with the lowest validation loss;
76
+ - linear and MLP divide each layer's activations by `input_scale`, and axial subtracts `input_mean`
77
+ and then divides by `input_scale`. These are computed per layer from up to 512 training rows and are
78
+ stored in `model.pt`. EFC normalises each token by its own RMS and stores no statistic;
79
+ - each published probe reproduces the trainer's validation loss and AUROC within 1e-4.
80
 
81
+ Validation AUROC on the 5,000 validation rows (ES: stopped early):
82
+
83
+ | Directory | linear | MLP | EFC | axial |
84
+ |---|---|---|---|---|
85
+ | `qwen3.5-2b` | 0.926 | 0.968 | 0.987 | 0.991 ES |
86
+ | `qwen3.5-9b` | 0.965 | 0.986 | 0.996 | 0.999 |
87
+ | `qwen3.5-27b` | 0.972 | 0.988 | 0.998 | 0.999 |
88
+ | `qwen3.5-122b-a10b` | 0.938 | 0.972 | 0.998 ES | 0.999 |
89
+ | `qwen3.5-397b-a17b` | 0.948 | 0.977 | 0.999 | 0.999 |
90
+ | `nemotron-3-nano-30b-a3b` | 0.951 | 0.980 | 0.993 | 0.996 ES |
91
+ | `nemotron-3-super-120b-a12b` | 0.968 | 0.986 | 0.998 | 0.998 ES |
92
+ | `nemotron-3-ultra-550b-a55b` | 0.982 | 0.993 | 0.999 | 0.999 ES |
93
+ | `kimi-k3` | 0.948 | 0.959 | 0.999 | 0.999 ES |
94
 
95
  ## Activations the probes expect
96
 
97
  - Layer `k` is the output of decoder block `k`, i.e. Hugging Face `hidden_states[k + 1]`, in bfloat16; the
98
+ probes run in float32. All activations were captured with vLLM at the decoder-layer outputs.
99
+ - Kimi K3's decoder layers use attention residuals, so a layer has no single residual stream. Its layer
100
+ `k` is the attention-residual mixture that layer `k + 1` reads, computed with the model's own
101
+ `attn_res` op. Kimi K3 ran from its released MXFP4 checkpoint, not a bf16 one; its activations are
102
+ bfloat16.
103
+ - The transcript ends with a final user turn and an assistant answer, closed by the end-of-turn token,
104
+ rendered with the model's chat template with thinking disabled. Nemotron-3: the template's default,
105
+ which removes the reasoning of every assistant turn before the final user turn. Kimi K3: its template
106
+ cannot disable thinking, so the final assistant turn has an empty reasoning block before the answer.
107
+ - Linear and MLP read one token: the token before the final end-of-turn token. For Qwen and Nemotron-3
108
+ that is the answer's last token. For Kimi K3 it is the `<|sep|>` that closes `<|close|>message`, after
109
+ the answer's `<|close|>response<|sep|>`. EFC and axial read every token from the start of the final user
110
+ turn through the end-of-turn token.
111
+
112
+ ```
113
+ Qwen3.5: <|im_start|>user\n{question}<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\nNo.<|im_end|>
114
+ Nemotron-3: <|im_start|>user\n{question}<|im_end|>\n<|im_start|>assistant\n<think></think>No.<|im_end|>
115
+ Kimi K3: <|open|>message role="user"<|sep|>{question}<|close|>message<|sep|><|end_of_msg|><|open|>message role="assistant"<|sep|><|open|>think<|sep|><|close|>think<|sep|><|open|>response<|sep|>No.<|close|>response<|sep|><|close|>message<|sep|><|end_of_msg|>
116
+ ```
117
 
118
  ## Licences and attribution
119
 
120
  The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`).
121
 
122
+ They are derived from the models below. Their licence texts ship here unchanged, and their attribution
123
+ notices are kept in `NOTICE`:
 
124
 
125
  | Model | Licence | Licence file |
126
  |---|---|---|
127
+ | Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B | Apache-2.0 | `LICENSE-QWEN-APACHE-2.0.txt` |
128
  | nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | NVIDIA Nemotron Open Model License (v. December 15, 2025) | `LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt` |
129
+ | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | OpenMDW License Agreement, version 1.1 | `LICENSE-OPENMDW-1.1.txt` |
130
+ | moonshotai/Kimi-K3 | Kimi K3 License | `LICENSE-KIMI-K3.txt` |
131
 
132
  Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.
kimi-k3/axial/config.json ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "AxialProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "num_layers": 6,
6
+ "d_proj": 256,
7
+ "n_attn_heads": 4,
8
+ "n_blocks": 4,
9
+ "d_ff": 1024,
10
+ "dropout": 0.0,
11
+ "use_layer_embed": true,
12
+ "pool_mode": "cls",
13
+ "use_checkpoint": true,
14
+ "normalize_input": "centered_unit_norm",
15
+ "proj_adapter_rank": 0,
16
+ "input_adapter_rank": 0,
17
+ "sliding_window": null,
18
+ "layer_keys": [
19
+ "28",
20
+ "39",
21
+ "51",
22
+ "62",
23
+ "74",
24
+ "84"
25
+ ]
26
+ }
27
+ }
kimi-k3/axial/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e9f58be05ae2a458a158a6a137b2781d67086bdc51998023bf7aa415e15f9eec
3
+ size 32730360
kimi-k3/axial/probe_metadata.json ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "moonshotai/Kimi-K3",
3
+ "model_revision": "f831ab66814297da540d832a5235f8e904f29d06",
4
+ "architecture": "axial",
5
+ "layers": [
6
+ 28,
7
+ 39,
8
+ 51,
9
+ 62,
10
+ 74,
11
+ 84
12
+ ],
13
+ "eval_sequence_aggregator": "last",
14
+ "obfuscate_over": "last-user-and-assistant-generation"
15
+ }
kimi-k3/efc/config.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "EFCProbe",
3
+ "init_args": {
4
+ "num_layers": 6,
5
+ "d_model": 7168,
6
+ "feature_mode": "combined",
7
+ "center": true,
8
+ "d_hidden": 64,
9
+ "d_probe": 256,
10
+ "normalization": "rms",
11
+ "normalization_eps": 1e-06,
12
+ "shrinkage": "fixed",
13
+ "shrinkage_alpha": 0.1,
14
+ "spectral_transform": "eigh",
15
+ "newton_schulz_iterations": 10,
16
+ "jitter": 1e-05,
17
+ "feature_chunk_size": 4,
18
+ "normalize_input": "none",
19
+ "layer_keys": [
20
+ "28",
21
+ "39",
22
+ "51",
23
+ "62",
24
+ "74",
25
+ "84"
26
+ ]
27
+ }
28
+ }
kimi-k3/efc/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ac7a0b7ecfc54147598c274e95b03b94eff9b2145270b29e671431941af31d0b
3
+ size 22156448
kimi-k3/efc/probe_metadata.json ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "moonshotai/Kimi-K3",
3
+ "model_revision": "f831ab66814297da540d832a5235f8e904f29d06",
4
+ "architecture": "efc",
5
+ "layers": [
6
+ 28,
7
+ 39,
8
+ 51,
9
+ 62,
10
+ 74,
11
+ 84
12
+ ],
13
+ "eval_sequence_aggregator": "mean",
14
+ "obfuscate_over": "last-user-and-assistant-generation"
15
+ }
kimi-k3/linear/layer_28/config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "LinearProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "nhead": 1,
6
+ "normalize_input": "unit_norm"
7
+ }
8
+ }
kimi-k3/linear/layer_28/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:911ba849ea162e8c2e156d4d79d5b7b284cb9ba66f85229d0c75749ac99472c8
3
+ size 31421
kimi-k3/linear/layer_39/config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "LinearProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "nhead": 1,
6
+ "normalize_input": "unit_norm"
7
+ }
8
+ }
kimi-k3/linear/layer_39/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b2d9f1ea91389baeeef2a53d5031d3cc4dafdc1d09c628e44e41716dfb14b704
3
+ size 31421
kimi-k3/linear/layer_51/config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "LinearProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "nhead": 1,
6
+ "normalize_input": "unit_norm"
7
+ }
8
+ }
kimi-k3/linear/layer_51/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bdce0165c3fc3327d76598c0ab6f02417e9dec4e9fbf1d4390a2348dd63e182c
3
+ size 31421
kimi-k3/linear/layer_62/config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "LinearProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "nhead": 1,
6
+ "normalize_input": "unit_norm"
7
+ }
8
+ }
kimi-k3/linear/layer_62/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:efe2e891fb0d1ade54648a18e33abb17941825ce04b4160724ce223945d7b985
3
+ size 31421
kimi-k3/linear/layer_74/config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "LinearProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "nhead": 1,
6
+ "normalize_input": "unit_norm"
7
+ }
8
+ }
kimi-k3/linear/layer_74/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:78c6e9e498c8eb19a012e0afb2ea0b08c6240daedd3d1059c98b16f6d3f43e88
3
+ size 31421
kimi-k3/linear/layer_84/config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "LinearProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "nhead": 1,
6
+ "normalize_input": "unit_norm"
7
+ }
8
+ }
kimi-k3/linear/layer_84/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6c315787f11b2da67c63e6597516c5727bb395604ea20bb53a7af55f513d16b8
3
+ size 31421
kimi-k3/linear/probe_metadata.json ADDED
@@ -0,0 +1,25 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "moonshotai/Kimi-K3",
3
+ "model_revision": "f831ab66814297da540d832a5235f8e904f29d06",
4
+ "architecture": "linear",
5
+ "layers": [
6
+ 28,
7
+ 39,
8
+ 51,
9
+ 62,
10
+ 74,
11
+ 84
12
+ ],
13
+ "eval_sequence_aggregator": "mean",
14
+ "obfuscate_over": "second-last-token-generation",
15
+ "layer_rule": {
16
+ "used_layers": [
17
+ 28,
18
+ 39,
19
+ 51,
20
+ 62,
21
+ 74,
22
+ 84
23
+ ]
24
+ }
25
+ }
kimi-k3/mlp/layer_28/config.json ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "MLPProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "d_mlp": 256,
6
+ "nhead": 1,
7
+ "normalize_input": "unit_norm",
8
+ "activation": "relu",
9
+ "d_mlp2": null
10
+ }
11
+ }
kimi-k3/mlp/layer_28/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:42354c090cba70cccedcd68f538dc22a1fce8168724c582b21dbdb227db1ce83
3
+ size 7345201
kimi-k3/mlp/layer_39/config.json ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "MLPProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "d_mlp": 256,
6
+ "nhead": 1,
7
+ "normalize_input": "unit_norm",
8
+ "activation": "relu",
9
+ "d_mlp2": null
10
+ }
11
+ }
kimi-k3/mlp/layer_39/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b6769d33b59fce353f360588b6e5e8c54a6dd4f4624948830e852760ff8a01ee
3
+ size 7345201
kimi-k3/mlp/layer_51/config.json ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "MLPProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "d_mlp": 256,
6
+ "nhead": 1,
7
+ "normalize_input": "unit_norm",
8
+ "activation": "relu",
9
+ "d_mlp2": null
10
+ }
11
+ }
kimi-k3/mlp/layer_51/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c6604a356b9e5ac866f49dc24213cc08b332bd771321f31083f2dfdf33b25432
3
+ size 7345201
kimi-k3/mlp/layer_62/config.json ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "MLPProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "d_mlp": 256,
6
+ "nhead": 1,
7
+ "normalize_input": "unit_norm",
8
+ "activation": "relu",
9
+ "d_mlp2": null
10
+ }
11
+ }
kimi-k3/mlp/layer_62/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7d4dda19ec1c77b9fda81bb7fa6697cbf484f1a76849b7080657fa882b0259ee
3
+ size 7345201
kimi-k3/mlp/layer_74/config.json ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "MLPProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "d_mlp": 256,
6
+ "nhead": 1,
7
+ "normalize_input": "unit_norm",
8
+ "activation": "relu",
9
+ "d_mlp2": null
10
+ }
11
+ }
kimi-k3/mlp/layer_74/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:28bd24aaa1f446b193ab82b1ac9d1df16957ebdb10dc1c3cc6013c68d2cf3edd
3
+ size 7345201
kimi-k3/mlp/layer_84/config.json ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "class_name": "MLPProbe",
3
+ "init_args": {
4
+ "d_model": 7168,
5
+ "d_mlp": 256,
6
+ "nhead": 1,
7
+ "normalize_input": "unit_norm",
8
+ "activation": "relu",
9
+ "d_mlp2": null
10
+ }
11
+ }
kimi-k3/mlp/layer_84/model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ce505c1fdb735b7d1fe93113768cc52e0259a4940acf77489d35a8149e26e1d8
3
+ size 7345201
kimi-k3/mlp/probe_metadata.json ADDED
@@ -0,0 +1,25 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "moonshotai/Kimi-K3",
3
+ "model_revision": "f831ab66814297da540d832a5235f8e904f29d06",
4
+ "architecture": "mlp",
5
+ "layers": [
6
+ 28,
7
+ 39,
8
+ 51,
9
+ 62,
10
+ 74,
11
+ 84
12
+ ],
13
+ "eval_sequence_aggregator": "mean",
14
+ "obfuscate_over": "second-last-token-generation",
15
+ "layer_rule": {
16
+ "used_layers": [
17
+ 28,
18
+ 39,
19
+ 51,
20
+ 62,
21
+ 74,
22
+ 84
23
+ ]
24
+ }
25
+ }
nemotron-3-nano-30b-a3b/axial/model.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:110882082b46efef53d87f04530d910e1cb30903324d52d8d9ac1881f0ee5a93
3
  size 28035320
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:345567d256088641608f0babda239691ebe8f3ad68418e4e386afe7ab20b8441
3
  size 28035320
nemotron-3-nano-30b-a3b/axial/probe_metadata.json CHANGED
@@ -1,5 +1,4 @@
1
  {
2
- "schema_version": 1,
3
  "model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
4
  "model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
5
  "architecture": "axial",
@@ -11,6 +10,6 @@
11
  42,
12
  47
13
  ],
14
- "input_window": "final_exchange",
15
- "readout": "last_token_logit"
16
  }
 
1
  {
 
2
  "model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
3
  "model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
4
  "architecture": "axial",
 
10
  42,
11
  47
12
  ],
13
+ "eval_sequence_aggregator": "last",
14
+ "obfuscate_over": "last-user-and-assistant-generation"
15
  }
nemotron-3-nano-30b-a3b/efc/model.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a99eaef4206efd9881653e45ddc07c4eda443609cff9d0767f2a56e32f97cee7
3
  size 8393888
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:42f946be32dcdead9c673be7f397bb18b221cd05cbd2f37fb7d9dc36e7e8f392
3
  size 8393888
nemotron-3-nano-30b-a3b/efc/probe_metadata.json CHANGED
@@ -1,5 +1,4 @@
1
  {
2
- "schema_version": 1,
3
  "model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
4
  "model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
5
  "architecture": "efc",
@@ -11,6 +10,6 @@
11
  42,
12
  47
13
  ],
14
- "input_window": "final_exchange",
15
- "readout": "pooled_logit"
16
  }
 
1
  {
 
2
  "model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
3
  "model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
4
  "architecture": "efc",
 
10
  42,
11
  47
12
  ],
13
+ "eval_sequence_aggregator": "mean",
14
+ "obfuscate_over": "last-user-and-assistant-generation"
15
  }
nemotron-3-nano-30b-a3b/linear/layer_16/model.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:eb6d069a758fc3abf09afcbf806bd0270e4d14b72edc0a6bfe05909b806cc6a0
3
  size 13501
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8ee8243e4e836f290d471616b9cc6a7391c5d908e3d18ae3222b1ac7b816c763
3
  size 13501
nemotron-3-nano-30b-a3b/linear/layer_22/model.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:7c4ed443b764006b76f99f3c91856e9d642c9338908250d5dc44b2ee8c0dc73a
3
  size 13501
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d0f1deda31003d2a6e7b1af3231a1d714cf2df1bdf27f697359db30d31b3e988
3
  size 13501
nemotron-3-nano-30b-a3b/linear/layer_29/model.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:13793b0c18d0770bc6ef0bcf79e9e8a43fc477e6b48161054bf2dc25a27c2eaf
3
  size 13501
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:61c9fb00b39d038fca5e854d82d1dd1f0f7b8598e56b17f4b5f4578861aff6fd
3
  size 13501
nemotron-3-nano-30b-a3b/linear/layer_35/model.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:25e5a52a5445a1667488cd02dd868683000154c53994c081ecbb5010e191108a
3
  size 13501
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7c92f889fae2e52832f25b96b25133c07e1a5c84e5d279c8cb047e732b933ee4
3
  size 13501
nemotron-3-nano-30b-a3b/linear/layer_42/model.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:cd4106fef9b007b81f66093342a5588e71f7101ee8b976a26601fc82be7301ea
3
  size 13501
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ffc79f4fccbae06ada63dd0925f53ab33963a8a89e5aaf969d7a1764bdf3fde4
3
  size 13501
nemotron-3-nano-30b-a3b/linear/layer_47/model.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:71fdb470fb96a9d1b242d7acd3efc8b82eb80bd00e14e20d6229dc8db50f89ef
3
  size 13501
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8c6543a8a2428395d9bb59fea8d4ec56d58bfc47dafd4783970af4d4d72de7fe
3
  size 13501
nemotron-3-nano-30b-a3b/linear/probe_metadata.json CHANGED
@@ -1,5 +1,4 @@
1
  {
2
- "schema_version": 1,
3
  "model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
4
  "model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
5
  "architecture": "linear",
@@ -11,6 +10,16 @@
11
  42,
12
  47
13
  ],
14
- "input_window": "answer_last_token",
15
- "readout": "layer_mean_probability"
 
 
 
 
 
 
 
 
 
 
16
  }
 
1
  {
 
2
  "model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
3
  "model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
4
  "architecture": "linear",
 
10
  42,
11
  47
12
  ],
13
+ "eval_sequence_aggregator": "mean",
14
+ "obfuscate_over": "second-last-token-generation",
15
+ "layer_rule": {
16
+ "used_layers": [
17
+ 16,
18
+ 22,
19
+ 29,
20
+ 35,
21
+ 42,
22
+ 47
23
+ ]
24
+ }
25
  }
nemotron-3-nano-30b-a3b/mlp/layer_16/model.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:964f9352f96a02340f84efee4ebbefea63a078f997e3adfa47b5e19e8c05c0df
3
  size 2757681
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:86bc24e945d5397d1fa96bdf09976dc3c67d04afbe0d395f138c82fe9ed4bf85
3
  size 2757681
nemotron-3-nano-30b-a3b/mlp/layer_22/model.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a209e6499d9bab5d744912f261887ab1e653275d40d764cf45d551304058dd5b
3
  size 2757681
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e44cbd9a75b74730c316b09ce3c1d807f4b735efa8d13b19c094291472cdd6c8
3
  size 2757681
nemotron-3-nano-30b-a3b/mlp/layer_29/model.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:6a5572ef46021ce52870781f5ad56b8b04a780126a07daad269e8cec77abd989
3
  size 2757681
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bcc9eae57322e9c86f7f256425add809beb28eba92c52da4652e583417d8ef88
3
  size 2757681