Camera-ready probes: 9 models x {linear, mlp, efc, axial} trained on FIBSv1 @ 65dccf12
Browse filesReplaces every probe directory with the camera-ready E5 banks (deception PR #1078): Qwen3.5 2B, 9B, 27B,
122B-A10B and 397B-A17B; Nemotron-3 Nano-30B-A3B, Super-120B-A12B and Ultra-550B-A55B; Kimi K3. Trained on
AlignmentResearch/fibs-v1 at 65dccf12, seed 0, linear/MLP up to 6 epochs, EFC/axial up to 4, early stopping,
input_scale included. Adds qwen3.5-27b, nemotron-3-ultra-550b-a55b and kimi-k3; removes qwen3.6-27b (still at
22a7a341 and 0e3d386b). probe_metadata.json is in the format probe-inference 0.1.0 reads (layers,
eval_sequence_aggregator, obfuscate_over, layer_rule). Card, NOTICE and the OpenMDW-1.1 and Kimi K3 licence
texts updated.
This view is limited to 50 files because it contains too many changes. See raw diff
- LICENSE-KIMI-K3.txt +52 -0
- LICENSE-OPENMDW-1.1.txt +49 -0
- NOTICE +8 -1
- README.md +87 -27
- kimi-k3/axial/config.json +27 -0
- kimi-k3/axial/model.pt +3 -0
- kimi-k3/axial/probe_metadata.json +15 -0
- kimi-k3/efc/config.json +28 -0
- kimi-k3/efc/model.pt +3 -0
- kimi-k3/efc/probe_metadata.json +15 -0
- kimi-k3/linear/layer_28/config.json +8 -0
- kimi-k3/linear/layer_28/model.pt +3 -0
- kimi-k3/linear/layer_39/config.json +8 -0
- kimi-k3/linear/layer_39/model.pt +3 -0
- kimi-k3/linear/layer_51/config.json +8 -0
- kimi-k3/linear/layer_51/model.pt +3 -0
- kimi-k3/linear/layer_62/config.json +8 -0
- kimi-k3/linear/layer_62/model.pt +3 -0
- kimi-k3/linear/layer_74/config.json +8 -0
- kimi-k3/linear/layer_74/model.pt +3 -0
- kimi-k3/linear/layer_84/config.json +8 -0
- kimi-k3/linear/layer_84/model.pt +3 -0
- kimi-k3/linear/probe_metadata.json +25 -0
- kimi-k3/mlp/layer_28/config.json +11 -0
- kimi-k3/mlp/layer_28/model.pt +3 -0
- kimi-k3/mlp/layer_39/config.json +11 -0
- kimi-k3/mlp/layer_39/model.pt +3 -0
- kimi-k3/mlp/layer_51/config.json +11 -0
- kimi-k3/mlp/layer_51/model.pt +3 -0
- kimi-k3/mlp/layer_62/config.json +11 -0
- kimi-k3/mlp/layer_62/model.pt +3 -0
- kimi-k3/mlp/layer_74/config.json +11 -0
- kimi-k3/mlp/layer_74/model.pt +3 -0
- kimi-k3/mlp/layer_84/config.json +11 -0
- kimi-k3/mlp/layer_84/model.pt +3 -0
- kimi-k3/mlp/probe_metadata.json +25 -0
- nemotron-3-nano-30b-a3b/axial/model.pt +1 -1
- nemotron-3-nano-30b-a3b/axial/probe_metadata.json +2 -3
- nemotron-3-nano-30b-a3b/efc/model.pt +1 -1
- nemotron-3-nano-30b-a3b/efc/probe_metadata.json +2 -3
- nemotron-3-nano-30b-a3b/linear/layer_16/model.pt +1 -1
- nemotron-3-nano-30b-a3b/linear/layer_22/model.pt +1 -1
- nemotron-3-nano-30b-a3b/linear/layer_29/model.pt +1 -1
- nemotron-3-nano-30b-a3b/linear/layer_35/model.pt +1 -1
- nemotron-3-nano-30b-a3b/linear/layer_42/model.pt +1 -1
- nemotron-3-nano-30b-a3b/linear/layer_47/model.pt +1 -1
- nemotron-3-nano-30b-a3b/linear/probe_metadata.json +12 -3
- nemotron-3-nano-30b-a3b/mlp/layer_16/model.pt +1 -1
- nemotron-3-nano-30b-a3b/mlp/layer_22/model.pt +1 -1
- nemotron-3-nano-30b-a3b/mlp/layer_29/model.pt +1 -1
LICENSE-KIMI-K3.txt
ADDED
|
@@ -0,0 +1,52 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Kimi K3 License
|
| 2 |
+
|
| 3 |
+
Copyright (c) 2026 Moonshot AI
|
| 4 |
+
|
| 5 |
+
Permission is hereby granted, free of charge, to any person (the "Licensee")
|
| 6 |
+
obtaining a copy of this software — including the model weights, parameters,
|
| 7 |
+
configuration files, inference and training code, and associated documentation
|
| 8 |
+
(collectively, the "Software") — to deal in the Software without restriction.
|
| 9 |
+
This includes, without limitation, the rights to use, copy, modify, merge,
|
| 10 |
+
publish, distribute, sublicense, and/or sell copies of the Software; to run,
|
| 11 |
+
deploy, fine-tune, or otherwise modify the Software and create derivative works
|
| 12 |
+
from it; and to permit persons to whom the Software is furnished to do so, in
|
| 13 |
+
each case subject to the following conditions:
|
| 14 |
+
|
| 15 |
+
1. The above copyright notice and this permission notice shall be included in
|
| 16 |
+
all copies or substantial portions of the Software. Licensee's use of the
|
| 17 |
+
Software must comply with applicable laws and regulations.
|
| 18 |
+
|
| 19 |
+
2. "Model as a Service" means giving a third party access to language model
|
| 20 |
+
inference or fine-tuning (e.g., via API) in a manner that allows such third
|
| 21 |
+
party to exercise meaningful control over the inputs, parameters, or training
|
| 22 |
+
data. This does not include (a) end-user products with model capabilities solely
|
| 23 |
+
embedded within specific features or harnesses, or (b) mere relaying of requests
|
| 24 |
+
to models hosted by others.
|
| 25 |
+
|
| 26 |
+
If the Licensee or any of its affiliates operates a Model as a Service business,
|
| 27 |
+
and the aggregate revenue of the Licensee and its affiliates exceeds 20 million
|
| 28 |
+
US dollars (or the equivalent in other currencies) in total over any consecutive
|
| 29 |
+
12 months, the Licensee must enter into a separate agreement with Moonshot AI
|
| 30 |
+
before using the Software or its derivative works for any commercial purpose.
|
| 31 |
+
|
| 32 |
+
3. If the Software (or any derivative works thereof) is used for any of the
|
| 33 |
+
Licensee's commercial products or services that have more than 100 million
|
| 34 |
+
monthly active users, or more than 20 million US dollars (or equivalent in other
|
| 35 |
+
currencies) in monthly revenue, "Kimi K3" must be prominently displayed on the
|
| 36 |
+
user interface of such product or service.
|
| 37 |
+
|
| 38 |
+
4. The requirements set forth in Sections 2 and 3 do not apply to: (a) internal
|
| 39 |
+
use of the Software, defined as any use that does not make the Software, its
|
| 40 |
+
outputs, or its underlying capabilities available to third parties; or (b) any
|
| 41 |
+
use of the Software accessed through Moonshot AI's official products or
|
| 42 |
+
certified inference partners.
|
| 43 |
+
|
| 44 |
+
5. THE SOFTWARE AND ANY OUTPUT AND RESULTS THEREFROM ARE PROVIDED ON AN “AS IS”
|
| 45 |
+
BASIS, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT
|
| 46 |
+
LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE
|
| 47 |
+
AND NONINFRINGEMENT. IN NO EVENT SHALL MOONSHOT AI OR ITS AFFILIATES OR
|
| 48 |
+
COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER
|
| 49 |
+
IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN
|
| 50 |
+
CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
|
| 51 |
+
|
| 52 |
+
For any questions regarding this license, please contact <license@moonshot.ai>.
|
LICENSE-OPENMDW-1.1.txt
ADDED
|
@@ -0,0 +1,49 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
OpenMDW License Agreement, version 1.1 (OpenMDW-1.1)
|
| 2 |
+
|
| 3 |
+
By exercising rights granted to you under this agreement, you accept and agree
|
| 4 |
+
to its terms.
|
| 5 |
+
|
| 6 |
+
As used in this agreement, "Model Materials" means the materials provided to
|
| 7 |
+
you under this agreement, consisting of: (1) one or more machine learning
|
| 8 |
+
models (including architecture and parameters); and (2) all related artifacts
|
| 9 |
+
(including associated data, documentation and software) that are provided to
|
| 10 |
+
you hereunder.
|
| 11 |
+
|
| 12 |
+
Subject to your compliance with this agreement, permission is hereby granted,
|
| 13 |
+
free of charge, to deal in the Model Materials without restriction, including
|
| 14 |
+
under all copyright, patent, database, and trade secret rights included or
|
| 15 |
+
embodied therein.
|
| 16 |
+
|
| 17 |
+
If you distribute any portion of the Model Materials, you shall retain in your
|
| 18 |
+
distribution (1) a copy of this agreement, and (2) all copyright notices and
|
| 19 |
+
other notices of origin included in the Model Materials that are applicable to
|
| 20 |
+
your distribution.
|
| 21 |
+
|
| 22 |
+
If you file, maintain, or voluntarily participate in a lawsuit against any
|
| 23 |
+
person or entity asserting that the Model Materials directly or indirectly
|
| 24 |
+
infringe any patent or copyright, then all rights and grants made to you
|
| 25 |
+
hereunder are terminated, unless that lawsuit was in response to a
|
| 26 |
+
corresponding lawsuit first brought against you.
|
| 27 |
+
|
| 28 |
+
This agreement does not impose any restrictions or obligations with respect to
|
| 29 |
+
any use, modification, or sharing of any outputs generated by using the Model
|
| 30 |
+
Materials.
|
| 31 |
+
|
| 32 |
+
THE MODEL MATERIALS ARE PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS
|
| 33 |
+
OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
| 34 |
+
FITNESS FOR A PARTICULAR PURPOSE, TITLE, NONINFRINGEMENT, ACCURACY, OR THE
|
| 35 |
+
ABSENCE OF LATENT OR OTHER DEFECTS OR ERRORS, WHETHER OR NOT DISCOVERABLE, ALL
|
| 36 |
+
TO THE GREATEST EXTENT PERMISSIBLE UNDER APPLICABLE LAW.
|
| 37 |
+
|
| 38 |
+
YOU ARE SOLELY RESPONSIBLE FOR (1) CLEARING RIGHTS OF OTHER PERSONS THAT MAY
|
| 39 |
+
APPLY TO THE MODEL MATERIALS OR ANY USE THEREOF, INCLUDING WITHOUT LIMITATION
|
| 40 |
+
ANY PERSON'S COPYRIGHTS OR OTHER RIGHTS INCLUDED OR EMBODIED IN THE MODEL
|
| 41 |
+
MATERIALS; (2) OBTAINING ANY NECESSARY CONSENTS, PERMISSIONS OR OTHER RIGHTS
|
| 42 |
+
REQUIRED FOR ANY USE OF THE MODEL MATERIALS; OR (3) PERFORMING ANY DUE
|
| 43 |
+
DILIGENCE OR UNDERTAKING ANY OTHER INVESTIGATIONS INTO THE MODEL MATERIALS OR
|
| 44 |
+
ANYTHING INCORPORATED OR EMBODIED THEREIN.
|
| 45 |
+
|
| 46 |
+
IN NO EVENT SHALL THE PROVIDERS OF THE MODEL MATERIALS BE LIABLE FOR ANY CLAIM,
|
| 47 |
+
DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR
|
| 48 |
+
OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE MODEL MATERIALS, THE
|
| 49 |
+
USE THEREOF OR OTHER DEALINGS THEREIN.
|
NOTICE
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
This repository contains activations and/or probe parameters derived from the following models.
|
| 2 |
|
| 3 |
-
Qwen3.5
|
| 4 |
Qwen/Qwen3.5-397B-A17B): Copyright Alibaba Cloud (Qwen team). Licensed under the Apache License,
|
| 5 |
Version 2.0. A copy of the license is in LICENSE-QWEN-APACHE-2.0.txt.
|
| 6 |
|
|
@@ -8,3 +8,10 @@ NVIDIA Nemotron-3 models (nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16,
|
|
| 8 |
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16):
|
| 9 |
Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.
|
| 10 |
A copy of the NVIDIA Nemotron Open Model License is in LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
This repository contains activations and/or probe parameters derived from the following models.
|
| 2 |
|
| 3 |
+
Qwen3.5 models (Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-27B, Qwen/Qwen3.5-122B-A10B,
|
| 4 |
Qwen/Qwen3.5-397B-A17B): Copyright Alibaba Cloud (Qwen team). Licensed under the Apache License,
|
| 5 |
Version 2.0. A copy of the license is in LICENSE-QWEN-APACHE-2.0.txt.
|
| 6 |
|
|
|
|
| 8 |
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16):
|
| 9 |
Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.
|
| 10 |
A copy of the NVIDIA Nemotron Open Model License is in LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt.
|
| 11 |
+
|
| 12 |
+
NVIDIA Nemotron-3 Ultra (nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16):
|
| 13 |
+
Governed by the OpenMDW License Agreement, version 1.1 (OpenMDW-1.1).
|
| 14 |
+
A copy of the license is in LICENSE-OPENMDW-1.1.txt.
|
| 15 |
+
|
| 16 |
+
Kimi K3 (moonshotai/Kimi-K3): Copyright (c) 2026 Moonshot AI. Licensed under the Kimi K3 License.
|
| 17 |
+
A copy of the license, including its copyright and permission notice, is in LICENSE-KIMI-K3.txt.
|
README.md
CHANGED
|
@@ -4,69 +4,129 @@ tags:
|
|
| 4 |
- probes
|
| 5 |
- activation-probes
|
| 6 |
- interpretability
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 7 |
---
|
| 8 |
|
| 9 |
# probe-inference weights
|
| 10 |
|
| 11 |
Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial)
|
| 12 |
-
for
|
| 13 |
-
layers and returns one score per transcript. Load them with the `probe-inference` package
|
|
|
|
| 14 |
|
| 15 |
```python
|
| 16 |
from probe_inference import load_probe_from_hub
|
| 17 |
|
| 18 |
probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package's pinned revision
|
| 19 |
-
|
| 20 |
-
score = probe.score(acts[:, probe.find_window(input_ids, tokenizer)])
|
| 21 |
```
|
| 22 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 23 |
## Layout
|
| 24 |
|
| 25 |
-
`<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A
|
| 26 |
-
small probe per layer (`layer_<L>/config.json`, `layer_<L>/model.pt`); an EFC or
|
| 27 |
-
that reads all its layers at once (`config.json`, `model.pt`). Every probe has
|
| 28 |
-
|
| 29 |
-
(`
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
probe-inference 0.1.0 reads the same weights with the older metadata format, at commit
|
| 33 |
-
`22a7a341078ba1722cfad73ef7e40bdc25aa74c2`.
|
| 34 |
|
| 35 |
| Directory | Model (revision) | Layers | linear MiB | MLP MiB | EFC MiB | axial MiB |
|
| 36 |
|---|---|---|---|---|---|---|
|
| 37 |
| `qwen3.5-2b` | `Qwen/Qwen3.5-2B` (`15852e8c`) | 7, 10, 13, 16, 19, 22 | 0.1 | 12.0 | 6.1 | 26.1 |
|
| 38 |
| `qwen3.5-9b` | `Qwen/Qwen3.5-9B` (`c2022362`) | 10, 13, 18, 21, 26, 29 | 0.1 | 24.0 | 12.1 | 28.1 |
|
| 39 |
-
| `qwen3.
|
| 40 |
| `qwen3.5-122b-a10b` | `Qwen/Qwen3.5-122B-A10B` (`dc4d3484`) | 14, 20, 26, 32, 38, 43 | 0.1 | 18.0 | 9.1 | 27.1 |
|
| 41 |
| `qwen3.5-397b-a17b` | `Qwen/Qwen3.5-397B-A17B` (`84726181`) | 18, 25, 33, 40, 48, 54 | 0.1 | 24.0 | 12.1 | 28.1 |
|
| 42 |
| `nemotron-3-nano-30b-a3b` | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` (`bf77c317`) | 16, 22, 29, 35, 42, 47 | 0.1 | 15.8 | 8.0 | 26.7 |
|
| 43 |
| `nemotron-3-super-120b-a12b` | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16` (`2dc98e2a`) | 26, 37, 48, 59, 70, 79 | 0.1 | 24.0 | 12.1 | 28.1 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 44 |
|
| 45 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 46 |
|
| 47 |
## Activations the probes expect
|
| 48 |
|
| 49 |
- Layer `k` is the output of decoder block `k`, i.e. Hugging Face `hidden_states[k + 1]`, in bfloat16; the
|
| 50 |
-
probes run in float32.
|
| 51 |
-
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 58 |
|
| 59 |
## Licences and attribution
|
| 60 |
|
| 61 |
The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`).
|
| 62 |
|
| 63 |
-
They are derived from the models below.
|
| 64 |
-
|
| 65 |
-
unchanged (see `NOTICE`):
|
| 66 |
|
| 67 |
| Model | Licence | Licence file |
|
| 68 |
|---|---|---|
|
| 69 |
-
| Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.
|
| 70 |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | NVIDIA Nemotron Open Model License (v. December 15, 2025) | `LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt` |
|
|
|
|
|
|
|
| 71 |
|
| 72 |
Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.
|
|
|
|
| 4 |
- probes
|
| 5 |
- activation-probes
|
| 6 |
- interpretability
|
| 7 |
+
datasets:
|
| 8 |
+
- AlignmentResearch/fibs-v1
|
| 9 |
+
base_model:
|
| 10 |
+
- Qwen/Qwen3.5-2B
|
| 11 |
+
- Qwen/Qwen3.5-9B
|
| 12 |
+
- Qwen/Qwen3.5-27B
|
| 13 |
+
- Qwen/Qwen3.5-122B-A10B
|
| 14 |
+
- Qwen/Qwen3.5-397B-A17B
|
| 15 |
+
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
|
| 16 |
+
- nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
|
| 17 |
+
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
|
| 18 |
+
- moonshotai/Kimi-K3
|
| 19 |
---
|
| 20 |
|
| 21 |
# probe-inference weights
|
| 22 |
|
| 23 |
Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial)
|
| 24 |
+
for nine open-weight models. Each probe reads a model's residual-stream activations at six decoder
|
| 25 |
+
layers and returns one score per transcript. Load them with the `probe-inference` package
|
| 26 |
+
([AlignmentResearch/caught-in-the-act-probes](https://github.com/AlignmentResearch/caught-in-the-act-probes)):
|
| 27 |
|
| 28 |
```python
|
| 29 |
from probe_inference import load_probe_from_hub
|
| 30 |
|
| 31 |
probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package's pinned revision
|
| 32 |
+
score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))
|
|
|
|
| 33 |
```
|
| 34 |
|
| 35 |
+
These are the camera-ready probes, tagged `camera-ready`. The earlier probes (seven models, Qwen3.6-27B
|
| 36 |
+
instead of Qwen3.5-27B, no Nemotron-3 Ultra or Kimi K3) stay available at commit
|
| 37 |
+
`22a7a341078ba1722cfad73ef7e40bdc25aa74c2`, and in schema-1 metadata at commit
|
| 38 |
+
`0e3d386b1f1486d472e316a64966a7a94bcd5137`.
|
| 39 |
+
|
| 40 |
## Layout
|
| 41 |
|
| 42 |
+
`<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A
|
| 43 |
+
linear or MLP probe is one small probe per layer (`layer_<L>/config.json`, `layer_<L>/model.pt`); an EFC or
|
| 44 |
+
axial probe is one module that reads all its layers at once (`config.json`, `model.pt`). Every probe has
|
| 45 |
+
`probe_metadata.json`: the model and revision, the architecture, the layers, the read window
|
| 46 |
+
(`obfuscate_over`), the token aggregation (`eval_sequence_aggregator`) and, for linear and MLP, the layers
|
| 47 |
+
whose sigmoids are averaged (`layer_rule.used_layers`, all six). `model.pt` files are plain float32 state
|
| 48 |
+
dicts, with the input normaliser (`input_scale`, and `input_mean` for axial) the probe was trained with.
|
|
|
|
|
|
|
| 49 |
|
| 50 |
| Directory | Model (revision) | Layers | linear MiB | MLP MiB | EFC MiB | axial MiB |
|
| 51 |
|---|---|---|---|---|---|---|
|
| 52 |
| `qwen3.5-2b` | `Qwen/Qwen3.5-2B` (`15852e8c`) | 7, 10, 13, 16, 19, 22 | 0.1 | 12.0 | 6.1 | 26.1 |
|
| 53 |
| `qwen3.5-9b` | `Qwen/Qwen3.5-9B` (`c2022362`) | 10, 13, 18, 21, 26, 29 | 0.1 | 24.0 | 12.1 | 28.1 |
|
| 54 |
+
| `qwen3.5-27b` | `Qwen/Qwen3.5-27B` (`fc05daec`) | 19, 27, 35, 43, 51, 58 | 0.1 | 30.0 | 15.1 | 29.2 |
|
| 55 |
| `qwen3.5-122b-a10b` | `Qwen/Qwen3.5-122B-A10B` (`dc4d3484`) | 14, 20, 26, 32, 38, 43 | 0.1 | 18.0 | 9.1 | 27.1 |
|
| 56 |
| `qwen3.5-397b-a17b` | `Qwen/Qwen3.5-397B-A17B` (`84726181`) | 18, 25, 33, 40, 48, 54 | 0.1 | 24.0 | 12.1 | 28.1 |
|
| 57 |
| `nemotron-3-nano-30b-a3b` | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` (`bf77c317`) | 16, 22, 29, 35, 42, 47 | 0.1 | 15.8 | 8.0 | 26.7 |
|
| 58 |
| `nemotron-3-super-120b-a12b` | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16` (`2dc98e2a`) | 26, 37, 48, 59, 70, 79 | 0.1 | 24.0 | 12.1 | 28.1 |
|
| 59 |
+
| `nemotron-3-ultra-550b-a55b` | `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16` (`77df655d`) | 32, 45, 59, 72, 86, 97 | 0.2 | 48.0 | 24.1 | 32.2 |
|
| 60 |
+
| `kimi-k3` | `moonshotai/Kimi-K3` (`f831ab66`) | 28, 39, 51, 62, 74, 84 | 0.2 | 42.0 | 21.1 | 31.2 |
|
| 61 |
+
|
| 62 |
+
The layers are at depth fractions 0.3, 0.42, 0.55, 0.67, 0.8 and 0.9 of the model's decoder blocks,
|
| 63 |
+
`round(f * num_blocks)`. Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and
|
| 64 |
+
axial scores are logits.
|
| 65 |
+
|
| 66 |
+
## Training
|
| 67 |
+
|
| 68 |
+
Every probe was trained on [FIBSv1](https://huggingface.co/datasets/AlignmentResearch/fibs-v1)
|
| 69 |
+
(`AlignmentResearch/fibs-v1`) at revision `65dccf12934620bfe01db665ea3d7d0e3c9f1355`, on the model's own
|
| 70 |
+
activations over FIBSv1's transcripts:
|
| 71 |
+
|
| 72 |
+
- 152,980 training rows and 5,000 validation rows, the same rows for every model; seed 0;
|
| 73 |
+
- linear and MLP: up to 6 epochs; EFC and axial: up to 4 epochs;
|
| 74 |
+
- early stopping: training stops after 4 checkpoints (one epoch) in a row without a lower validation
|
| 75 |
+
loss. The published probe is the checkpoint with the lowest validation loss;
|
| 76 |
+
- linear and MLP divide each layer's activations by `input_scale`, and axial subtracts `input_mean`
|
| 77 |
+
and then divides by `input_scale`. These are computed per layer from up to 512 training rows and are
|
| 78 |
+
stored in `model.pt`. EFC normalises each token by its own RMS and stores no statistic;
|
| 79 |
+
- each published probe reproduces the trainer's validation loss and AUROC within 1e-4.
|
| 80 |
|
| 81 |
+
Validation AUROC on the 5,000 validation rows (ES: stopped early):
|
| 82 |
+
|
| 83 |
+
| Directory | linear | MLP | EFC | axial |
|
| 84 |
+
|---|---|---|---|---|
|
| 85 |
+
| `qwen3.5-2b` | 0.926 | 0.968 | 0.987 | 0.991 ES |
|
| 86 |
+
| `qwen3.5-9b` | 0.965 | 0.986 | 0.996 | 0.999 |
|
| 87 |
+
| `qwen3.5-27b` | 0.972 | 0.988 | 0.998 | 0.999 |
|
| 88 |
+
| `qwen3.5-122b-a10b` | 0.938 | 0.972 | 0.998 ES | 0.999 |
|
| 89 |
+
| `qwen3.5-397b-a17b` | 0.948 | 0.977 | 0.999 | 0.999 |
|
| 90 |
+
| `nemotron-3-nano-30b-a3b` | 0.951 | 0.980 | 0.993 | 0.996 ES |
|
| 91 |
+
| `nemotron-3-super-120b-a12b` | 0.968 | 0.986 | 0.998 | 0.998 ES |
|
| 92 |
+
| `nemotron-3-ultra-550b-a55b` | 0.982 | 0.993 | 0.999 | 0.999 ES |
|
| 93 |
+
| `kimi-k3` | 0.948 | 0.959 | 0.999 | 0.999 ES |
|
| 94 |
|
| 95 |
## Activations the probes expect
|
| 96 |
|
| 97 |
- Layer `k` is the output of decoder block `k`, i.e. Hugging Face `hidden_states[k + 1]`, in bfloat16; the
|
| 98 |
+
probes run in float32. All activations were captured with vLLM at the decoder-layer outputs.
|
| 99 |
+
- Kimi K3's decoder layers use attention residuals, so a layer has no single residual stream. Its layer
|
| 100 |
+
`k` is the attention-residual mixture that layer `k + 1` reads, computed with the model's own
|
| 101 |
+
`attn_res` op. Kimi K3 ran from its released MXFP4 checkpoint, not a bf16 one; its activations are
|
| 102 |
+
bfloat16.
|
| 103 |
+
- The transcript ends with a final user turn and an assistant answer, closed by the end-of-turn token,
|
| 104 |
+
rendered with the model's chat template with thinking disabled. Nemotron-3: the template's default,
|
| 105 |
+
which removes the reasoning of every assistant turn before the final user turn. Kimi K3: its template
|
| 106 |
+
cannot disable thinking, so the final assistant turn has an empty reasoning block before the answer.
|
| 107 |
+
- Linear and MLP read one token: the token before the final end-of-turn token. For Qwen and Nemotron-3
|
| 108 |
+
that is the answer's last token. For Kimi K3 it is the `<|sep|>` that closes `<|close|>message`, after
|
| 109 |
+
the answer's `<|close|>response<|sep|>`. EFC and axial read every token from the start of the final user
|
| 110 |
+
turn through the end-of-turn token.
|
| 111 |
+
|
| 112 |
+
```
|
| 113 |
+
Qwen3.5: <|im_start|>user\n{question}<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\nNo.<|im_end|>
|
| 114 |
+
Nemotron-3: <|im_start|>user\n{question}<|im_end|>\n<|im_start|>assistant\n<think></think>No.<|im_end|>
|
| 115 |
+
Kimi K3: <|open|>message role="user"<|sep|>{question}<|close|>message<|sep|><|end_of_msg|><|open|>message role="assistant"<|sep|><|open|>think<|sep|><|close|>think<|sep|><|open|>response<|sep|>No.<|close|>response<|sep|><|close|>message<|sep|><|end_of_msg|>
|
| 116 |
+
```
|
| 117 |
|
| 118 |
## Licences and attribution
|
| 119 |
|
| 120 |
The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`).
|
| 121 |
|
| 122 |
+
They are derived from the models below. Their licence texts ship here unchanged, and their attribution
|
| 123 |
+
notices are kept in `NOTICE`:
|
|
|
|
| 124 |
|
| 125 |
| Model | Licence | Licence file |
|
| 126 |
|---|---|---|
|
| 127 |
+
| Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B | Apache-2.0 | `LICENSE-QWEN-APACHE-2.0.txt` |
|
| 128 |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | NVIDIA Nemotron Open Model License (v. December 15, 2025) | `LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt` |
|
| 129 |
+
| nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | OpenMDW License Agreement, version 1.1 | `LICENSE-OPENMDW-1.1.txt` |
|
| 130 |
+
| moonshotai/Kimi-K3 | Kimi K3 License | `LICENSE-KIMI-K3.txt` |
|
| 131 |
|
| 132 |
Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.
|
kimi-k3/axial/config.json
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "AxialProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"num_layers": 6,
|
| 6 |
+
"d_proj": 256,
|
| 7 |
+
"n_attn_heads": 4,
|
| 8 |
+
"n_blocks": 4,
|
| 9 |
+
"d_ff": 1024,
|
| 10 |
+
"dropout": 0.0,
|
| 11 |
+
"use_layer_embed": true,
|
| 12 |
+
"pool_mode": "cls",
|
| 13 |
+
"use_checkpoint": true,
|
| 14 |
+
"normalize_input": "centered_unit_norm",
|
| 15 |
+
"proj_adapter_rank": 0,
|
| 16 |
+
"input_adapter_rank": 0,
|
| 17 |
+
"sliding_window": null,
|
| 18 |
+
"layer_keys": [
|
| 19 |
+
"28",
|
| 20 |
+
"39",
|
| 21 |
+
"51",
|
| 22 |
+
"62",
|
| 23 |
+
"74",
|
| 24 |
+
"84"
|
| 25 |
+
]
|
| 26 |
+
}
|
| 27 |
+
}
|
kimi-k3/axial/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e9f58be05ae2a458a158a6a137b2781d67086bdc51998023bf7aa415e15f9eec
|
| 3 |
+
size 32730360
|
kimi-k3/axial/probe_metadata.json
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model": "moonshotai/Kimi-K3",
|
| 3 |
+
"model_revision": "f831ab66814297da540d832a5235f8e904f29d06",
|
| 4 |
+
"architecture": "axial",
|
| 5 |
+
"layers": [
|
| 6 |
+
28,
|
| 7 |
+
39,
|
| 8 |
+
51,
|
| 9 |
+
62,
|
| 10 |
+
74,
|
| 11 |
+
84
|
| 12 |
+
],
|
| 13 |
+
"eval_sequence_aggregator": "last",
|
| 14 |
+
"obfuscate_over": "last-user-and-assistant-generation"
|
| 15 |
+
}
|
kimi-k3/efc/config.json
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "EFCProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"num_layers": 6,
|
| 5 |
+
"d_model": 7168,
|
| 6 |
+
"feature_mode": "combined",
|
| 7 |
+
"center": true,
|
| 8 |
+
"d_hidden": 64,
|
| 9 |
+
"d_probe": 256,
|
| 10 |
+
"normalization": "rms",
|
| 11 |
+
"normalization_eps": 1e-06,
|
| 12 |
+
"shrinkage": "fixed",
|
| 13 |
+
"shrinkage_alpha": 0.1,
|
| 14 |
+
"spectral_transform": "eigh",
|
| 15 |
+
"newton_schulz_iterations": 10,
|
| 16 |
+
"jitter": 1e-05,
|
| 17 |
+
"feature_chunk_size": 4,
|
| 18 |
+
"normalize_input": "none",
|
| 19 |
+
"layer_keys": [
|
| 20 |
+
"28",
|
| 21 |
+
"39",
|
| 22 |
+
"51",
|
| 23 |
+
"62",
|
| 24 |
+
"74",
|
| 25 |
+
"84"
|
| 26 |
+
]
|
| 27 |
+
}
|
| 28 |
+
}
|
kimi-k3/efc/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ac7a0b7ecfc54147598c274e95b03b94eff9b2145270b29e671431941af31d0b
|
| 3 |
+
size 22156448
|
kimi-k3/efc/probe_metadata.json
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model": "moonshotai/Kimi-K3",
|
| 3 |
+
"model_revision": "f831ab66814297da540d832a5235f8e904f29d06",
|
| 4 |
+
"architecture": "efc",
|
| 5 |
+
"layers": [
|
| 6 |
+
28,
|
| 7 |
+
39,
|
| 8 |
+
51,
|
| 9 |
+
62,
|
| 10 |
+
74,
|
| 11 |
+
84
|
| 12 |
+
],
|
| 13 |
+
"eval_sequence_aggregator": "mean",
|
| 14 |
+
"obfuscate_over": "last-user-and-assistant-generation"
|
| 15 |
+
}
|
kimi-k3/linear/layer_28/config.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "LinearProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"nhead": 1,
|
| 6 |
+
"normalize_input": "unit_norm"
|
| 7 |
+
}
|
| 8 |
+
}
|
kimi-k3/linear/layer_28/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:911ba849ea162e8c2e156d4d79d5b7b284cb9ba66f85229d0c75749ac99472c8
|
| 3 |
+
size 31421
|
kimi-k3/linear/layer_39/config.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "LinearProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"nhead": 1,
|
| 6 |
+
"normalize_input": "unit_norm"
|
| 7 |
+
}
|
| 8 |
+
}
|
kimi-k3/linear/layer_39/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b2d9f1ea91389baeeef2a53d5031d3cc4dafdc1d09c628e44e41716dfb14b704
|
| 3 |
+
size 31421
|
kimi-k3/linear/layer_51/config.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "LinearProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"nhead": 1,
|
| 6 |
+
"normalize_input": "unit_norm"
|
| 7 |
+
}
|
| 8 |
+
}
|
kimi-k3/linear/layer_51/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:bdce0165c3fc3327d76598c0ab6f02417e9dec4e9fbf1d4390a2348dd63e182c
|
| 3 |
+
size 31421
|
kimi-k3/linear/layer_62/config.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "LinearProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"nhead": 1,
|
| 6 |
+
"normalize_input": "unit_norm"
|
| 7 |
+
}
|
| 8 |
+
}
|
kimi-k3/linear/layer_62/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:efe2e891fb0d1ade54648a18e33abb17941825ce04b4160724ce223945d7b985
|
| 3 |
+
size 31421
|
kimi-k3/linear/layer_74/config.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "LinearProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"nhead": 1,
|
| 6 |
+
"normalize_input": "unit_norm"
|
| 7 |
+
}
|
| 8 |
+
}
|
kimi-k3/linear/layer_74/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:78c6e9e498c8eb19a012e0afb2ea0b08c6240daedd3d1059c98b16f6d3f43e88
|
| 3 |
+
size 31421
|
kimi-k3/linear/layer_84/config.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "LinearProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"nhead": 1,
|
| 6 |
+
"normalize_input": "unit_norm"
|
| 7 |
+
}
|
| 8 |
+
}
|
kimi-k3/linear/layer_84/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6c315787f11b2da67c63e6597516c5727bb395604ea20bb53a7af55f513d16b8
|
| 3 |
+
size 31421
|
kimi-k3/linear/probe_metadata.json
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model": "moonshotai/Kimi-K3",
|
| 3 |
+
"model_revision": "f831ab66814297da540d832a5235f8e904f29d06",
|
| 4 |
+
"architecture": "linear",
|
| 5 |
+
"layers": [
|
| 6 |
+
28,
|
| 7 |
+
39,
|
| 8 |
+
51,
|
| 9 |
+
62,
|
| 10 |
+
74,
|
| 11 |
+
84
|
| 12 |
+
],
|
| 13 |
+
"eval_sequence_aggregator": "mean",
|
| 14 |
+
"obfuscate_over": "second-last-token-generation",
|
| 15 |
+
"layer_rule": {
|
| 16 |
+
"used_layers": [
|
| 17 |
+
28,
|
| 18 |
+
39,
|
| 19 |
+
51,
|
| 20 |
+
62,
|
| 21 |
+
74,
|
| 22 |
+
84
|
| 23 |
+
]
|
| 24 |
+
}
|
| 25 |
+
}
|
kimi-k3/mlp/layer_28/config.json
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "MLPProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"d_mlp": 256,
|
| 6 |
+
"nhead": 1,
|
| 7 |
+
"normalize_input": "unit_norm",
|
| 8 |
+
"activation": "relu",
|
| 9 |
+
"d_mlp2": null
|
| 10 |
+
}
|
| 11 |
+
}
|
kimi-k3/mlp/layer_28/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:42354c090cba70cccedcd68f538dc22a1fce8168724c582b21dbdb227db1ce83
|
| 3 |
+
size 7345201
|
kimi-k3/mlp/layer_39/config.json
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "MLPProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"d_mlp": 256,
|
| 6 |
+
"nhead": 1,
|
| 7 |
+
"normalize_input": "unit_norm",
|
| 8 |
+
"activation": "relu",
|
| 9 |
+
"d_mlp2": null
|
| 10 |
+
}
|
| 11 |
+
}
|
kimi-k3/mlp/layer_39/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b6769d33b59fce353f360588b6e5e8c54a6dd4f4624948830e852760ff8a01ee
|
| 3 |
+
size 7345201
|
kimi-k3/mlp/layer_51/config.json
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "MLPProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"d_mlp": 256,
|
| 6 |
+
"nhead": 1,
|
| 7 |
+
"normalize_input": "unit_norm",
|
| 8 |
+
"activation": "relu",
|
| 9 |
+
"d_mlp2": null
|
| 10 |
+
}
|
| 11 |
+
}
|
kimi-k3/mlp/layer_51/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c6604a356b9e5ac866f49dc24213cc08b332bd771321f31083f2dfdf33b25432
|
| 3 |
+
size 7345201
|
kimi-k3/mlp/layer_62/config.json
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "MLPProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"d_mlp": 256,
|
| 6 |
+
"nhead": 1,
|
| 7 |
+
"normalize_input": "unit_norm",
|
| 8 |
+
"activation": "relu",
|
| 9 |
+
"d_mlp2": null
|
| 10 |
+
}
|
| 11 |
+
}
|
kimi-k3/mlp/layer_62/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7d4dda19ec1c77b9fda81bb7fa6697cbf484f1a76849b7080657fa882b0259ee
|
| 3 |
+
size 7345201
|
kimi-k3/mlp/layer_74/config.json
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "MLPProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"d_mlp": 256,
|
| 6 |
+
"nhead": 1,
|
| 7 |
+
"normalize_input": "unit_norm",
|
| 8 |
+
"activation": "relu",
|
| 9 |
+
"d_mlp2": null
|
| 10 |
+
}
|
| 11 |
+
}
|
kimi-k3/mlp/layer_74/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:28bd24aaa1f446b193ab82b1ac9d1df16957ebdb10dc1c3cc6013c68d2cf3edd
|
| 3 |
+
size 7345201
|
kimi-k3/mlp/layer_84/config.json
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"class_name": "MLPProbe",
|
| 3 |
+
"init_args": {
|
| 4 |
+
"d_model": 7168,
|
| 5 |
+
"d_mlp": 256,
|
| 6 |
+
"nhead": 1,
|
| 7 |
+
"normalize_input": "unit_norm",
|
| 8 |
+
"activation": "relu",
|
| 9 |
+
"d_mlp2": null
|
| 10 |
+
}
|
| 11 |
+
}
|
kimi-k3/mlp/layer_84/model.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ce505c1fdb735b7d1fe93113768cc52e0259a4940acf77489d35a8149e26e1d8
|
| 3 |
+
size 7345201
|
kimi-k3/mlp/probe_metadata.json
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model": "moonshotai/Kimi-K3",
|
| 3 |
+
"model_revision": "f831ab66814297da540d832a5235f8e904f29d06",
|
| 4 |
+
"architecture": "mlp",
|
| 5 |
+
"layers": [
|
| 6 |
+
28,
|
| 7 |
+
39,
|
| 8 |
+
51,
|
| 9 |
+
62,
|
| 10 |
+
74,
|
| 11 |
+
84
|
| 12 |
+
],
|
| 13 |
+
"eval_sequence_aggregator": "mean",
|
| 14 |
+
"obfuscate_over": "second-last-token-generation",
|
| 15 |
+
"layer_rule": {
|
| 16 |
+
"used_layers": [
|
| 17 |
+
28,
|
| 18 |
+
39,
|
| 19 |
+
51,
|
| 20 |
+
62,
|
| 21 |
+
74,
|
| 22 |
+
84
|
| 23 |
+
]
|
| 24 |
+
}
|
| 25 |
+
}
|
nemotron-3-nano-30b-a3b/axial/model.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 28035320
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:345567d256088641608f0babda239691ebe8f3ad68418e4e386afe7ab20b8441
|
| 3 |
size 28035320
|
nemotron-3-nano-30b-a3b/axial/probe_metadata.json
CHANGED
|
@@ -1,5 +1,4 @@
|
|
| 1 |
{
|
| 2 |
-
"schema_version": 1,
|
| 3 |
"model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
|
| 4 |
"model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
|
| 5 |
"architecture": "axial",
|
|
@@ -11,6 +10,6 @@
|
|
| 11 |
42,
|
| 12 |
47
|
| 13 |
],
|
| 14 |
-
"
|
| 15 |
-
"
|
| 16 |
}
|
|
|
|
| 1 |
{
|
|
|
|
| 2 |
"model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
|
| 3 |
"model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
|
| 4 |
"architecture": "axial",
|
|
|
|
| 10 |
42,
|
| 11 |
47
|
| 12 |
],
|
| 13 |
+
"eval_sequence_aggregator": "last",
|
| 14 |
+
"obfuscate_over": "last-user-and-assistant-generation"
|
| 15 |
}
|
nemotron-3-nano-30b-a3b/efc/model.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 8393888
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:42f946be32dcdead9c673be7f397bb18b221cd05cbd2f37fb7d9dc36e7e8f392
|
| 3 |
size 8393888
|
nemotron-3-nano-30b-a3b/efc/probe_metadata.json
CHANGED
|
@@ -1,5 +1,4 @@
|
|
| 1 |
{
|
| 2 |
-
"schema_version": 1,
|
| 3 |
"model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
|
| 4 |
"model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
|
| 5 |
"architecture": "efc",
|
|
@@ -11,6 +10,6 @@
|
|
| 11 |
42,
|
| 12 |
47
|
| 13 |
],
|
| 14 |
-
"
|
| 15 |
-
"
|
| 16 |
}
|
|
|
|
| 1 |
{
|
|
|
|
| 2 |
"model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
|
| 3 |
"model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
|
| 4 |
"architecture": "efc",
|
|
|
|
| 10 |
42,
|
| 11 |
47
|
| 12 |
],
|
| 13 |
+
"eval_sequence_aggregator": "mean",
|
| 14 |
+
"obfuscate_over": "last-user-and-assistant-generation"
|
| 15 |
}
|
nemotron-3-nano-30b-a3b/linear/layer_16/model.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 13501
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:8ee8243e4e836f290d471616b9cc6a7391c5d908e3d18ae3222b1ac7b816c763
|
| 3 |
size 13501
|
nemotron-3-nano-30b-a3b/linear/layer_22/model.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 13501
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:d0f1deda31003d2a6e7b1af3231a1d714cf2df1bdf27f697359db30d31b3e988
|
| 3 |
size 13501
|
nemotron-3-nano-30b-a3b/linear/layer_29/model.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 13501
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:61c9fb00b39d038fca5e854d82d1dd1f0f7b8598e56b17f4b5f4578861aff6fd
|
| 3 |
size 13501
|
nemotron-3-nano-30b-a3b/linear/layer_35/model.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 13501
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7c92f889fae2e52832f25b96b25133c07e1a5c84e5d279c8cb047e732b933ee4
|
| 3 |
size 13501
|
nemotron-3-nano-30b-a3b/linear/layer_42/model.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 13501
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ffc79f4fccbae06ada63dd0925f53ab33963a8a89e5aaf969d7a1764bdf3fde4
|
| 3 |
size 13501
|
nemotron-3-nano-30b-a3b/linear/layer_47/model.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 13501
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:8c6543a8a2428395d9bb59fea8d4ec56d58bfc47dafd4783970af4d4d72de7fe
|
| 3 |
size 13501
|
nemotron-3-nano-30b-a3b/linear/probe_metadata.json
CHANGED
|
@@ -1,5 +1,4 @@
|
|
| 1 |
{
|
| 2 |
-
"schema_version": 1,
|
| 3 |
"model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
|
| 4 |
"model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
|
| 5 |
"architecture": "linear",
|
|
@@ -11,6 +10,16 @@
|
|
| 11 |
42,
|
| 12 |
47
|
| 13 |
],
|
| 14 |
-
"
|
| 15 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
}
|
|
|
|
| 1 |
{
|
|
|
|
| 2 |
"model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16",
|
| 3 |
"model_revision": "bf77c3174f68ad409e1c2aa60daeb46e32d1c606",
|
| 4 |
"architecture": "linear",
|
|
|
|
| 10 |
42,
|
| 11 |
47
|
| 12 |
],
|
| 13 |
+
"eval_sequence_aggregator": "mean",
|
| 14 |
+
"obfuscate_over": "second-last-token-generation",
|
| 15 |
+
"layer_rule": {
|
| 16 |
+
"used_layers": [
|
| 17 |
+
16,
|
| 18 |
+
22,
|
| 19 |
+
29,
|
| 20 |
+
35,
|
| 21 |
+
42,
|
| 22 |
+
47
|
| 23 |
+
]
|
| 24 |
+
}
|
| 25 |
}
|
nemotron-3-nano-30b-a3b/mlp/layer_16/model.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 2757681
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:86bc24e945d5397d1fa96bdf09976dc3c67d04afbe0d395f138c82fe9ed4bf85
|
| 3 |
size 2757681
|
nemotron-3-nano-30b-a3b/mlp/layer_22/model.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 2757681
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e44cbd9a75b74730c316b09ce3c1d807f4b735efa8d13b19c094291472cdd6c8
|
| 3 |
size 2757681
|
nemotron-3-nano-30b-a3b/mlp/layer_29/model.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 2757681
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:bcc9eae57322e9c86f7f256425add809beb28eba92c52da4652e583417d8ef88
|
| 3 |
size 2757681
|