PEFT
Safetensors
English
sev
research
cybersecurity
agent-activity
decision-model
lora
File size: 8,140 Bytes
8c19e93
 
 
 
 
 
 
 
 
a1824ae
8c19e93
 
 
 
 
 
 
 
da0a131
a1824ae
 
7f9f3f5
 
 
 
 
 
 
 
 
 
 
 
8c19e93
 
 
da0a131
 
a1824ae
8c19e93
 
 
 
 
7f9f3f5
8c19e93
 
da0a131
 
 
 
a1824ae
7f9f3f5
 
 
 
 
da0a131
 
 
8c19e93
da0a131
8c19e93
 
 
 
da0a131
 
 
 
 
8c19e93
a1824ae
8c19e93
a1824ae
 
da0a131
a1824ae
 
da0a131
 
 
 
 
 
a1824ae
7f9f3f5
 
da0a131
 
 
 
a1824ae
da0a131
a1824ae
7f9f3f5
a1824ae
7f9f3f5
 
 
 
 
 
 
 
da0a131
7f9f3f5
da0a131
7f9f3f5
 
da0a131
7f9f3f5
da0a131
 
 
a1824ae
 
 
 
da0a131
 
a1824ae
da0a131
 
 
 
a1824ae
7f9f3f5
 
 
 
a1824ae
da0a131
a1824ae
da0a131
 
 
 
 
a1824ae
da0a131
 
7f9f3f5
 
da0a131
 
 
a1824ae
 
 
7f9f3f5
 
da0a131
a1824ae
da0a131
 
 
 
7f9f3f5
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
---
language:
- en
license: apache-2.0
base_model: Qwen/Qwen3.5-4B-Base
base_model_relation: adapter
library_name: peft
datasets:
- macmacmacmac/Sev-behavioral-research-v1
- macmacmacmac/openai-agent-swarmtraces
tags:
- sev
- research
- cybersecurity
- agent-activity
- decision-model
- lora
---
# Sev-4B: security evidence and response policies

Sev-4B scores supplied answers to questions about security logs and recovered
programs. **v0.3.1-html-contrast-research** continues
`v0.3.0-response-policy-research`. It is a Qwen3.5-4B LoRA adapter with a
decision head. It does not generate text.

On 175 policy questions from 35 source programs held out of training, this
checkpoint answers **158/175, 90.3%**, up from 147/175 on v0.3.0. At the
calibration-selected alert cut it catches **35/35** required approvals and
raises **5/140** false alerts, **3.6%**, under the 5% ceiling. v0.3.0 raised
14/140 and failed that ceiling. Twelve of those fourteen were allowed HTML
reads. This pass trains the contrast on the same program: writing `innerText`
requires approval, and parsing HTML does not. The DNS diagnostic was not rerun.
The trial's isolation-and-packing check failed. v0.3.0 remains available.

## Use

Install the [Sev runtime](https://github.com/maceip/Sev). A generic text-generation
pipeline does not load the decision head.

```bash
git clone https://github.com/maceip/Sev.git
cd Sev
uv sync --extra serve
uv run python -m kev.serve \
  --run macmacmacmac/Sev-4B@v0.3.1-html-contrast-research --port 8009
```

Send TypeSafe-shaped requests to `POST /v1/systemone`. Supply the observed
evidence, the applicable policy, and explicit answer choices. The model reads
supplied code without executing it. Its scores support analyst review; they do
not establish that a program ran, identify its human owner, or prove agent origin.

The checkpoint ships with temperature **1.0**. The false-alert cut for the
policy panel is **1.0** on that raw approval probability: alert only when the
approval option is certain. That cut was chosen on the 140 policy calibration
questions, at a 5% false-alert budget, and then applied to development.
Reported evaluation uses
`KEV_BACKEND=torch KEV_DTYPE=fp32 KEV_MERGE=1` on CUDA. Accelerated serving can
differ numerically. The API's `confidence` rescales the largest probability above
chance; it is not an independently verified correctness probability.

## Training and lineage

| Setting | Value |
| --- | --- |
| Backbone | `Qwen/Qwen3.5-4B-Base@1001bb4d826a52d1f399e183466143f4da7b741b` |
| Immediate parent | `macmacmacmac/Sev-4B@a1824aefba305fda86e3503b895a7d9b3871f79a` |
| Selected run | `sev-r2-response-policy-4b-v2/00-trial-0` |
| Training records / questions | 6,577 / 9,456 |
| Retained curriculum records | 6,438 |
| Added source programs / policy questions | 139 / 278 |
| Epochs / seed | 1 / 4 |
| Learning rate | `2.5e-6` |
| Batch / accumulation | 4 / 2 |
| LoRA / head | Rank 16, all targets / 256 dimensions |
| Precision | fp32 frozen weights, bf16 autocast |
| Updates / forward tokens | 823 / 1,505,379 |
| Rejected or truncated training records | 0 |

All previous curriculum records remain byte-identical. They include native
Sysmon, ExCyTIn, GUIDE, public classification and authored-rule replay, and 446
SwarmTraces static-program records. The added programs have authored API-specific
policies and checked labels. Their source text is real recovered code, not a
native execution trace. Source-closure groups separate training, calibration
and development. No locked test was read.

[Training configuration](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/training_config.json), [provenance](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/provenance.json),
and [data lineage](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/DATA_PROVENANCE.md) pin the inputs. The original
[`v0.1.0-research`](https://huggingface.co/macmacmacmac/Sev-4B/tree/v0.1.0-research)
synthetic-origin preview is a separate historical lineage. The previous
[`v0.2.0-swarmtraces-research`](https://huggingface.co/macmacmacmac/Sev-4B/tree/v0.2.0-swarmtraces-research)
release remains available unchanged.

## Matched development results

| Policy decision | v0.3.0 | v0.3.1 |
| --- | ---: | ---: |
| Correct argmax answers | 147/175 | 158/175 |
| Required approvals detected at the alert cut | 25/35 | 35/35 |
| False alerts at that cut | 14/140 | 5/140 |
| Allowed HTML questions flagged | 12/21 | 0/21 |

Each alert cut was selected on calibration only, with a 5% false-alert budget.
v0.3.1's cut is 1.0 on the raw approval probability. The 5 remaining false
alerts are 4 console questions and 1 plain-text question. These are explicit
policy decisions on a development panel, not measured detection rates in live
networks.

The retained-task table and the DNS result below were measured on v0.3.0.
They were not rerun for v0.3.1.

| Retained panel | v0.2.0 | v0.3.0, not rerun here |
| --- | ---: | ---: |
| Manual SwarmTraces | 65/83 | 66/83 |
| Static SwarmTraces | 345/354 | 347/354 |
| Native Sysmon | 409/410 | 409/410 |
| ExCyTIn | 398/398 | 398/398 |
| Original Sysmon | 62/64 | 62/64 |
| General | 86/105 | 86/105 |
| GUIDE triage | 130/231 | 131/231 |
| GUIDE detector | 638/1,064 | 639/1,064 |

All 25 correctness-retention checks pass. Unchanged totals do not imply unchanged
individual answers. GUIDE detector calibrated negative log loss worsens by
0.0116 and Brier by 0.0077. No matched continuation without the new component was
run, so the isolated causal contribution of SwarmTraces is not established.

On v0.3.0 the DNS diagnostic was **21/32** and failed the zero-new-error check.
v0.3.1 did not rerun it.
[Policy evaluation](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/evaluation.json), [DNS evaluation](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/dns-evaluation.json), and
[registered screen](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/screen.json) retain the complete comparisons.

## Calibration and artifact verification

The shipped temperature minimizes question-weighted negative log loss on 2,674
calibration questions from 1,920 records. Five-fold diagnostics keep all 465
source groups intact across task families. Raw calibration ECE is 0.07626 and
out-of-fold ECE is 0.03386. This is a calibration diagnostic, not fresh field
validation. Development and DNS observations were excluded from fitting.

Only temperature metadata changed when assembling the serving copy. Learned
head tensors, adapter and tokenizer bytes match the evaluated checkpoint.
[Calibration](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/calibration.json), [integrity evidence](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/checkpoint-integrity.json),
and [SHA256SUMS](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/SHA256SUMS) identify the release files.

Research iteration is paused following this release. The 0.8B and 9B checkpoints
are not updated by this publication.

## License and attribution

Source and adapter/head weights carry Apache-2.0 notices. Preserve [LICENSE](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/LICENSE),
[NOTICE](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/NOTICE) and the Qwen [BASE_LICENSE](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/BASE_LICENSE). Sev builds on Kev by
Jared Palmer and Qwen3.5 by the Qwen team.

Each dataset retains its own terms. Upstream SwarmTraces reuse terms remain
unverified; the model license does not relicense those artifacts. The collection
includes metadata-only entries, and membership does not imply training use.
The historical synthetic source's notice applies only to that source. See
[data provenance](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/DATA_PROVENANCE.md).