Instructions to use macmacmacmac/Sev-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use macmacmacmac/Sev-4B with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download README.md from macmacmacmac/Sev-4B: direct link, hf CLI and curl.
- Browser
- Download file 8.14 kB
-
https://huggingface.co/macmacmacmac/Sev-4B/resolve/main/README.md
- Command line
-
hf download hf://macmacmacmac/Sev-4B/README.md
-
curl -L -o README.md https://huggingface.co/macmacmacmac/Sev-4B/resolve/main/README.md
8.14 kB
| language: | |
| - en | |
| license: apache-2.0 | |
| base_model: Qwen/Qwen3.5-4B-Base | |
| base_model_relation: adapter | |
| library_name: peft | |
| datasets: | |
| - macmacmacmac/Sev-behavioral-research-v1 | |
| - macmacmacmac/openai-agent-swarmtraces | |
| tags: | |
| - sev | |
| - research | |
| - cybersecurity | |
| - agent-activity | |
| - decision-model | |
| - lora | |
| # Sev-4B: security evidence and response policies | |
| Sev-4B scores supplied answers to questions about security logs and recovered | |
| programs. **v0.3.1-html-contrast-research** continues | |
| `v0.3.0-response-policy-research`. It is a Qwen3.5-4B LoRA adapter with a | |
| decision head. It does not generate text. | |
| On 175 policy questions from 35 source programs held out of training, this | |
| checkpoint answers **158/175, 90.3%**, up from 147/175 on v0.3.0. At the | |
| calibration-selected alert cut it catches **35/35** required approvals and | |
| raises **5/140** false alerts, **3.6%**, under the 5% ceiling. v0.3.0 raised | |
| 14/140 and failed that ceiling. Twelve of those fourteen were allowed HTML | |
| reads. This pass trains the contrast on the same program: writing `innerText` | |
| requires approval, and parsing HTML does not. The DNS diagnostic was not rerun. | |
| The trial's isolation-and-packing check failed. v0.3.0 remains available. | |
| ## Use | |
| Install the [Sev runtime](https://github.com/maceip/Sev). A generic text-generation | |
| pipeline does not load the decision head. | |
| ```bash | |
| git clone https://github.com/maceip/Sev.git | |
| cd Sev | |
| uv sync --extra serve | |
| uv run python -m kev.serve \ | |
| --run macmacmacmac/Sev-4B@v0.3.1-html-contrast-research --port 8009 | |
| ``` | |
| Send TypeSafe-shaped requests to `POST /v1/systemone`. Supply the observed | |
| evidence, the applicable policy, and explicit answer choices. The model reads | |
| supplied code without executing it. Its scores support analyst review; they do | |
| not establish that a program ran, identify its human owner, or prove agent origin. | |
| The checkpoint ships with temperature **1.0**. The false-alert cut for the | |
| policy panel is **1.0** on that raw approval probability: alert only when the | |
| approval option is certain. That cut was chosen on the 140 policy calibration | |
| questions, at a 5% false-alert budget, and then applied to development. | |
| Reported evaluation uses | |
| `KEV_BACKEND=torch KEV_DTYPE=fp32 KEV_MERGE=1` on CUDA. Accelerated serving can | |
| differ numerically. The API's `confidence` rescales the largest probability above | |
| chance; it is not an independently verified correctness probability. | |
| ## Training and lineage | |
| | Setting | Value | | |
| | --- | --- | | |
| | Backbone | `Qwen/Qwen3.5-4B-Base@1001bb4d826a52d1f399e183466143f4da7b741b` | | |
| | Immediate parent | `macmacmacmac/Sev-4B@a1824aefba305fda86e3503b895a7d9b3871f79a` | | |
| | Selected run | `sev-r2-response-policy-4b-v2/00-trial-0` | | |
| | Training records / questions | 6,577 / 9,456 | | |
| | Retained curriculum records | 6,438 | | |
| | Added source programs / policy questions | 139 / 278 | | |
| | Epochs / seed | 1 / 4 | | |
| | Learning rate | `2.5e-6` | | |
| | Batch / accumulation | 4 / 2 | | |
| | LoRA / head | Rank 16, all targets / 256 dimensions | | |
| | Precision | fp32 frozen weights, bf16 autocast | | |
| | Updates / forward tokens | 823 / 1,505,379 | | |
| | Rejected or truncated training records | 0 | | |
| All previous curriculum records remain byte-identical. They include native | |
| Sysmon, ExCyTIn, GUIDE, public classification and authored-rule replay, and 446 | |
| SwarmTraces static-program records. The added programs have authored API-specific | |
| policies and checked labels. Their source text is real recovered code, not a | |
| native execution trace. Source-closure groups separate training, calibration | |
| and development. No locked test was read. | |
| [Training configuration](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/training_config.json), [provenance](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/provenance.json), | |
| and [data lineage](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/DATA_PROVENANCE.md) pin the inputs. The original | |
| [`v0.1.0-research`](https://huggingface.co/macmacmacmac/Sev-4B/tree/v0.1.0-research) | |
| synthetic-origin preview is a separate historical lineage. The previous | |
| [`v0.2.0-swarmtraces-research`](https://huggingface.co/macmacmacmac/Sev-4B/tree/v0.2.0-swarmtraces-research) | |
| release remains available unchanged. | |
| ## Matched development results | |
| | Policy decision | v0.3.0 | v0.3.1 | | |
| | --- | ---: | ---: | | |
| | Correct argmax answers | 147/175 | 158/175 | | |
| | Required approvals detected at the alert cut | 25/35 | 35/35 | | |
| | False alerts at that cut | 14/140 | 5/140 | | |
| | Allowed HTML questions flagged | 12/21 | 0/21 | | |
| Each alert cut was selected on calibration only, with a 5% false-alert budget. | |
| v0.3.1's cut is 1.0 on the raw approval probability. The 5 remaining false | |
| alerts are 4 console questions and 1 plain-text question. These are explicit | |
| policy decisions on a development panel, not measured detection rates in live | |
| networks. | |
| The retained-task table and the DNS result below were measured on v0.3.0. | |
| They were not rerun for v0.3.1. | |
| | Retained panel | v0.2.0 | v0.3.0, not rerun here | | |
| | --- | ---: | ---: | | |
| | Manual SwarmTraces | 65/83 | 66/83 | | |
| | Static SwarmTraces | 345/354 | 347/354 | | |
| | Native Sysmon | 409/410 | 409/410 | | |
| | ExCyTIn | 398/398 | 398/398 | | |
| | Original Sysmon | 62/64 | 62/64 | | |
| | General | 86/105 | 86/105 | | |
| | GUIDE triage | 130/231 | 131/231 | | |
| | GUIDE detector | 638/1,064 | 639/1,064 | | |
| All 25 correctness-retention checks pass. Unchanged totals do not imply unchanged | |
| individual answers. GUIDE detector calibrated negative log loss worsens by | |
| 0.0116 and Brier by 0.0077. No matched continuation without the new component was | |
| run, so the isolated causal contribution of SwarmTraces is not established. | |
| On v0.3.0 the DNS diagnostic was **21/32** and failed the zero-new-error check. | |
| v0.3.1 did not rerun it. | |
| [Policy evaluation](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/evaluation.json), [DNS evaluation](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/dns-evaluation.json), and | |
| [registered screen](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/screen.json) retain the complete comparisons. | |
| ## Calibration and artifact verification | |
| The shipped temperature minimizes question-weighted negative log loss on 2,674 | |
| calibration questions from 1,920 records. Five-fold diagnostics keep all 465 | |
| source groups intact across task families. Raw calibration ECE is 0.07626 and | |
| out-of-fold ECE is 0.03386. This is a calibration diagnostic, not fresh field | |
| validation. Development and DNS observations were excluded from fitting. | |
| Only temperature metadata changed when assembling the serving copy. Learned | |
| head tensors, adapter and tokenizer bytes match the evaluated checkpoint. | |
| [Calibration](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/calibration.json), [integrity evidence](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/checkpoint-integrity.json), | |
| and [SHA256SUMS](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/SHA256SUMS) identify the release files. | |
| Research iteration is paused following this release. The 0.8B and 9B checkpoints | |
| are not updated by this publication. | |
| ## License and attribution | |
| Source and adapter/head weights carry Apache-2.0 notices. Preserve [LICENSE](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/LICENSE), | |
| [NOTICE](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/NOTICE) and the Qwen [BASE_LICENSE](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/BASE_LICENSE). Sev builds on Kev by | |
| Jared Palmer and Qwen3.5 by the Qwen team. | |
| Each dataset retains its own terms. Upstream SwarmTraces reuse terms remain | |
| unverified; the model license does not relicense those artifacts. The collection | |
| includes metadata-only entries, and membership does not imply training use. | |
| The historical synthetic source's notice applies only to that source. See | |
| [data provenance](https://huggingface.co/macmacmacmac/Sev-4B/blob/v0.3.0-response-policy-research/DATA_PROVENANCE.md). | |