|
Download README.md from billytesterman/secjev-encoder: direct link, hf CLI and curl.
- Browser
- Download file 7.5 kB
-
https://huggingface.co/billytesterman/secjev-encoder/resolve/main/README.md
- Command line
-
hf download hf://billytesterman/secjev-encoder/README.md
-
curl -L -o README.md https://huggingface.co/billytesterman/secjev-encoder/resolve/main/README.md
7.5 kB
| license: other | |
| license_name: lfm1.0 | |
| license_link: LICENSE | |
| base_model: LiquidAI/LFM2.5-Encoder-350M | |
| pipeline_tag: text-classification | |
| library_name: onnx | |
| tags: | |
| - security | |
| - code | |
| - cwe | |
| - vulnerability-detection | |
| - lfm2 | |
| - onnx | |
| - webgpu | |
| # secjev encoder | |
| [LFM2.5-Encoder-350M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M) fine-tuned to | |
| answer the 25 questions of the CWE Top 25 about a window of up to 240 lines of source code. | |
| One forward pass reads a window and returns P(yes) for all 25 questions, for example | |
| "Is a SQL query built here by concatenating or interpolating data, rather than by binding | |
| parameters?" (`cwe_89`). It reads 28 languages, from C and Go to PHP and Solidity. | |
| ``` | |
| window text ──► LFM2.5 bidirectional encoder ──► attention pool (one per question) ──► linear ──► sigmoid(logit / T) | |
| ``` | |
| It is a triage aid that ranks where to look first. It does not prove that a flaw exists. | |
| ## Files | |
| | path | what | | |
| |---|---| | |
| | `model/model.safetensors`, `model/secjev.json` | PyTorch weights (1.4 GB): the encoder plus the per-question attention pool and the linear head; labels, questions, temperature | | |
| | `onnx/model.onnx` | fp32 ONNX graph (1.4 GB): matches PyTorch to 1e-5 in the logits | | |
| | `onnx/model_q8.onnx` + `onnx/model_q8.weights` | 8-bit weights in blocks of 32 with fp32 maths (593 MB), for ONNX Runtime on CPU or CUDA, and WebGPU | | |
| | `onnx/model_q8_wasm.onnx` | the same weights file, with attention written as per-head operators, for onnxruntime-web on WebAssembly | | |
| | `onnx/bundle.json` | labels, questions, per-language limits, temperature, window rules and the file-extension map | | |
| | `*/tokenizer.json` | the base model's tokenizer | | |
| | `onnx/validate-validation.json` | the validation numbers below | | |
| Each ONNX graph takes `input_ids` (int64, `[1, n]`, n ≤ 8192, BOS first, no padding) for | |
| one window and returns `logits` (`[1, 25]`). P(yes) is `sigmoid(logit / temperature)`. | |
| Score windows one at a time, because the graphs use ONNX Runtime's `MultiHeadAttention` | |
| without a padding mask. That keeps an 8,192-token window from building a 4.3 GB score | |
| matrix. The two 8-bit graphs share `model_q8.weights`, so download it next to whichever | |
| graph you use. | |
| ## Input format | |
| The model was trained on windows of this exact shape: a header, then each line numbered | |
| and cut at 400 characters. Files are tiled every 240 lines from line 1, and a tile whose | |
| text is over 33,084 characters is halved until it fits. `bundle.json` → `window` holds | |
| these constants. `limits` says which languages each question applies to; the memory-safety | |
| questions, for example, are asked only of C and C++. | |
| ``` | |
| // ==== src/app.py:1-240 of 512 ==== | |
| 1 | import os, subprocess | |
| 2 | from flask import request | |
| ... | |
| ``` | |
| ## Usage (ONNX Runtime, Python) | |
| ```python | |
| import json | |
| import numpy as np | |
| import onnxruntime as ort | |
| from tokenizers import Tokenizer | |
| d = "onnx/" | |
| meta = json.load(open(d + "bundle.json")) | |
| tok = Tokenizer.from_file(d + "tokenizer.json") | |
| sess = ort.InferenceSession(d + "model_q8.onnx", providers=["CPUExecutionProvider"]) | |
| def window(path, lines, start, total): # the format the model was trained on | |
| body = "\n".join(f"{start + i} | {l[:400] + ' ... [long line cut]' if len(l) > 400 else l}" | |
| for i, l in enumerate(lines)) | |
| return f"// ==== {path}:{start}-{start + len(lines) - 1} of {total} ====\n" + body | |
| src = open("app.py").read().split("\n")[:240] | |
| ids = tok.encode(window("app.py", src, 1, len(src))).ids[: meta["max_length"]] | |
| logits = sess.run(None, {"input_ids": np.array([ids], dtype=np.int64)})[0][0] | |
| p = 1 / (1 + np.exp(-logits / meta["temperature"])) | |
| for i in np.argsort(-p)[:5]: | |
| cwe = meta["labels"][i] | |
| print(f"{p[i]:.3f} {cwe} {meta['questions'][cwe]}") | |
| ``` | |
| In a browser, load the graphs through the `onnxruntime-web/webgpu` entry, which runs 8-bit | |
| `MatMulNBits`. The package's default entry does not. Use `model_q8_wasm.onnx` when WebGPU | |
| is missing. Multi-threaded WebAssembly needs COOP/COEP headers. | |
| ## Training | |
| - **Real labels only.** Positives are the window before a CVE/GHSA fix, labelled on the | |
| questions its advisory names. The same place after the fix is the matching negative. Other | |
| questions on fix windows are masked. Windows from a sweep of ordinary open-source code are | |
| negatives on every question, at weight 0.25. No LLM-generated labels are used. | |
| - **Pair loss.** Within each fix, the window before the fix should score above the window | |
| after it (a logistic loss on the gap). This term made the model learn the flaw itself | |
| rather than "looks like code that gets fixed". | |
| - **Attention pooling.** Each question learns its own weighting of the window's tokens. The | |
| head learns at 5e-4 and the encoder at 3e-5, in bf16, up to 8,192 tokens. | |
| - **Selection and calibration.** The checkpoint was chosen by `auroc` on the fix validation | |
| windows. One temperature (1.6356) was fitted on held-out calibration windows. | |
| - **Data.** The permissively licensed part of the secjev dataset: 52,673 fix windows and | |
| 200,000 ordinary windows. Each fix window is seen twice per epoch. | |
| ## Evaluation | |
| Validation split: 3,060 fix windows plus 10,000 ordinary windows, 209,900 | |
| (window, question) scores in all. | |
| | | auroc (before vs after fix) | pre_above_post | before fix vs ordinary | fixed vs ordinary | ordinary flagged (p ≥ 0.5) | | |
| |---|---|---|---|---|---| | |
| | PyTorch, bf16 | 0.5830 | 0.6560 | 0.9369 | 0.8996 | 0.75% | | |
| | `model.onnx` (fp32) | 0.5832 | 0.6680 | 0.9368 | 0.8994 | 0.75% | | |
| | `model_q8.onnx` | 0.5830 | 0.6691 | 0.9370 | 0.8999 | 0.76% | | |
| | `model_q8_wasm.onnx` | 0.5830 | 0.6674 | 0.9370 | 0.8999 | 0.76% | | |
| What the columns mean: | |
| - **before fix vs ordinary (≈0.94):** the model separates code that later needed a security | |
| fix from ordinary code well. | |
| - **auroc and pre_above_post:** it separates the vulnerable version from its own fixed | |
| version, often a few changed lines, only modestly. | |
| - **fixed vs ordinary (0.90):** much of the signal is "this is the kind of code where flaws | |
| of this class live". Read a high score as "look here", not "this is vulnerable". | |
| Against fp32, the 8-bit graphs move a probability by 0.0005 on average (99th percentile | |
| 0.009, max 0.056). They keep 98.4% of the top 500 scores and 99.97% of the p ≥ 0.5 calls. | |
| Speed per window, q8 (RTX 4070 Ti SUPER, i9-14900KF): | |
| | tokens | CUDA | CPU, 32 threads | WebGPU (Chrome) | WebAssembly (Chrome) | | |
| |---|---|---|---|---| | |
| | 2,048 | 0.05 s | 1.3 s | 0.17-0.33 s | 3.0 s | | |
| | 8,192 | 0.3 s | 5.9 s | 1.5 s | 15 s | | |
| ## Limitations | |
| - It sees one window and nothing else: no callers, no configuration, no data flow across | |
| files. Questions such as authorization (`cwe_862`) or CSRF (`cwe_352`) often hinge on | |
| context outside the window. | |
| - Most training positives come from public CVE fixes, so the scores lean toward flaws that | |
| were found and fixed in open-source projects. | |
| - Expect false positives on ordinary code at roughly the rate in the table. Treat the | |
| output as a ranking for human review. | |
| ## License | |
| This model is a derivative of LiquidAI/LFM2.5-Encoder-350M and is distributed under the | |
| [LFM Open License v1.0](LICENSE). Commercial use is licensed only to organisations below | |
| US$10M in annual revenue; see section 5 of the license. | |
| Changes from the base model: the encoder weights were fine-tuned, a per-question attention | |
| pool and a 25-way linear head were added, and the result was exported to ONNX, including an | |
| 8-bit quantised copy. | |