CVE → CVSS v3.1 vector predictor (ModernBERT-base)

Given an English vulnerability description, this model predicts all eight CVSS v3.1 base metrics (AV, AC, PR, UI, S, C, I, A) and, from them, the base score and severity. It returns per-metric probabilities, a severity distribution and a list of uncertain metrics, so a calling system (or an agent) knows when to trust the estimate and when to escalate.

  • Architecture: fine-tuned ModernBERT-base encoder (≈150M parameters), masked mean pooling, one linear head per CVSS metric.
  • Training data: ≈107,000 CVE Records from the CVE List published before July 2025.
  • Evaluation: 57,611 CVEs published January–October 2026, all after the training period (temporal split).
  • Output: always a valid CVSS v3.1 base vector, scored with the official v3.1 formula.
  • Use it for: fast triage and a first estimate. It is not an authoritative score: CVSS is assigned by the CNA, NVD or your own analysts.

Built by AXONVERTEX AI Research. The pipeline-integration guide is in PIPELINE_INTEGRATION.md.

Results at a glance

Test set: CVEs published 2026-01-01 to early October 2026, excluding Oracle (n = 55,131; see Evaluation for why). 95% bootstrap intervals in brackets.

What Result
Exact match, all 8 metrics 31.3% [30.9, 31.7]
Base-score mean absolute error (expected score) 1.02 [1.02, 1.03]
Base score within ±1.0 59.0%
Severity accuracy (most likely severity) 63.2% [62.8, 63.6]
Mean per-metric accuracy (8 metrics) 82.6%

For comparison on the same rows: a model that knows only which CNA published the CVE gets 14.3% exact match; TF-IDF + logistic regression gets 29.9%.

Quick start

pip install -U torch transformers huggingface_hub safetensors numpy
import os, sys
from huggingface_hub import hf_hub_download

REPO = "AXONVERTEX-AI-RESEARCH/cve-cvss-vector-modernbert-base"
sys.path.insert(0, os.path.dirname(hf_hub_download(REPO, "cvss_predictor.py")))
from cvss_predictor import CVSSPredictor

predictor = CVSSPredictor(REPO)          # CPU or GPU, picked automatically
r = predictor.predict(
    "SQL injection in the login form of Acme Portal 2.3 allows unauthenticated remote attackers "
    "to execute arbitrary SQL commands via the username parameter."
)
print(r["vector"], r["base_score"], r["most_likely_severity"], r["uncertain_metrics"])

Command line: python cvss_predictor.py "description one" "description two". HTTP API: uvicorn serve:app --port 8000, then POST /predict with {"descriptions": [...]} (see serve.py).

Example predictions from the trained model:

Description (abridged) Predicted vector Score Severity
SQL injection in a login form, unauthenticated remote attackers AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H 9.8 CRITICAL
Stored XSS in a comment editor, authenticated users AV:N/AC:L/PR:L/UI:R/S:C/C:L/I:L/A:N 5.4 MEDIUM
Race condition in a device driver, local low-privileged user, use-after-free AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H 7.0 HIGH

Output contract

predict(text) returns one JSON-serialisable dict per description:

Field Type Meaning
vector string Most probable valid CVSS v3.1 base vector, e.g. CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H. Vectors with no impact (C:N/I:N/A:N) are never returned.
base_score float CVSS v3.1 base score of vector.
severity string Severity of vector: LOW, MEDIUM, HIGH or CRITICAL.
vector_probability float Model probability of that exact vector among all 2,592 valid vectors. Often low (many vectors are close); compare across findings rather than reading it as a percentage chance.
most_likely_severity string Severity with the highest total probability across all vectors. Use this for triage queues; it is more accurate than severity (63.2% vs 61.7%).
severity_probabilities dict Probability of LOW / MEDIUM / HIGH / CRITICAL.
expected_score float Probability-weighted base score; the lowest-error point estimate (MAE 1.02).
metrics dict Per metric: name, predicted value, value_name and probabilities over its values.
uncertain_metrics list Metrics whose top probability is below 0.7 (configurable). Candidates for review or for an LLM that can see the code.

score_vector(vector) scores a vector proposed elsewhere (e.g. by an LLM) with the same formula, for like-for-like comparison.

The probabilities come from a softmax per metric; their calibration has not been measured. The 0.7 threshold is a heuristic.

Use as a tool in an agent pipeline

Function-calling definition (works with OpenAI-style and most open-model tool formats):

{
  "name": "predict_cvss",
  "description": "Estimate the CVSS v3.1 base vector, base score and severity of a vulnerability from an English, CVE-style description. Returns per-metric probabilities and the metrics the model is unsure about. The result is an estimate for triage, not an official score.",
  "parameters": {
    "type": "object",
    "properties": {
      "descriptions": {
        "type": "array",
        "items": {"type": "string"},
        "description": "One description per vulnerability: affected component, vulnerability type, who can exploit it (remote/local, authentication needed), whether user interaction is needed, and the impact."
      }
    },
    "required": ["descriptions"]
  }
}

Rules for an agent using this tool:

  1. Describe the vulnerability the way a CVE description does. Do not include a CVSS score or vector in the text.
  2. Report most_likely_severity and expected_score for prioritisation, and vector for detail. Label them as model estimates.
  3. If uncertain_metrics is non-empty, decide those metrics from evidence (code, configuration, exploit preconditions), keep the confident ones, and rescore with score_vector.
  4. Never lower a finding's priority on this model's output alone when there is evidence of active exploitation.
  5. Send HIGH and CRITICAL findings to human review.

PIPELINE_INTEGRATION.md covers the full integration with a three-model secure-coding pipeline (Antares-1B, Foundation-Sec-8B-Reasoning, Gemma 12B): reference flow, prompts, escalation logic, deployment and how to evaluate it.

Training data and preprocessing

  • Source: CVE List, cvelistV5 repository snapshot of October 2026, CVE JSON 5.x records with state PUBLISHED.
  • Text: the CNA's English description, whitespace collapsed. Descriptions shorter than 30 characters were dropped.
  • Label: the CVSS v3.x base vector. The CNA's vector is used first, with an ADP vector (mainly CISA's vulnrichment container) as fallback. v3.1 is preferred over v3.0. Temporal and environmental metrics are ignored. NVD's own scores are not in the CVE List and were not used.
  • Leakage removal: some CNAs write the answer into the description, e.g. Oracle (CVSS 3.1 Base Score 4.9 ... CVSS Vector: (CVSS:3.1/...)), Concrete CMS and dotCMS. Scores, vectors and CVSS-calculator links are stripped from all descriptions before training, and the loader applies the same cleaning to inputs.
  • Deduplication: exact-duplicate descriptions removed; the earliest record is kept.
  • Temporal split by datePublished:
Split Published Size
Train before 2025-07-01 ≈107,000
Validation 2025-07-01 to 2025-12-31 used for checkpoint selection
Test 2026-01-01 to early October 2026 57,611

Test labels come from CNAs (46,441), CISA-ADP (10,962) and Red Hat (208).

Training procedure

Setting Value
Base model answerdotai/ModernBERT-base
Objective mean of 8 cross-entropy losses, square-root inverse-frequency class weights
Max length 512 tokens
Batch size / learning rate 64 / 5e-5, AdamW, weight decay 0.01, linear schedule with 6% warm-up
Epochs 4; checkpoint with the best validation mean macro-F1 kept (epoch 3)
Precision / hardware bf16 autocast, one NVIDIA A100; ≈2.5 minutes per epoch

Validation at the selected epoch: mean macro-F1 0.797, exact match 43.1%. Validation plateaued after the first epoch (0.789); results/training_history.csv has the full curve.

The training notebook is in training/CVE_finetune_A100.ipynb.

Evaluation

Why Oracle is excluded from the headline numbers

Oracle's descriptions are generated from the CVSS vector itself: "Easily exploitable" means AC:L, "high privileged attacker" means PR:H, "scope change" means S:C, "complete DOS" means A:H. The model reaches 99.0% exact match on Oracle CVEs even after the explicit vector strings are removed. That is template decoding, not vulnerability understanding, so headline figures exclude Oracle (2,450 test CVEs). Results with Oracle are shown alongside.

Decoding

Choosing each metric independently can produce a no-impact vector (score 0), which almost never occurs in real CVEs; the independent decoder did so 680 times on the test set. The released loader uses joint decoding over all 2,592 valid vectors instead.

Test subset Decoding Exact vector Score MAE Within ±1.0 Severity acc Severity macro-F1
Excl. Oracle (n = 55,131) per-metric argmax 31.2% 1.141 58.4% 61.2% 0.524
joint, impact > 0 (vector, severity) 31.3% 1.084 58.8% 61.7% 0.525
marginal (expected_score, most_likely_severity) — 1.024 59.0% 63.2% 0.520
All (n = 57,581) joint, impact > 0 34.2% 1.038 60.5% 63.4% 0.547
marginal — 0.984 60.8% 64.8% 0.545

Severity macro-F1 is over LOW/MEDIUM/HIGH/CRITICAL; the 30 test CVEs whose true vector has no impact are excluded from score and severity metrics.

Baselines (same test rows, excluding Oracle, same decoding)

Method Exact vector Mean metric macro-F1 Score MAE Severity acc Severity macro-F1
Most frequent training vector 5.8% 0.303 2.85 11.8% 0.053
CNA prior: the assigner's most frequent vector (no text) 14.3% 0.499 1.67 42.7% 0.297
TF-IDF (word 1–2-grams) + logistic regression 29.9% 0.721 1.12 59.6% 0.489
This model 31.3% 0.742 1.08 61.7% 0.525

The description text more than doubles exact match over knowing the assigner alone. A linear model captures most of the signal; the encoder's clearest gains are on the rarer classes (severity macro-F1 +3.7 points).

Per metric (all 57,611 test CVEs, per-metric argmax)

Metric Accuracy Macro-F1 Majority-class accuracy Weakest class
AV Attack Vector 90.9% 0.716 80.1% (N) Adjacent (F1 0.49)
AC Attack Complexity 86.0% 0.697 85.8% (L) High (recall 0.44)
PR Privileges Required 78.3% 0.715 55.3% (N) High (F1 0.56)
UI User Interaction 90.6% 0.865 76.7% (N) Required (recall 0.76)
S Scope 87.2% 0.787 79.9% (U) Changed (recall 0.60)
C Confidentiality 77.9% 0.759 51.0% (H) Low (recall 0.63)
I Integrity 77.6% 0.769 42.5% (H) Low (recall 0.68)
A Availability 77.7% 0.721 43.1% (H) Low (recall 0.45)

Attack Complexity barely beats always answering "Low": descriptions rarely state the conditions that make an attack complex. With per-metric decoding, CRITICAL is recovered 49% of the time (mostly predicted HIGH, usually one metric away) and LOW 33% of the time.

By CNA (15 largest in the test set, per-metric argmax)

CNA Test CVEs Exact vector Score MAE
GitHub_M 7,389 16.1% 1.49
VulnCheck 6,678 28.4% 1.14
VulDB 5,463 34.4% 1.07
Patchstack 4,114 30.1% 1.16
mitre 3,292 32.2% 1.29
Wordfence 3,241 67.4% 0.57
Linux 3,053 40.3% 0.71
Chrome 2,782 42.1% 1.01
oracle 2,450 99.0% 0.01
microsoft 1,743 45.7% 0.67
WPScan 1,510 31.0% 1.18
ibm 1,150 27.0% 1.28
redhat 1,062 16.6% 1.33
apache 828 26.8% 1.43
apple 755 22.4% 1.10

The model learns each CNA's scoring conventions as well as the vulnerability. Consistent scorers are predicted well, such as Wordfence and the Linux kernel. Advisories scored by many different maintainers (GitHub_M) are predicted poorly. CNA-scored and CISA-ADP-scored labels are equally predictable (34.3% vs 33.5% exact).

Limitations and appropriate use

  • An estimate, not a score. Exact match is 31%, and 59% of scores land within ±1.0. Use it to sort and pre-fill, not to publish.
  • Scoring conventions differ between CNAs, and human scorers disagree. Part of the remaining error is label inconsistency rather than model error.
  • It reads text, not code. Its quality depends on how well the description states attack vector, privileges, user interaction and impact. In a code-scanning pipeline, have an LLM write the finding in CVE style first (see the integration guide).
  • Weak spots: Attack Complexity, Scope, the Low value of C/I/A, and the LOW and CRITICAL severity bands.
  • English only; CVSS v3.1 only (not v4.0). Trained on CVEs published up to June 2025; vulnerability language drifts, so retrain periodically.
  • Probabilities are uncalibrated. Use them for ranking and for spotting uncertainty.
  • Not for dependency scanning. For "is my library version affected", use OSV.dev or the GitHub Advisory Database.

Files

File Contents
model.safetensors, config.json, tokenizer*.json Fine-tuned ModernBERT encoder and tokenizer (AutoModel.from_pretrained loads the encoder alone)
heads.safetensors The eight linear heads
cvss_config.json Label order per head, max length, pooling, decoding and training settings
cvss_predictor.py Loader, joint decoding, CVSS 3.1 calculator, CLI
serve.py FastAPI server: /predict, /score, /health
PIPELINE_INTEGRATION.md How to use the model in an agentic secure-coding pipeline
results/ Test metrics, decoding and baseline comparisons, confidence intervals, training history, per-CVE test predictions
training/CVE_finetune_A100.ipynb Data parsing, training and evaluation notebook
CVE_TERMS_OF_USE.md Notice for the CVE content used

License and data notice

Model weights and code: Apache 2.0. The training data and results/test_predictions.parquet come from the CVE List. CVE® is a registered trademark of The MITRE Corporation; CVE content is used under the CVE Terms of Use. See CVE_TERMS_OF_USE.md.

Citation

@misc{axonvertex2026cvssvector,
  title  = {CVE to CVSS v3.1 Vector Prediction with a Fine-Tuned ModernBERT Encoder},
  author = {Dasgupta, Krishnendu},
  year   = {2026},
  publisher = {AXONVERTEX AI Research},
  howpublished = {\url{https://huggingface.co/AXONVERTEX-AI-RESEARCH/cve-cvss-vector-modernbert-base}}
}
Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AXONVERTEX-AI-RESEARCH/cve-cvss-vector-modernbert-base

Finetuned
(1526)
this model