sec-embedding / README.md
DuyTa's picture
Model card: add acceptance (nghiệm thu) LLM/quality/task KPIs, training set size & hardware
f39e9e5 verified
|
Raw History Blame Contribute Delete
3.68 kB
---
license: apache-2.0
base_model: BAAI/bge-m3
base_model_relation: finetune
library_name: sentence-transformers
pipeline_tag: feature-extraction
language:
- en
- vi
pretty_name: sec-embedding (fine-tuned BGE-M3)
tags:
- sentence-transformers
- feature-extraction
- embedding
- dense-retrieval
- contrastive-learning
- cve
- cybersecurity
- qdrant
- secAI
datasets:
- DuyTa/Cyber_F1_v2
- DuyTa/cve-kgrag-db
---
# sec-embedding
**This is a fine-tuned version of [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3)** for CVE / cybersecurity dense retrieval.
It was trained on a **CVE investigation-trajectory dataset** with **hard-negative mining** from a **local Qdrant** collection (`cve_kb`, NVD/MITRE core chunks). It is not a raw copy of the base checkpoint. It is the retriever component of the **secAI** stack, paired with [`DuyTa/sec-rerank`](https://huggingface.co/DuyTa/sec-rerank) and [`DuyTa/Cyber-F1-AWQ`](https://huggingface.co/DuyTa/Cyber-F1-AWQ).
## Training
From `notebooks/BGE_M3_Colab.ipynb`:
| | |
|---|---|
| Base | `BAAI/bge-m3` via Unsloth `FastSentenceTransformer` (`unsloth/bge-m3`) |
| Role | Bi-encoder / **dense retriever** (1024-d, same geometry as bge-m3) |
| Adapter | LoRA, `r=32`, modules `key`, `query`, `value`, `dense` |
| Loss | `CachedMultipleNegativesRankingLoss` (InfoNCE, in-batch hard negatives) |
| Engine | `sentence-transformers` `SentenceTransformerTrainer` |
| Max sequence length | 1024 |
| Learning rate | 2e-5, bf16 |
Each example is a `(query, positive)` pair:
- **Query** — CVE investigation trajectory (Vietnamese or English) over CVE-ID, CWE, product, severity, year, CAPEC / ATT&CK, filled from real KB metadata.
- **Positive** — matching CVE passage from local Qdrant `cve_kb`.
- **Hard negatives** — other CVE documents in the same mini-batch, all mined from that Qdrant index (near-miss CVEs: similar wording, wrong ID).
Dataset source field: `Qdrant cve_kb (NVD/MITRE)`. Split: 40k train / 5k validation.
### Training corpus
Built from **five years of authoritative cybersecurity sources**: NVD (173,473 CVEs), MITRE CWE (768 weakness types, mapped to ~92% of CVEs), CAPEC/ATT&CK (443/174 entries) and Exploit-DB (3,139 exploits, 2021–2026). Public datasets: [`DuyTa/Cyber_F1_v2`](https://huggingface.co/datasets/DuyTa/Cyber_F1_v2), [`DuyTa/cve-kgrag-db`](https://huggingface.co/datasets/DuyTa/cve-kgrag-db).
Training hardware: 2×A100 80GB.
## Acceptance (nghiệm thu) — reported KPIs
Measured on **NVIDIA A100 80GB** in the full production chatflow (Hybrid Search → Rerank → LLM), on a 1,000-sample security test set (40% CVE identification/classification, 40% remediation advice, 20% real-world scenario reasoning):
| Metric | Result | Target | Pass |
|---|---|---|---|
| Retrieval quality — **Hit Rate@10** | **98.78%** | > 96% | ✅ |
| Throughput | **2,662 emb/s** (concurrency 32) | ≥ 1,200 emb/s | ✅ |
Raw per-sample logs (`embedding-hit-rate-at-10.jsonl`) and evaluation code are delivered with the acceptance package.
## Usage
```python
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("DuyTa/sec-embedding")
query_emb = model.encode("CVE-2021-44228 impact on log4j", normalize_embeddings=True)
doc_emb = model.encode(passage, normalize_embeddings=True)
```
Rebuild the Qdrant index with **this** checkpoint. Mixing vectors with vanilla `BAAI/bge-m3` drops recall.
## Attribution & license
Released under **Apache-2.0**. Derived from [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3) (**MIT License**); the MIT notice of the base model is retained and credit for the base weights belongs to the BAAI authors.