Feature Extraction
sentence-transformers
PyTorch
English
Vietnamese
xlm-roberta
embedding
dense-retrieval
contrastive-learning
cve
cybersecurity
qdrant
secAI
text-embeddings-inference
Instructions to use DuyTa/sec-embedding with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use DuyTa/sec-embedding with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("DuyTa/sec-embedding") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
|
Download README.md from DuyTa/sec-embedding: direct link, hf CLI and curl.
- Browser
- Download file 3.68 kB
-
https://huggingface.co/DuyTa/sec-embedding/resolve/main/README.md
- Command line
-
hf download hf://DuyTa/sec-embedding/README.md
-
curl -L -o README.md https://huggingface.co/DuyTa/sec-embedding/resolve/main/README.md
3.68 kB
| license: apache-2.0 | |
| base_model: BAAI/bge-m3 | |
| base_model_relation: finetune | |
| library_name: sentence-transformers | |
| pipeline_tag: feature-extraction | |
| language: | |
| - en | |
| - vi | |
| pretty_name: sec-embedding (fine-tuned BGE-M3) | |
| tags: | |
| - sentence-transformers | |
| - feature-extraction | |
| - embedding | |
| - dense-retrieval | |
| - contrastive-learning | |
| - cve | |
| - cybersecurity | |
| - qdrant | |
| - secAI | |
| datasets: | |
| - DuyTa/Cyber_F1_v2 | |
| - DuyTa/cve-kgrag-db | |
| # sec-embedding | |
| **This is a fine-tuned version of [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3)** for CVE / cybersecurity dense retrieval. | |
| It was trained on a **CVE investigation-trajectory dataset** with **hard-negative mining** from a **local Qdrant** collection (`cve_kb`, NVD/MITRE core chunks). It is not a raw copy of the base checkpoint. It is the retriever component of the **secAI** stack, paired with [`DuyTa/sec-rerank`](https://huggingface.co/DuyTa/sec-rerank) and [`DuyTa/Cyber-F1-AWQ`](https://huggingface.co/DuyTa/Cyber-F1-AWQ). | |
| ## Training | |
| From `notebooks/BGE_M3_Colab.ipynb`: | |
| | | | | |
| |---|---| | |
| | Base | `BAAI/bge-m3` via Unsloth `FastSentenceTransformer` (`unsloth/bge-m3`) | | |
| | Role | Bi-encoder / **dense retriever** (1024-d, same geometry as bge-m3) | | |
| | Adapter | LoRA, `r=32`, modules `key`, `query`, `value`, `dense` | | |
| | Loss | `CachedMultipleNegativesRankingLoss` (InfoNCE, in-batch hard negatives) | | |
| | Engine | `sentence-transformers` `SentenceTransformerTrainer` | | |
| | Max sequence length | 1024 | | |
| | Learning rate | 2e-5, bf16 | | |
| Each example is a `(query, positive)` pair: | |
| - **Query** — CVE investigation trajectory (Vietnamese or English) over CVE-ID, CWE, product, severity, year, CAPEC / ATT&CK, filled from real KB metadata. | |
| - **Positive** — matching CVE passage from local Qdrant `cve_kb`. | |
| - **Hard negatives** — other CVE documents in the same mini-batch, all mined from that Qdrant index (near-miss CVEs: similar wording, wrong ID). | |
| Dataset source field: `Qdrant cve_kb (NVD/MITRE)`. Split: 40k train / 5k validation. | |
| ### Training corpus | |
| Built from **five years of authoritative cybersecurity sources**: NVD (173,473 CVEs), MITRE CWE (768 weakness types, mapped to ~92% of CVEs), CAPEC/ATT&CK (443/174 entries) and Exploit-DB (3,139 exploits, 2021–2026). Public datasets: [`DuyTa/Cyber_F1_v2`](https://huggingface.co/datasets/DuyTa/Cyber_F1_v2), [`DuyTa/cve-kgrag-db`](https://huggingface.co/datasets/DuyTa/cve-kgrag-db). | |
| Training hardware: 2×A100 80GB. | |
| ## Acceptance (nghiệm thu) — reported KPIs | |
| Measured on **NVIDIA A100 80GB** in the full production chatflow (Hybrid Search → Rerank → LLM), on a 1,000-sample security test set (40% CVE identification/classification, 40% remediation advice, 20% real-world scenario reasoning): | |
| | Metric | Result | Target | Pass | | |
| |---|---|---|---| | |
| | Retrieval quality — **Hit Rate@10** | **98.78%** | > 96% | ✅ | | |
| | Throughput | **2,662 emb/s** (concurrency 32) | ≥ 1,200 emb/s | ✅ | | |
| Raw per-sample logs (`embedding-hit-rate-at-10.jsonl`) and evaluation code are delivered with the acceptance package. | |
| ## Usage | |
| ```python | |
| from sentence_transformers import SentenceTransformer | |
| model = SentenceTransformer("DuyTa/sec-embedding") | |
| query_emb = model.encode("CVE-2021-44228 impact on log4j", normalize_embeddings=True) | |
| doc_emb = model.encode(passage, normalize_embeddings=True) | |
| ``` | |
| Rebuild the Qdrant index with **this** checkpoint. Mixing vectors with vanilla `BAAI/bge-m3` drops recall. | |
| ## Attribution & license | |
| Released under **Apache-2.0**. Derived from [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3) (**MIT License**); the MIT notice of the base model is retained and credit for the base weights belongs to the BAAI authors. | |