Buckets:
Open-MalSec
Open-MalSec is an open cybersecurity dataset built for practical defensive security, ML and security-agent research.
The current corpus contains 20 independent subsets and 1,104 synthetic/curated security examples covering phishing, malware, ransomware, scams, cloud identity, API security, software supply chain, GenAI security, mobile, IoT/OT and more.
DOI: 10.57967/hf/5104
What's in it
| Subset | Rows | Focus |
|---|---|---|
common-malware-vectors |
51 | common malware delivery and execution vectors |
ai-agents-social-phishing |
50 | AI-assisted social phishing and impersonation |
ai-llm-security-threats |
50 | LLM, RAG and agent security risks |
botnet-ddos-misc |
61 | botnets, DDoS and network disruption |
ceo-hr-phish-invoice-scam |
52 | BEC, payroll, HR and invoice fraud |
cloud-saas-identity-attacks |
50 | cloud accounts, MFA, OAuth and SaaS identity |
defi-meme-crypto-token-scams |
60 | DeFi, wallet and token scams |
discord-social-engineering |
60 | Discord phishing and community scams |
facebook-romance-scams |
60 | romance, identity and advance-fee scams |
iot-ot-ics-threats |
50 | IoT, OT and industrial control threats |
living-off-the-land-lotl |
50 | native-tool and signed-binary abuse |
mobile-threats-detection |
60 | Android/iOS malicious and suspicious behaviour |
onlyfans-subscription-fake-scams |
60 | subscription, creator and payment scams |
phishing-email-inbound |
60 | inbound email phishing |
ransomware-as-a-service-raas |
60 | RaaS ecosystem and extortion |
ransomware-cases |
60 | ransomware incident scenarios |
runescape-wow-diablo-mmo-scams |
60 | gaming/MMO phishing and trade scams |
secrets-tokens-credential-exposure |
50 | secrets, tokens, sessions and credential exposure |
software-supply-chain-devsecops |
50 | package, CI/CD and release supply-chain threats |
web-api-attacks |
50 | API security weaknesses and abuse |
Load it
from datasets import load_dataset
dataset = load_dataset(
"tegridydev/open-malsec",
"cloud-saas-identity-attacks",
)
The default config is common-malware-vectors.
List every subset:
from datasets import get_dataset_config_names
print(get_dataset_config_names("tegridydev/open-malsec"))
Structure
The refreshed and newly added subsets use a structured record shape:
{
"id": "phishing-email-inbound-001",
"instruction": "Analyse this email security scenario...",
"input": {
"source": "email",
"actor": "synthetic-password-expiry-1",
"subject": "Password expires today",
"message": "Your account expires today...",
"artifacts": [],
"context": {
"case_reference": "INC-2401",
"organisation": "Northbridge Labs",
"subject_user": "Alex Morgan",
"device": "WS-021"
}
},
"output": {
"classification": "malicious",
"threat_type": "credential_phishing",
"summary": "Password-expiry phishing...",
"indicators": [],
"attack_techniques": [],
"recommended_actions": []
},
"metadata": {
"provenance": "synthetic_curated",
"difficulty": "easy",
"platforms": ["cross-platform"],
"tags": ["cybersecurity", "open-malsec"],
"frameworks": [],
"synthetic": true,
"record_family": "password-expiry",
"variant": 1
}
}
Each subset is isolated as its own Hugging Face config so schema differences do not affect other subsets.
Dataset design
Focuses on a few simple rules:
- stable namespaced record IDs
- valid JSON with one independently loadable config per subset
- defensive analysis and recommended actions rather than offensive instructions
malicious,suspiciousandbenigncontrols- OWASP framework mappings for API and GenAI-specific subsets
The data is scenario based. It is useful for experiments, classifiers, evaluation and instruction tuning, but it is not a live threat-intelligence feed or a substitute for current incident data.
Uses
Good fits include:
- malicious/security text classification
- phishing and social-engineering detection
- security-agent evaluation
- structured-output experiments
- defensive security fine-tuning
- threat and indicator extraction
- dataset/model benchmarking
- cybersecurity education
Models trained on Open-MalSec should still be tested against separate, current and task specific data before realworld use.
Frameworks
Some records include mappings to external security frameworks where useful:
- MITRE ATT&CK Enterprise / Mobile / ICS
- OWASP API Security Top 10 2023
- OWASP GenAI LLM Top 10 2026
Framework IDs are annotations for research and analysis. They should not be treated as exhaustive mappings for every possible interpretation of a scenario.
Research
Open-MalSec has been used as an experimental cybersecurity corpus in peer-reviewed research:
Toward Equitable Arabic Cybersecurity Literacy: A Rubric-Constrained LLM Framework for Phishing Detection and Bilingual Translation Fidelity
Ghazal et al., Mathematical and Computational Applications, 31(5), 168, 2026.
The paper evaluates its SECURE-A²RC framework using three cybersecurity corpora, including Open MalSec.
Citation
@dataset{tegridydev_open_malsec_2025,
title = {Open-MalSec},
author = {{TegridyDev}},
year = {2025},
publisher = {Hugging Face},
doi = {10.57967/hf/5104},
url = {https://huggingface.co/datasets/tegridydev/open-malsec}
}
Research using Open-MalSec:
@article{ghazal2026equitable,
author = {Ghazal, Taher M. and Anwar, Fareeha and Al-Ghuribi, Sumaia Mohammed and Ahmed, Amjed A. and Najim, Ali Hamzah and Almomani, Omar and Pachiyannan, Prabu and Sakr, Hesham A.},
title = {Toward Equitable Arabic Cybersecurity Literacy: A Rubric-Constrained LLM Framework for Phishing Detection and Bilingual Translation Fidelity},
journal = {Mathematical and Computational Applications},
year = {2026},
volume = {31},
number = {5},
pages = {168},
doi = {10.3390/mca31050168}
}
Licence
MIT.
Links
- Hugging Face: https://huggingface.co/tegridydev
- Dataset: https://huggingface.co/datasets/tegridydev/open-malsec
- GitHub: https://github.com/tegridydev/open-malsec
- Total size
- 1.72 MB
- Files
- 25
- Last updated
- Sep 12
- Pre-warmed CDN
- US EU US EU