1.72 MB
25 files
Updated 22 days ago
README.md

Open-MalSec

Open-MalSec is an open cybersecurity dataset built for practical defensive security, ML and security-agent research.

The current corpus contains 20 independent subsets and 1,104 synthetic/curated security examples covering phishing, malware, ransomware, scams, cloud identity, API security, software supply chain, GenAI security, mobile, IoT/OT and more.

DOI: 10.57967/hf/5104

What's in it

Subset Rows Focus
common-malware-vectors 51 common malware delivery and execution vectors
ai-agents-social-phishing 50 AI-assisted social phishing and impersonation
ai-llm-security-threats 50 LLM, RAG and agent security risks
botnet-ddos-misc 61 botnets, DDoS and network disruption
ceo-hr-phish-invoice-scam 52 BEC, payroll, HR and invoice fraud
cloud-saas-identity-attacks 50 cloud accounts, MFA, OAuth and SaaS identity
defi-meme-crypto-token-scams 60 DeFi, wallet and token scams
discord-social-engineering 60 Discord phishing and community scams
facebook-romance-scams 60 romance, identity and advance-fee scams
iot-ot-ics-threats 50 IoT, OT and industrial control threats
living-off-the-land-lotl 50 native-tool and signed-binary abuse
mobile-threats-detection 60 Android/iOS malicious and suspicious behaviour
onlyfans-subscription-fake-scams 60 subscription, creator and payment scams
phishing-email-inbound 60 inbound email phishing
ransomware-as-a-service-raas 60 RaaS ecosystem and extortion
ransomware-cases 60 ransomware incident scenarios
runescape-wow-diablo-mmo-scams 60 gaming/MMO phishing and trade scams
secrets-tokens-credential-exposure 50 secrets, tokens, sessions and credential exposure
software-supply-chain-devsecops 50 package, CI/CD and release supply-chain threats
web-api-attacks 50 API security weaknesses and abuse

Load it

from datasets import load_dataset

dataset = load_dataset(
    "tegridydev/open-malsec",
    "cloud-saas-identity-attacks",
)

The default config is common-malware-vectors.

List every subset:

from datasets import get_dataset_config_names

print(get_dataset_config_names("tegridydev/open-malsec"))

Structure

The refreshed and newly added subsets use a structured record shape:

{
  "id": "phishing-email-inbound-001",
  "instruction": "Analyse this email security scenario...",
  "input": {
    "source": "email",
    "actor": "synthetic-password-expiry-1",
    "subject": "Password expires today",
    "message": "Your account expires today...",
    "artifacts": [],
    "context": {
      "case_reference": "INC-2401",
      "organisation": "Northbridge Labs",
      "subject_user": "Alex Morgan",
      "device": "WS-021"
    }
  },
  "output": {
    "classification": "malicious",
    "threat_type": "credential_phishing",
    "summary": "Password-expiry phishing...",
    "indicators": [],
    "attack_techniques": [],
    "recommended_actions": []
  },
  "metadata": {
    "provenance": "synthetic_curated",
    "difficulty": "easy",
    "platforms": ["cross-platform"],
    "tags": ["cybersecurity", "open-malsec"],
    "frameworks": [],
    "synthetic": true,
    "record_family": "password-expiry",
    "variant": 1
  }
}

Each subset is isolated as its own Hugging Face config so schema differences do not affect other subsets.

Dataset design

Focuses on a few simple rules:

  • stable namespaced record IDs
  • valid JSON with one independently loadable config per subset
  • defensive analysis and recommended actions rather than offensive instructions
  • malicious, suspicious and benign controls
  • OWASP framework mappings for API and GenAI-specific subsets

The data is scenario based. It is useful for experiments, classifiers, evaluation and instruction tuning, but it is not a live threat-intelligence feed or a substitute for current incident data.

Uses

Good fits include:

  • malicious/security text classification
  • phishing and social-engineering detection
  • security-agent evaluation
  • structured-output experiments
  • defensive security fine-tuning
  • threat and indicator extraction
  • dataset/model benchmarking
  • cybersecurity education

Models trained on Open-MalSec should still be tested against separate, current and task specific data before realworld use.

Frameworks

Some records include mappings to external security frameworks where useful:

  • MITRE ATT&CK Enterprise / Mobile / ICS
  • OWASP API Security Top 10 2023
  • OWASP GenAI LLM Top 10 2026

Framework IDs are annotations for research and analysis. They should not be treated as exhaustive mappings for every possible interpretation of a scenario.

Research

Open-MalSec has been used as an experimental cybersecurity corpus in peer-reviewed research:

Toward Equitable Arabic Cybersecurity Literacy: A Rubric-Constrained LLM Framework for Phishing Detection and Bilingual Translation Fidelity

Ghazal et al., Mathematical and Computational Applications, 31(5), 168, 2026.

The paper evaluates its SECURE-A²RC framework using three cybersecurity corpora, including Open MalSec.

Citation

@dataset{tegridydev_open_malsec_2025,
  title     = {Open-MalSec},
  author    = {{TegridyDev}},
  year      = {2025},
  publisher = {Hugging Face},
  doi       = {10.57967/hf/5104},
  url       = {https://huggingface.co/datasets/tegridydev/open-malsec}
}

Research using Open-MalSec:

@article{ghazal2026equitable,
  author  = {Ghazal, Taher M. and Anwar, Fareeha and Al-Ghuribi, Sumaia Mohammed and Ahmed, Amjed A. and Najim, Ali Hamzah and Almomani, Omar and Pachiyannan, Prabu and Sakr, Hesham A.},
  title   = {Toward Equitable Arabic Cybersecurity Literacy: A Rubric-Constrained LLM Framework for Phishing Detection and Bilingual Translation Fidelity},
  journal = {Mathematical and Computational Applications},
  year    = {2026},
  volume  = {31},
  number  = {5},
  pages   = {168},
  doi     = {10.3390/mca31050168}
}

Licence

MIT.

Links

Total size
1.72 MB
Files
25
Last updated
Sep 12
Pre-warmed CDN
US EU US EU

Contributors