File size: 4,885 Bytes
adb5082 2f5acbd adb5082 2f5acbd adb5082 2f5acbd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 | ---
title: Malware Detection Pipeline
emoji: π¬
colorFrom: red
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
---
# Two-Stage Malware Detection Pipeline
### MLP + Vision Transformer
**Student:** Amine Garaali | **Deep Learning Project** | **2026**
---
## How it works
This demo combines two deep learning models into a sequential detection pipeline β the same architecture used in real-world security products.
```
Upload .exe or .dll
β
βββββββββββββββββββββββββββββββββββββββββββ
β Stage 1 β MLP (EMBER features) β
β Extracts 2381 static PE features β
β Outputs: malware probability β
βββββββββββββββββββββββββββββββββββββββββββ
β
probability < 0.5 β β
BENIGN β stop
probability β₯ 0.5 β π΄ MALWARE β Stage 2
β
βββββββββββββββββββββββββββββββββββββββββββ
β Stage 2 β Vision Transformer (ViT) β
β Converts binary to grayscale image β
β Classifies into 1 of 25 malware β
β families using self-attention β
βββββββββββββββββββββββββββββββββββββββββββ
β
Family name + confidence score
```
No file is executed at any point. Both models perform static analysis only.
---
## Stage 1 β MLP Malware Detector
- **Architecture:** 4-layer MLP (2381 β 512 β 256 β 128 β 1)
- **Input:** 2381-dimensional EMBER feature vector (imports, byte histogram, sections, strings, header info)
- **Task:** Binary classification β malware vs benign
- **Dataset:** EMBER 2018 β 800,000 labeled Windows PE files
- **Performance:** 95% accuracy, ROC-AUC 0.9878
- **Key components:** BatchNorm, ReLU, Dropout(0.3), Sigmoid output, Adam optimizer
The MLP acts as a fast, high-accuracy first filter. If the file scores below 50% malware probability it is immediately classified as benign and Stage 2 is skipped.
---
## Stage 2 β Vision Transformer (ViT) Family Classifier
- **Architecture:** ViT built from scratch β PatchEmbedding + 6 TransformerBlocks + classification head
- **Input:** 64Γ64 grayscale image derived from raw file bytes (each byte = one pixel)
- **Task:** 25-class malware family identification
- **Dataset:** MalImg β 9,339 grayscale malware images
- **Performance:** 98% accuracy across 25 families
- **Key components:** Multi-head self-attention (8 heads), GELU activation, CLS token, positional embeddings, AdamW optimizer
The ViT processes the file's binary content as a sequence of 64 image patches, computing attention between all patch pairs simultaneously. This allows it to detect structural relationships across the entire file that convolutional networks cannot capture in a single layer.
---
## The 25 malware families
```
Adialer.C Agent.FYI Allaple.A Allaple.L Alueron.gen!J
Autorun.K C2LOP.P C2LOP.gen!g Dialplatform.B Dontovo.A
Fakerean Instantaccess Lolyda.AA1 Lolyda.AA2 Lolyda.AA3
Lolyda.AT Malex.gen!J Obfuscator.AD Rbot!gen Skintrim.N
Swizzor.gen!E Swizzor.gen!I VB.AT Wintrim.BX Yuner.A
```
---
## Why two stages?
| | Stage 1 (MLP) | Stage 2 (ViT) |
|---|---|---|
| Question answered | Is this malware? | What kind of malware? |
| Input | 2381 engineered features | Raw binary as image |
| Feature learning | Human-designed | Learned by model |
| Speed | Very fast | Slower |
| Can detect benign? | Yes | No |
The MLP is fast and handles the binary decision efficiently. The ViT only runs when needed, providing deeper analysis for confirmed malware. This mirrors how production security tools work β cheap detection first, expensive classification second.
---
## Limitations
- Stage 1 was trained on PE files from 2018 β very recent malware using novel evasion may score lower than expected
- Stage 2 recognizes only the 25 MalImg families β novel malware will be classified into the closest-looking known family
- Neither model executes the file β dynamic behavior (e.g. fileless malware) is not analyzed
- For research and educational purposes only
---
## References
- Anderson & Roth (2018). *EMBER: An Open Dataset for Training Static PE Malware Machine Learning Models.* arXiv:1804.04637
- Nataraj et al. (2011). *Malware Images: Visualization and Automatic Classification.* VizSec 2011
- Dosovitskiy et al. (2020). *An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.* arXiv:2010.11929
- Khan & Nauman (2024). *Interpretable Detection of Malicious Behavior in Windows PE using Multi-Head 2D Transformers.* Big Data Mining and Analytics
|