Text Classification
Safetensors
PyTorch
English
phishbyte
phishing-detection
email-security
cybersecurity
security
from-scratch
no-pretrained-weights
cascading-inference
lightweight
explainable-ai
nlp
phishing
spam-detection
malware-detection
threat-detection
email-classification
feature-engineering
interpretable-ml
tfidf
residual-network
context-gating
calibrated-probabilities
dmarc-alignment
Eval Results (legacy)
File size: 17,909 Bytes
7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b e3d8566 7bb380b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 | ---
language:
- en
license: mit
library_name: phishbyte
pipeline_tag: text-classification
tags:
- phishing-detection
- email-security
- cybersecurity
- security
- pytorch
- from-scratch
- no-pretrained-weights
- cascading-inference
- lightweight
- explainable-ai
- nlp
- phishing
- spam-detection
- malware-detection
- threat-detection
- email-classification
- text-classification
- feature-engineering
- interpretable-ml
- tfidf
- residual-network
- context-gating
- calibrated-probabilities
- dmarc-alignment
datasets:
- ceas-2008
- enron-email
- spamassassin
- ling-spam
- nazario-phishing
- nigerian-fraud
metrics:
- f1
- precision
- recall
- accuracy
model-index:
- name: phishbyte
results:
- task:
type: text-classification
name: Phishing Email Detection
dataset:
name: 7-source benchmark (CEAS, Enron, SpamAssassin, Ling-Spam, Nazario, Nigerian, farshad72, puyang2025)
type: ceas-2008
metrics:
- type: f1
value: 0.9445
name: F1 Score
- type: accuracy
value: 0.9477
name: Accuracy
- type: precision
value: 0.9402
name: Precision
- type: recall
value: 0.9489
name: Recall
widget:
- text: "From: PayPal Security <security@paypa1-alert.tk>\nReply-To: attacker@evil-domain.ru\nSubject: URGENT: Your account will be suspended\n\nDear Customer, your PayPal account has been suspended. Verify now at http://paypal-login.tk/verify"
example_title: "Phishing email example"
- text: "From: alice@company.com\nReply-To: alice@company.com\nSubject: Team lunch tomorrow\n\nHi everyone, lunch is at noon in the usual spot. See you there!"
example_title: "Legitimate email example"
---
# Phish_Byte v10
A from-scratch PyTorch model for **email phishing detection** โ no pretrained language model, no transformer, no fine-tuning.
**743,571 parameters.** Every signal is a feature computed directly from the email itself, fed through a **context-gating architecture** that learns how one piece of evidence should change the interpretation of another โ rather than a fixed hand-written rule deciding that for it.
> The only non-transformer phishing detection model on HuggingFace.
---
## Table of contents
- [Install](#install--no-pypi-package-yet)
- [Usage](#usage)
- [How it works, stage by stage](#how-it-works-stage-by-stage)
- [Architecture](#architecture)
- [Version history](#version-history)
- [Feature groups](#feature-groups)
- [Training data](#training-data)
- [Benchmarks](#benchmarks)
- [Limitations](#limitations)
---
## Install โ no PyPI package yet
`pip install phishbyte` does **not** work yet. The only supported path is cloning the source repository. Five steps, in order:
**Step 1 โ Clone the repository**
```bash
git clone https://github.com/AnonymousSingh-007/Phish_Byte.git
cd Phish_Byte
```
**Step 2 โ Create a virtual environment**
```bash
python -m venv venv
```
**Step 3 โ Activate it**
```bash
# Windows (PowerShell)
.\venv\Scripts\Activate.ps1
# Mac / Linux
source venv/bin/activate
```
**Step 4 โ Install dependencies**
```bash
pip install -r requirements.txt
```
Minimal set: `torch`, `huggingface_hub`, `safetensors`, `dnspython`, `numpy`, `pandas`.
For GPU acceleration on RTX 50-series (Blackwell) cards:
```bash
pip install torch --index-url https://download.pytorch.org/whl/cu128
```
**Step 5 โ Verify the install**
```bash
python verify_install.py
```
This checks every Python package and every source file is present, then does a live test download of the model weights from this Hub repo. If anything is missing, it tells you exactly what โ not a confusing traceback. Expected output when everything is correct:
```
โ
Python 3.11.x
โ
torch
โ
huggingface_hub
โ
safetensors
โ
dns
โ
numpy
โ
pandas
โ
phishbyte/__init__.py
... (all source files)
โ
from phishbyte import PhishByteEngine โ works
โ
Model loaded from Hub successfully
โ
INSTALLATION VERIFIED
```
---
## Usage
### Basic usage โ analyze any raw email
```python
from phishbyte import PhishByteEngine
# First call downloads ~3 MB (weights + thresholds + vocabulary) from this
# Hub repo and caches it locally. Every call after is instant.
engine = PhishByteEngine.from_pretrained("SamSec007/phishbyte")
verdict = engine.analyze(raw_email_string)
print(verdict.label) # "phishing" or "legitimate"
print(verdict.probability) # calibrated confidence, 0.0 to 1.0
print(verdict.confidence) # "high" / "medium" / "low"
print(verdict.layer_used) # 1 = a fast rule made the call, 2 = the full network did
print(verdict.feature_weights) # every signal computed for this specific email
```
### Analyze a real email from your own Gmail
**Step 1.** Open the suspicious email in Gmail.
**Step 2.** Click the **โฎ** menu in the top right of the email, then click **Show original**. This opens a new tab with the complete raw email, including every header.
**Step 3.** Select all the text (Ctrl+A) and copy it (Ctrl+C).
**Step 4.** Run the CLI:
```bash
python cli.py
```
**Step 5.** Paste the email when prompted, then press Enter followed by Ctrl+Z on Windows (or Ctrl+D on Mac/Linux) to submit it.
### Analyze a saved `.eml` file
```bash
python cli.py --file suspicious.eml
```
### Quick demo โ no files needed
```bash
python cli.py --demo phish # a representative phishing example
python cli.py --demo legit # a representative legitimate example
```
### Reading the verdict object
```python
PhishVerdict(
label = "phishing",
probability = 0.9735,
confidence = "high",
layer_used = 2,
feature_weights = {
"display_name_mismatch": 1.00, # "PayPal" in display name, unrelated domain
"mcld_mismatch": 1.00, # most-linked domain isn't the sender's
"auth_alignment_score": 0.95, # authentication does NOT validate this sender
"coercive_urgency_score": 0.82, # pressure language: "verify now", "suspended"
"professional_formality_score": 0.02, # essentially none โ this isn't formal writing
...
},
detail = "MLP probability (calibrated): 97.35%. Trust consistency: 0.10. Auth alignment: 0.95.",
)
```
---
## How it works, stage by stage
**Stage 1 โ Parsing.** The raw email string is split into its headers (`From`, `Reply-To`, `Return-Path`, `Subject`, `Authentication-Results`) and its body, using Python's standard email parser. This handles both plain-text and HTML/multipart emails.
**Stage 2 โ Domain analysis.** Checks whether the From, Reply-To, and Return-Path addresses are consistent with each other; whether the display name claims a known brand (like "PayPal Security") while the actual domain is unrelated; and whether the domain itself looks auto-generated, based on digit density, hyphen count, and length.
**Stage 3 โ URL and body analysis.** Extracts every link in the email, checks whether visible link text matches where the link actually points, measures how many distinct destination domains the links spread across, and looks at structural characteristics of the body like unusual capitalization density.
**Stage 4 โ Authentication validation (SPF, DKIM, DMARC).** Reads the `Authentication-Results` header that the receiving mail server already computed, extracting whether SPF passed, whether DKIM signed the message with a domain that matches the sender, and what DMARC โ the policy that ties SPF and DKIM together โ concluded. This is a live, structural check, not a keyword guess.
**Stage 5 โ Subject line analysis.** The same kind of pattern-matching as the body, scoped to the subject: brand names, currency symbols, ALL-CAPS shouting, fake "RE:" prefixes designed to look like an ongoing conversation.
**Stage 6 โ Link and form forensics.** Finds the single most common destination domain across every link in the email and compares it to the sender. Checks whether any form on the page submits directly to a raw IP address instead of a domain โ legitimate sites essentially never do this. Detects "open redirect" URL patterns commonly used to disguise a final destination.
**Stage 7 โ Lexical domain analysis.** A character-by-character look at domain names: does the domain have an unusual run of digits? Does it read like a real word or a randomly generated string? Is it a near-miss spelling of a known brand, once common digit-for-letter substitutions are normalized (`micros0ft` โ `microsoft`)?
**Stage 8 โ Cross-signal agreement check.** Looks at whether the independent modules above agree with each other. Three modules independently raising concern is much stronger evidence than one module alone.
**Stage 9 โ Context feature computation.** This is where v10 diverges most from earlier versions. Urgency language in the body is split into two independent numbers โ how *coercive* it is ("verify immediately or your account will be suspended") versus how *professionally formal* it is ("we kindly ask for your commitment to this important task") โ because a single blended urgency score cannot tell these apart, and they mean very different things. A separate signal captures how strongly DMARC and DKIM validate the sender, independent of what the link-forensics stage found โ so the network can weigh "the links go somewhere else" differently depending on whether the sender proved its identity or not. A similar context signal exists for Reply-To addresses that use free email providers like Gmail.
**Stage 10 โ Fusion.** All the evidence from stages 2 through 9 is split into two groups โ *raw evidence* (facts about this email) and *context evidence* (facts that should change how the raw evidence is read) โ and handed to a small neural layer whose only job is learning a gate: for this particular email, how much should the context evidence turn up or down the weight given to each piece of raw evidence. This gate is learned from real data, not hardcoded.
**Stage 11 โ Decision.** The fused representation, plus 50 word-frequency signals learned from the training corpus, feeds a residual neural network. Its output passes through a learned temperature parameter before being converted into a final probability, so that "80% confident" is empirically close to being right 80% of the time.
---
## Architecture
```
raw email
โ
โผ
parsing โ domain / URL / auth (SPF+DKIM+DMARC) / subject / link-forensics / lexical
โ
โผ
cross-signal agreement check
โ
โผ
context feature computation
(coercive vs. professional urgency, auth-verified sender context,
freemail Reply-To context, observational tracking-pattern signal)
โ
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโ
โผ โผ โ
RAW evidence CONTEXT evidence โ
(38 numbers) (13 numbers) โ
โ โ โ
โโโโโโโโฌโโโโโโโโ โ
โผ โ
Context Fusion Layer โ
(learns a gate: how much โ
should context reweight โ
each raw signal, per email) โ
โ โ
โผ โ
fused representation (64) โโโโโโโ
โ TF-IDF (50)
โโโโโโฌโโโโโโโโโ
โผ
residual MLP: 620 โ 310 (ร2 residual blocks) โ 155 โ 76 โ 1
โ
โผ
temperature-calibrated confidence score
โ
โผ
PhishVerdict โ label, confidence, and every signal that fired
```
**743,571 parameters total.** For comparison, DistilBERT-based phishing detectors on HuggingFace use 66,000,000+ parameters โ roughly 90ร more.
---
## Version history
**v2 โ 12,545 parameters, 29 features, single dataset (CEAS-2008, ~39K emails).** The original prototype. A small MLP over hand-picked domain, URL, and subject features, with a cascading Layer 1 (cheap rules) โ Layer 2 (neural network) design that every later version kept.
**v7 โ 254K parameters, 85 features, 83K emails across 6 datasets.** Added a TF-IDF vocabulary learned directly from the training corpus, and Body Domain Identification โ checking the most common link destination against the sender. First version tested for generalization beyond a single dataset.
**v8 โ 716K parameters, 104 features, 166K emails across 7 datasets.** Added character-level lexical domain analysis (digit runs, entropy, typosquat distance) and a cross-signal layer that checks whether independent modules agree with each other rather than scoring each one in isolation.
**v9 โ 718K parameters, 107 features.** Replaced a naive SPF-only authentication check โ which misfired constantly on legitimate marketing platforms like Marketo and Mailgun that relay mail on a brand's behalf โ with a proper DMARC/DKIM alignment check that reads what the receiving mail server already validated.
**v10 (current) โ 743,571 parameters, context-gated fusion architecture.** Split urgency detection into two independent signals โ coercive pressure versus professional formality โ that were previously conflated into a single blended number. Restructured the network so raw evidence and context evidence enter as separate inputs to a learned fusion layer, instead of being concatenated together and left for the network to disentangle unaided.
---
## Feature groups
| Group | Count | What it measures |
|-------|:-----:|-------------------|
| Domain (raw) | 7 | header consistency, brand impersonation, display-name spoofing |
| URL + body (raw) | 4 | link security, anchor/href mismatch, link density |
| SPF (raw) | 3 | basic sender authorization signal |
| Subject (raw) | 7 | brand mentions, currency, formatting, fake reply prefixes |
| Char-level (raw) | 5 | capitalization, digit density, HTML/text ratio |
| BDI (raw) | 5 | most-common-link-domain mismatch, IP-target forms, open redirects |
| Lexical โ sender domain (raw) | 6 | digit runs, hyphen runs, entropy, typosquat distance |
| Coercive urgency (raw) | 1 | pressure/threat language, independent of formal tone |
| Auth-verified ESP context | 1 | how strongly DMARC/DKIM validate the sender |
| Professional formality (context) | 1 | formal/professional tone, independent of coercive urgency |
| Freemail Reply-To context | 2 | raw fact + how much authentication offsets it |
| Tracking pattern (observational) | 1 | structural resemblance to ESP tracking infrastructure โ never used to exempt anything on its own |
| Cross-signal fusion (context) | 5 | agreement between independent modules |
| Auth alignment (context) | 3 | DMARC pass, DKIM alignment, composite score |
| TF-IDF | 50 | words learned directly from the training corpus |
**101 features total** (38 raw + 13 context + 50 TF-IDF).
---
## Training data
| Dataset | Source | Contribution |
|---------|--------|---------------|
| CEAS-2008 | Kaggle | ~39K emails, 2008-era phishing |
| Enron | Kaggle | ~29K emails, legitimate corporate correspondence |
| SpamAssassin | Kaggle | ~10K emails, mixed spam/legitimate |
| Nigerian Fraud | Kaggle | ~3.3K emails, advance-fee fraud |
| Nazario | Kaggle | ~1.5K emails, phishing corpus |
| Ling-Spam | Kaggle | ~2.8K emails |
| farshad72/spam_email | HuggingFace | 83K rows, includes modern notification-style legitimate email |
| puyang2025/seven-phishing-email-datasets | HuggingFace | 203K rows, unified 7-source corpus |
| **Combined, after deduplication** | โ | **~166,000 emails, ~56% phishing / 44% legitimate** |
---
## Benchmarks
Evaluated on 12,000 held-out samples, self-reported.
| Metric | Value |
|--------|:-----:|
| F1 score | 0.9445 |
| Accuracy | 94.77% |
| Precision | 0.9402 |
| Recall | 0.9489 |
| Parameters | 743,571 |
| Model size on disk | ~3 MB |
| Throughput (GPU) | ~630 emails/sec |
| GPU required | No |
---
## Limitations
Read this before deploying anywhere real.
- **Most training data predates 2010.** Modern phishing techniques โ OAuth abuse, QR code lures, redirect chains through legitimate cloud services โ are underrepresented even after adding modern HuggingFace datasets.
- **False positives on legitimate marketing and notification email are reduced but not eliminated.** The context-gating architecture measurably helps here, but this remains an active area of work rather than a solved problem.
- **No adversarial robustness testing has been performed.** An attacker aware of the exact feature set could plausibly craft targeted bypasses. Use as one signal in a defence-in-depth stack, not a standalone gate.
- **Benchmark numbers are self-reported** on a held-out split of the training corpus, not independently verified or peer-reviewed.
- **Not production-hardened** โ no retry logic, rate limiting, or async network handling.
- **English-language only.**
## Citation
No peer-reviewed paper exists yet. Until then, cite the repository directly:
```bibtex
@software{phishbyte2026,
author = {Singh, Samratth},
title = {Phish_Byte: Context-gated fusion architecture for from-scratch email phishing detection},
year = {2026},
url = {https://github.com/AnonymousSingh-007/Phish_Byte}
}
```
## License
MIT
|