Instructions to use Akash-Sakala/bert-phishing-classifier_student_4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Akash-Sakala/bert-phishing-classifier_student_4bit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Akash-Sakala/bert-phishing-classifier_student_4bit")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("Akash-Sakala/bert-phishing-classifier_student_4bit") model = AutoModelForSequenceClassification.from_pretrained("Akash-Sakala/bert-phishing-classifier_student_4bit", device_map="auto") - Notebooks
- Google Colab
- Kaggle
DistilBERT Phishing URL Classifier: 4-bit NF4 (Distilled + Quantized)
The last stage of a distil-then-quantise pipeline. This is the 4-layer DistilBERT student with its linear layers quantised to 4-bit NF4 (bitsandbytes, double quantisation). It classifies a URL as phishing or benign.
BERT-base teacher ──KD──▶ DistilBERT student ──NF4──▶ 4-bit student (this)
438 MB · acc 89.7% 211 MB · acc 96.0% 110 MB · acc 95.8%
Results
All three rows are on the full test split (33,000 URLs, balanced):
| Model | Size on disk | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|---|
| BERT teacher (fp32) | 438 MB | 0.8971 | 0.9136 | 0.8763 | 0.8945 |
| DistilBERT student (fp32) | 211 MB | 0.9601 | 0.9710 | 0.9483 | 0.9595 |
| 4-bit NF4 student (this) | 110 MB | 0.9583 | 0.9577 | 0.9587 | 0.9582 |
- 4× smaller than the teacher and ~48% smaller than the fp32 student. Only the
nn.Linearweights are quantised. The 30,522 × 768 word-embedding table stays in fp32 and now makes up most of the file. - Quantisation costs only 0.18 points of accuracy (0.9601 → 0.9583) compared with the fp32 student. It also shifts the balance slightly toward recall (0.948 → 0.959) at the cost of precision (0.971 → 0.958), so fewer phishing URLs are missed.
Teacher test numbers come from the student card. The student and 4-bit numbers were re-measured on 27 Sep 2026 with max_length=128. The 4-bit row uses the stored NF4 weights, dequantised to bf16. This is the arithmetic bitsandbytes runs on GPU with bnb_4bit_compute_dtype=bfloat16.
Quantisation config
| Method | bitsandbytes load_in_4bit |
| Quant type | NF4 (4-bit NormalFloat) |
| Double quantisation | ✅ (quantises the quantisation constants too) |
| Compute dtype | bfloat16 |
| Architecture | DistilBertForSequenceClassification: 4 layers, 8 heads, hidden 768, ~52M params |
Distillation settings for the student (T = 3.0, α = 0.6 KL + 0.4 CE, 4 epochs) are documented on the student card.
Usage
Needs a CUDA GPU with bitsandbytes installed. The quantisation config is stored in config.json, so the model loads already quantised.
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
repo = "Akash-Sakala/bert-phishing-classifier_student_4bit"
tokenizer = AutoTokenizer.from_pretrained(repo) # standard distilbert-base-uncased vocab
model = AutoModelForSequenceClassification.from_pretrained(repo, device_map="cuda").eval()
urls = ["paypal.com.secure-login.verify-account.ru/webscr", "wikipedia.org/wiki/Phishing"]
inputs = tokenizer(urls, return_tensors="pt", truncation=True, padding=True, max_length=128).to(model.device)
with torch.no_grad():
probs = model(input_ids=inputs["input_ids"], attention_mask=inputs["attention_mask"]).logits.softmax(-1)[:, 1]
for u, p in zip(urls, probs):
print(f"{p:.2f} {'PHISHING' if p > 0.5 else 'benign'} {u}")
For CPU-only inference, use the fp32 student. bitsandbytes' 4-bit CPU kernels don't yet support this model's 2-class output layer.
Limitations
- The model sees only the URL string. Use it as one signal alongside reputation feeds and page analysis.
- Most training URLs have no
http(s)://scheme. Strip the scheme before inference. - Accuracy on new phishing campaigns will be lower than on this static test set.
Related
- Teacher:
bert-phishing-classifier_teacher - fp32 student:
bert-phishing-classifier_student - Dataset:
phishing-site-classification
- Downloads last month
- 29
Model tree for Akash-Sakala/bert-phishing-classifier_student_4bit
Base model
distilbert/distilbert-base-uncasedDataset used to train Akash-Sakala/bert-phishing-classifier_student_4bit
Evaluation results
- Accuracy on Akash-Sakala/phishing-site-classificationtest set self-reported0.958
- Precision on Akash-Sakala/phishing-site-classificationtest set self-reported0.958
- Recall on Akash-Sakala/phishing-site-classificationtest set self-reported0.959
- F1 on Akash-Sakala/phishing-site-classificationtest set self-reported0.958