Hadhari (حذارِ) - Arabic Spam Detection

Hadhari is a machine learning model designed to detect spam messages in Arabic text, targeting unsolicited advertisements.

Model Details

  • Algorithm: Linear Support Vector Classification (LinearSVC)
  • Feature Extraction: TF-IDF Vectorization (max_features=2000, ngram_range=(1,3))

Dataset

The model was trained on a human-in-the-loop (HITL) verified dataset collected from a live WhatsApp bot.

  • Size: ~1000 unique messages
  • Distribution: Balanced between genuine conversations and student-targeted spam.
  • Characteristics: Contains real-world Saudi colloquial Arabic, slang, emojis, and common local spam patterns.

Example Data

Class Message Example
Spam (1) حل واجـــبات المواد
بحــــــــــوثات علمــــــيةمشـــــــــــاريع تـخــــــرج*
https://wa.me/...
Spam (1) أعذار طبية(إجـازة مـرضـية)
مـعتـمد فـي تطبيق صـحـتي
(تقـبل لـجميع جهات العمل)

Performance Metrics

  • Overall Accuracy: 97.0%
  • Spam Precision: 99.0%

Usage

Live API endpoint and documentation available at: Hadhari API Space

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using mabosaimi/hadhari 1