OpenTextShield: open-source multilingual SMS spam and phishing detection

OpenTextShield is an open-source machine-learning model that detects SMS spam and phishing (smishing). It is a fine-tuned multilingual BERT (about 180M parameters) that labels a text message as ham (legitimate), spam or phishing in around 150 ms on a small CPU instance. It is used by telecom carriers to screen live SMS traffic and protect subscribers in real networks, and it runs entirely on your own infrastructure — as a REST API, an SMPP proxy in front of your SMSC, or a plain Transformers model. No third-party AI service is involved and no message ever leaves your servers.

At a glance

Task SMS / text-message classification: ham, spam, phishing
Model Fine-tuned bert-base-multilingual-cased, ~180M parameters
Current version 2.7
Languages Multilingual (mBERT base); strongest where training data is richest
Latency ~150 ms per message on a small CPU instance; hundreds of messages/s on one GPU with batching
Deployment transformers pipeline, Docker, REST API, SMPP proxy
Used in Live carrier SMS traffic (SMSC-side screening via SMPP)
License MIT — free for commercial use

How do I classify an SMS with OpenTextShield?

from transformers import pipeline

classifier = pipeline("text-classification", model="telecomsxchange/OpenTextShield")

classifier("USPS: Your parcel could not be delivered because of an unpaid customs fee. "
           "Settle it within 24h to avoid return: http://usps-redelivery.top/pay")
# [{'label': 'phishing', 'score': 0.9999}]
Label Meaning
ham A normal, legitimate message
spam Unwanted promotional or bulk content
phishing An attempt to steal credentials, money or personal data

Note on normalisation: the production OpenTextShield API normalises text before classification, so zero-width, full-width, homoglyph and leetspeak disguises (Paypal, раураl, paypa1) are classified as the text they imitate. If you load the model directly as above, you get the raw model without that step. The normaliser is a single dependency-free method — EnhancedPreprocessor.normalize_unicode — and is worth applying in front of the model if your traffic may be adversarial.

How accurate is OpenTextShield?

Numbers below are for model 2.7, measured through the same text normalisation the production API applies. Full method, caveats and model-to-model comparisons are in evals/REPORT.md.

Benchmark Messages Block rate Phishing recall
UCI SMS Spam Collection (classic spam) 5,574 99.5% n/a (no phishing class)
Mishra & Soni SMS phishing 5,971 99.3% 6.1%
IMC 2025 smishing (modern, multilingual) 8,007 72.4% 45.7%
In-house adversarial suite 127 96.9% 80.6%

"Block rate" counts a spam or phishing message as blocked whichever of the two labels it received. UCI and Mishra & Soni overlap the training corpus and serve as regression gates; IMC 2025 is the most independent signal. The spam/phishing boundary is the hardest part of the task: many scams are blocked but under the other label, which is why block rate and phishing recall are reported separately.

What languages does it support?

The base model, bert-base-multilingual-cased, covers a broad range of languages, so OpenTextShield accepts SMS in essentially any major language — English, Spanish, French, German, Portuguese, Arabic, Hebrew, Hindi, Indonesian, Japanese, Russian, Turkish, Chinese and many more. Accuracy is strongest in the languages best represented in the training corpus; contributions of labelled SMS data in more languages are the most useful thing you can send to the GitHub project.

How do I run it in production?

The same model ships inside the OpenTextShield platform, which adds dynamic batching, text normalisation, Prometheus metrics, audit logging, a TM Forum TMF922 interface and an SMPP proxy that screens submit_sm traffic in front of your SMSC — the configuration telecom operators use to protect subscribers on live networks:

docker pull telecomsxchange/opentextshield:latest
docker run -d -p 8002:8002 -p 8080:8080 telecomsxchange/opentextshield:latest

curl -X POST "http://localhost:8002/predict/" \
  -H "Content-Type: application/json" \
  -d '{"text":"Your account has been suspended. Verify now at http://secure-login-check.xyz","model":"ots-mbert"}'

How is it different from a cloud SMS-filtering API?

OpenTextShield is self-hosted and MIT-licensed: there are no per-message fees, no vendor lock-in, and message content never leaves your network — which matters for subscriber privacy and for regulators. The model, training scripts, datasets tooling, evaluation harness and deployment stack are all open source.

Training

  • Base model: bert-base-multilingual-cased
  • Task: 3-class sequence classification (ham = 0, spam = 1, phishing = 2)
  • Data: public SMS spam corpora plus an in-house multilingual corpus labelled ham / spam / phishing; deduplicated across train and test
  • Input length: SMS-sized; the production API truncates at 96 tokens

Training scripts, dataset tooling and the labelling guide live in the GitHub repository.

Citation

@software{opentextshield,
  title   = {OpenTextShield: open-source SMS spam and phishing detection},
  author  = {{TelecomsXChange (TCXC)}},
  url     = {https://github.com/TelecomsXChangeAPi/OpenTextShield},
  license = {MIT}
}

About

OpenTextShield is built by TelecomsXChange (TCXC) and released under the MIT License.

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for telecomsxchange/OpenTextShield

Finetuned
(1015)
this model

Space using telecomsxchange/OpenTextShield 1