Text Classification
Transformers
Safetensors
bert
sms
spam
phishing
smishing
sms-spam-detection
phishing-detection
fraud-detection
sms-firewall
a2p-messaging
telecom
cybersecurity
text-embeddings-inference
Instructions to use telecomsxchange/OpenTextShield with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use telecomsxchange/OpenTextShield with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="telecomsxchange/OpenTextShield")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("telecomsxchange/OpenTextShield") model = AutoModelForSequenceClassification.from_pretrained("telecomsxchange/OpenTextShield", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from telecomsxchange/OpenTextShield: direct link, hf CLI and curl.
- Browser
- Download file 7.39 kB
-
https://huggingface.co/telecomsxchange/OpenTextShield/resolve/main/README.md
- Command line
-
hf download hf://telecomsxchange/OpenTextShield/README.md
-
curl -L -o README.md https://huggingface.co/telecomsxchange/OpenTextShield/resolve/main/README.md
7.39 kB
| license: mit | |
| library_name: transformers | |
| pipeline_tag: text-classification | |
| base_model: google-bert/bert-base-multilingual-cased | |
| language: | |
| - multilingual | |
| - en | |
| - es | |
| - fr | |
| - de | |
| - pt | |
| - it | |
| - nl | |
| - ar | |
| - he | |
| - hi | |
| - id | |
| - ja | |
| - ru | |
| - tr | |
| - zh | |
| tags: | |
| - sms | |
| - spam | |
| - phishing | |
| - smishing | |
| - sms-spam-detection | |
| - phishing-detection | |
| - fraud-detection | |
| - sms-firewall | |
| - a2p-messaging | |
| - telecom | |
| - cybersecurity | |
| - bert | |
| widget: | |
| - text: "Your account has been suspended. Verify now at http://secure-login-check.xyz" | |
| example_title: Phishing | |
| - text: "Running about 15 min late, order me the usual? I'll grab the bill." | |
| example_title: Legitimate | |
| - text: "CONGRATULATIONS! Your number was picked for a $1,000 gift card. Reply YES to claim before midnight!" | |
| example_title: Spam | |
| # OpenTextShield: open-source multilingual SMS spam and phishing detection | |
| **OpenTextShield is an open-source machine-learning model that detects SMS spam and phishing (smishing).** It is a fine-tuned multilingual BERT (about 180M parameters) that labels a text message as `ham` (legitimate), `spam` or `phishing` in around 150 ms on a small CPU instance. It is used by telecom carriers to screen live SMS traffic and protect subscribers in real networks, and it runs entirely on your own infrastructure — as a REST API, an SMPP proxy in front of your SMSC, or a plain Transformers model. No third-party AI service is involved and no message ever leaves your servers. | |
| - **Try it now:** [Hugging Face Space](https://huggingface.co/spaces/telecomsxchange/OpenTextShield) · [ots.telecomsxchange.com](https://ots.telecomsxchange.com) | |
| - **Source, REST API and SMPP proxy:** [github.com/TelecomsXChangeAPi/OpenTextShield](https://github.com/TelecomsXChangeAPi/OpenTextShield) | |
| - **Docker image (API + model included):** [`telecomsxchange/opentextshield`](https://hub.docker.com/r/telecomsxchange/opentextshield) | |
| ## At a glance | |
| | | | | |
| |---|---| | |
| | Task | SMS / text-message classification: `ham`, `spam`, `phishing` | | |
| | Model | Fine-tuned `bert-base-multilingual-cased`, ~180M parameters | | |
| | Current version | 2.7 | | |
| | Languages | Multilingual (mBERT base); strongest where training data is richest | | |
| | Latency | ~150 ms per message on a small CPU instance; hundreds of messages/s on one GPU with batching | | |
| | Deployment | `transformers` pipeline, Docker, REST API, SMPP proxy | | |
| | Used in | Live carrier SMS traffic (SMSC-side screening via SMPP) | | |
| | License | MIT — free for commercial use | | |
| ## How do I classify an SMS with OpenTextShield? | |
| ```python | |
| from transformers import pipeline | |
| classifier = pipeline("text-classification", model="telecomsxchange/OpenTextShield") | |
| classifier("USPS: Your parcel could not be delivered because of an unpaid customs fee. " | |
| "Settle it within 24h to avoid return: http://usps-redelivery.top/pay") | |
| # [{'label': 'phishing', 'score': 0.9999}] | |
| ``` | |
| | Label | Meaning | | |
| |---|---| | |
| | `ham` | A normal, legitimate message | | |
| | `spam` | Unwanted promotional or bulk content | | |
| | `phishing` | An attempt to steal credentials, money or personal data | | |
| **Note on normalisation:** the production OpenTextShield API normalises text before classification, so zero-width, full-width, homoglyph and leetspeak disguises (`Paypal`, `раураl`, `paypa1`) are classified as the text they imitate. If you load the model directly as above, you get the raw model without that step. The normaliser is a single dependency-free method — [`EnhancedPreprocessor.normalize_unicode`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/src/api_interface/services/enhanced_preprocessing.py) — and is worth applying in front of the model if your traffic may be adversarial. | |
| ## How accurate is OpenTextShield? | |
| Numbers below are for model 2.7, measured through the same text normalisation the production API applies. Full method, caveats and model-to-model comparisons are in [`evals/REPORT.md`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/evals/REPORT.md). | |
| | Benchmark | Messages | Block rate | Phishing recall | | |
| |---|---|---|---| | |
| | UCI SMS Spam Collection (classic spam) | 5,574 | 99.5% | n/a (no phishing class) | | |
| | Mishra & Soni SMS phishing | 5,971 | 99.3% | 6.1% | | |
| | IMC 2025 smishing (modern, multilingual) | 8,007 | 72.4% | 45.7% | | |
| | In-house adversarial suite | 127 | 96.9% | 80.6% | | |
| "Block rate" counts a spam or phishing message as blocked whichever of the two labels it received. UCI and Mishra & Soni overlap the training corpus and serve as regression gates; IMC 2025 is the most independent signal. The spam/phishing boundary is the hardest part of the task: many scams are blocked but under the other label, which is why block rate and phishing recall are reported separately. | |
| ## What languages does it support? | |
| The base model, `bert-base-multilingual-cased`, covers a broad range of languages, so OpenTextShield accepts SMS in essentially any major language — English, Spanish, French, German, Portuguese, Arabic, Hebrew, Hindi, Indonesian, Japanese, Russian, Turkish, Chinese and many more. Accuracy is strongest in the languages best represented in the training corpus; contributions of labelled SMS data in more languages are the most useful thing you can send to the [GitHub project](https://github.com/TelecomsXChangeAPi/OpenTextShield). | |
| ## How do I run it in production? | |
| The same model ships inside the OpenTextShield platform, which adds dynamic batching, text normalisation, Prometheus metrics, audit logging, a TM Forum TMF922 interface and an SMPP proxy that screens `submit_sm` traffic in front of your SMSC — the configuration telecom operators use to protect subscribers on live networks: | |
| ```bash | |
| docker pull telecomsxchange/opentextshield:latest | |
| docker run -d -p 8002:8002 -p 8080:8080 telecomsxchange/opentextshield:latest | |
| curl -X POST "http://localhost:8002/predict/" \ | |
| -H "Content-Type: application/json" \ | |
| -d '{"text":"Your account has been suspended. Verify now at http://secure-login-check.xyz","model":"ots-mbert"}' | |
| ``` | |
| ## How is it different from a cloud SMS-filtering API? | |
| OpenTextShield is self-hosted and MIT-licensed: there are no per-message fees, no vendor lock-in, and message content never leaves your network — which matters for subscriber privacy and for regulators. The model, training scripts, datasets tooling, evaluation harness and deployment stack are all open source. | |
| ## Training | |
| - **Base model:** `bert-base-multilingual-cased` | |
| - **Task:** 3-class sequence classification (`ham` = 0, `spam` = 1, `phishing` = 2) | |
| - **Data:** public SMS spam corpora plus an in-house multilingual corpus labelled `ham` / `spam` / `phishing`; deduplicated across train and test | |
| - **Input length:** SMS-sized; the production API truncates at 96 tokens | |
| Training scripts, dataset tooling and the labelling guide live in the [GitHub repository](https://github.com/TelecomsXChangeAPi/OpenTextShield/tree/main/src/mBERT/training/model-training). | |
| ## Citation | |
| ```bibtex | |
| @software{opentextshield, | |
| title = {OpenTextShield: open-source SMS spam and phishing detection}, | |
| author = {{TelecomsXChange (TCXC)}}, | |
| url = {https://github.com/TelecomsXChangeAPi/OpenTextShield}, | |
| license = {MIT} | |
| } | |
| ``` | |
| ## About | |
| OpenTextShield is built by [TelecomsXChange (TCXC)](https://www.telecomsxchange.com) and released under the [MIT License](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/LICENSE). | |