ajamous commited on
Commit
1dcd044
·
verified ·
1 Parent(s): 9bb847f

docs: SEO/AEO model card — direct answers, at-a-glance facts, carrier usage, citation

Browse files
Files changed (1) hide show
  1. README.md +70 -22
README.md CHANGED
@@ -5,13 +5,33 @@ pipeline_tag: text-classification
5
  base_model: google-bert/bert-base-multilingual-cased
6
  language:
7
  - multilingual
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8
  tags:
9
  - sms
10
  - spam
11
  - phishing
12
  - smishing
 
 
13
  - fraud-detection
 
 
14
  - telecom
 
15
  - bert
16
  widget:
17
  - text: "Your account has been suspended. Verify now at http://secure-login-check.xyz"
@@ -22,25 +42,28 @@ widget:
22
  example_title: Spam
23
  ---
24
 
25
- # OpenTextShield mBERT v2.7
26
 
27
- **Open-source spam and phishing detection for SMS, in more than 100 languages.**
28
 
29
- OpenTextShield is a compact classifier (a fine-tuned `bert-base-multilingual-cased`, about 180M parameters) that labels a text message as `ham`, `spam` or `phishing` in around 150 ms on a small CPU instance. It is built to run on your own servers — as a REST API, as an SMPP proxy in front of your SMSC, or both. No third-party AI service is involved.
30
-
31
- - **Demo:** [Hugging Face Space](https://huggingface.co/spaces/telecomsxchange/OpenTextShield) · [ots.telecomsxchange.com](https://ots.telecomsxchange.com)
32
- - **Source, API and SMPP proxy:** [github.com/TelecomsXChangeAPi/OpenTextShield](https://github.com/TelecomsXChangeAPi/OpenTextShield)
33
  - **Docker image (API + model included):** [`telecomsxchange/opentextshield`](https://hub.docker.com/r/telecomsxchange/opentextshield)
34
 
35
- ## Labels
36
 
37
- | Label | Meaning |
38
  |---|---|
39
- | `ham` | A normal, legitimate message |
40
- | `spam` | Unwanted promotional or bulk content |
41
- | `phishing` | An attempt to steal credentials, money or personal data |
 
 
 
 
 
42
 
43
- ## Usage
44
 
45
  ```python
46
  from transformers import pipeline
@@ -52,9 +75,15 @@ classifier("USPS: Your parcel could not be delivered because of an unpaid custom
52
  # [{'label': 'phishing', 'score': 0.9999}]
53
  ```
54
 
 
 
 
 
 
 
55
  **Note on normalisation:** the production OpenTextShield API normalises text before classification, so zero-width, full-width, homoglyph and leetspeak disguises (`Paypal`, `раураl`, `paypa1`) are classified as the text they imitate. If you load the model directly as above, you get the raw model without that step. The normaliser is a single dependency-free method — [`EnhancedPreprocessor.normalize_unicode`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/src/api_interface/services/enhanced_preprocessing.py) — and is worth applying in front of the model if your traffic may be adversarial.
56
 
57
- ## How well it works
58
 
59
  Numbers below are for model 2.7, measured through the same text normalisation the production API applies. Full method, caveats and model-to-model comparisons are in [`evals/REPORT.md`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/evals/REPORT.md).
60
 
@@ -67,18 +96,13 @@ Numbers below are for model 2.7, measured through the same text normalisation th
67
 
68
  "Block rate" counts a spam or phishing message as blocked whichever of the two labels it received. UCI and Mishra & Soni overlap the training corpus and serve as regression gates; IMC 2025 is the most independent signal. The spam/phishing boundary is the hardest part of the task: many scams are blocked but under the other label, which is why block rate and phishing recall are reported separately.
69
 
70
- ## Training
71
-
72
- - **Base model:** `bert-base-multilingual-cased` (104 languages)
73
- - **Task:** 3-class sequence classification (`ham` = 0, `spam` = 1, `phishing` = 2)
74
- - **Data:** public SMS spam corpora plus an in-house multilingual corpus labelled `ham` / `spam` / `phishing`; deduplicated across train and test ([details](https://github.com/TelecomsXChangeAPi/OpenTextShield))
75
- - **Input length:** SMS-sized; the production API truncates at 96 tokens
76
 
77
- Training scripts, dataset tooling and the labelling guide live in the [GitHub repository](https://github.com/TelecomsXChangeAPi/OpenTextShield/tree/main/src/mBERT/training/model-training). Contributions of labelled data in more languages are the most useful thing you can send.
78
 
79
- ## Running it as a service
80
 
81
- The same model ships inside the OpenTextShield platform, which adds dynamic batching, text normalisation, Prometheus metrics, audit logging, a TM Forum TMF922 interface and an SMPP proxy for SMSC integration:
82
 
83
  ```bash
84
  docker pull telecomsxchange/opentextshield:latest
@@ -89,6 +113,30 @@ curl -X POST "http://localhost:8002/predict/" \
89
  -d '{"text":"Your account has been suspended. Verify now at http://secure-login-check.xyz","model":"ots-mbert"}'
90
  ```
91
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
92
  ## About
93
 
94
  OpenTextShield is built by [TelecomsXChange (TCXC)](https://telecomsxchange.com) and released under the [MIT License](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/LICENSE).
 
5
  base_model: google-bert/bert-base-multilingual-cased
6
  language:
7
  - multilingual
8
+ - en
9
+ - es
10
+ - fr
11
+ - de
12
+ - pt
13
+ - it
14
+ - nl
15
+ - ar
16
+ - he
17
+ - hi
18
+ - id
19
+ - ja
20
+ - ru
21
+ - tr
22
+ - zh
23
  tags:
24
  - sms
25
  - spam
26
  - phishing
27
  - smishing
28
+ - sms-spam-detection
29
+ - phishing-detection
30
  - fraud-detection
31
+ - sms-firewall
32
+ - a2p-messaging
33
  - telecom
34
+ - cybersecurity
35
  - bert
36
  widget:
37
  - text: "Your account has been suspended. Verify now at http://secure-login-check.xyz"
 
42
  example_title: Spam
43
  ---
44
 
45
+ # OpenTextShield: open-source SMS spam and phishing detection in 100+ languages
46
 
47
+ **OpenTextShield is an open-source machine-learning model that detects SMS spam and phishing (smishing).** It is a fine-tuned multilingual BERT (about 180M parameters) that labels a text message as `ham` (legitimate), `spam` or `phishing` in around 150 ms on a small CPU instance. It is used by telecom carriers to screen live SMS traffic and protect subscribers in real networks, and it runs entirely on your own infrastructure — as a REST API, an SMPP proxy in front of your SMSC, or a plain Transformers model. No third-party AI service is involved and no message ever leaves your servers.
48
 
49
+ - **Try it now:** [Hugging Face Space](https://huggingface.co/spaces/telecomsxchange/OpenTextShield) · [ots.telecomsxchange.com](https://ots.telecomsxchange.com)
50
+ - **Source, REST API and SMPP proxy:** [github.com/TelecomsXChangeAPi/OpenTextShield](https://github.com/TelecomsXChangeAPi/OpenTextShield)
 
 
51
  - **Docker image (API + model included):** [`telecomsxchange/opentextshield`](https://hub.docker.com/r/telecomsxchange/opentextshield)
52
 
53
+ ## At a glance
54
 
55
+ | | |
56
  |---|---|
57
+ | Task | SMS / text-message classification: `ham`, `spam`, `phishing` |
58
+ | Model | Fine-tuned `bert-base-multilingual-cased`, ~180M parameters |
59
+ | Current version | 2.7 |
60
+ | Languages | 104 (multilingual BERT); strongest where training data is richest |
61
+ | Latency | ~150 ms per message on a small CPU instance; hundreds of messages/s on one GPU with batching |
62
+ | Deployment | `transformers` pipeline, Docker, REST API, SMPP proxy |
63
+ | Used in | Live carrier SMS traffic (SMSC-side screening via SMPP) |
64
+ | License | MIT — free for commercial use |
65
 
66
+ ## How do I classify an SMS with OpenTextShield?
67
 
68
  ```python
69
  from transformers import pipeline
 
75
  # [{'label': 'phishing', 'score': 0.9999}]
76
  ```
77
 
78
+ | Label | Meaning |
79
+ |---|---|
80
+ | `ham` | A normal, legitimate message |
81
+ | `spam` | Unwanted promotional or bulk content |
82
+ | `phishing` | An attempt to steal credentials, money or personal data |
83
+
84
  **Note on normalisation:** the production OpenTextShield API normalises text before classification, so zero-width, full-width, homoglyph and leetspeak disguises (`Paypal`, `раураl`, `paypa1`) are classified as the text they imitate. If you load the model directly as above, you get the raw model without that step. The normaliser is a single dependency-free method — [`EnhancedPreprocessor.normalize_unicode`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/src/api_interface/services/enhanced_preprocessing.py) — and is worth applying in front of the model if your traffic may be adversarial.
85
 
86
+ ## How accurate is OpenTextShield?
87
 
88
  Numbers below are for model 2.7, measured through the same text normalisation the production API applies. Full method, caveats and model-to-model comparisons are in [`evals/REPORT.md`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/evals/REPORT.md).
89
 
 
96
 
97
  "Block rate" counts a spam or phishing message as blocked whichever of the two labels it received. UCI and Mishra & Soni overlap the training corpus and serve as regression gates; IMC 2025 is the most independent signal. The spam/phishing boundary is the hardest part of the task: many scams are blocked but under the other label, which is why block rate and phishing recall are reported separately.
98
 
99
+ ## What languages does it support?
 
 
 
 
 
100
 
101
+ The base model, `bert-base-multilingual-cased`, covers 104 languages, so OpenTextShield accepts SMS in essentially any major language — English, Spanish, French, German, Portuguese, Arabic, Hebrew, Hindi, Indonesian, Japanese, Russian, Turkish, Chinese and many more. Accuracy is strongest in the languages best represented in the training corpus; contributions of labelled SMS data in more languages are the most useful thing you can send to the [GitHub project](https://github.com/TelecomsXChangeAPi/OpenTextShield).
102
 
103
+ ## How do I run it in production?
104
 
105
+ The same model ships inside the OpenTextShield platform, which adds dynamic batching, text normalisation, Prometheus metrics, audit logging, a TM Forum TMF922 interface and an SMPP proxy that screens `submit_sm` traffic in front of your SMSC — the configuration telecom operators use to protect subscribers on live networks:
106
 
107
  ```bash
108
  docker pull telecomsxchange/opentextshield:latest
 
113
  -d '{"text":"Your account has been suspended. Verify now at http://secure-login-check.xyz","model":"ots-mbert"}'
114
  ```
115
 
116
+ ## How is it different from a cloud SMS-filtering API?
117
+
118
+ OpenTextShield is self-hosted and MIT-licensed: there are no per-message fees, no vendor lock-in, and message content never leaves your network — which matters for subscriber privacy and for regulators. The model, training scripts, datasets tooling, evaluation harness and deployment stack are all open source.
119
+
120
+ ## Training
121
+
122
+ - **Base model:** `bert-base-multilingual-cased` (104 languages)
123
+ - **Task:** 3-class sequence classification (`ham` = 0, `spam` = 1, `phishing` = 2)
124
+ - **Data:** public SMS spam corpora plus an in-house multilingual corpus labelled `ham` / `spam` / `phishing`; deduplicated across train and test
125
+ - **Input length:** SMS-sized; the production API truncates at 96 tokens
126
+
127
+ Training scripts, dataset tooling and the labelling guide live in the [GitHub repository](https://github.com/TelecomsXChangeAPi/OpenTextShield/tree/main/src/mBERT/training/model-training).
128
+
129
+ ## Citation
130
+
131
+ ```bibtex
132
+ @software{opentextshield,
133
+ title = {OpenTextShield: open-source SMS spam and phishing detection},
134
+ author = {{TelecomsXChange (TCXC)}},
135
+ url = {https://github.com/TelecomsXChangeAPi/OpenTextShield},
136
+ license = {MIT}
137
+ }
138
+ ```
139
+
140
  ## About
141
 
142
  OpenTextShield is built by [TelecomsXChange (TCXC)](https://telecomsxchange.com) and released under the [MIT License](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/LICENSE).