Instructions to use Horizon-Labs/multilingual-toxicity-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Horizon-Labs/multilingual-toxicity-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Horizon-Labs/multilingual-toxicity-base")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("Horizon-Labs/multilingual-toxicity-base") model = AutoModelForSequenceClassification.from_pretrained("Horizon-Labs/multilingual-toxicity-base", device_map="auto") - Transformers.js
How to use Horizon-Labs/multilingual-toxicity-base with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('text-classification', 'Horizon-Labs/multilingual-toxicity-base'); - Notebooks
- Google Colab
- Kaggle
Multilingual Toxicity (base, 308M): Detoxify labels in many languages
A multi-label toxic-comment classifier with the labels of Detoxify / unitary/unbiased-toxic-roberta: toxicity, severe_toxicity, obscene, threat, insult, identity_attack, sexual_explicit. It works on comments in English and 33 other languages. Built on mmBERT-base and trained on Civil Comments (CC0) plus translations, so it is a multilingual drop-in for Detoxify-style moderation. Apache-2.0. ONNX files for CPU and the browser (transformers.js) are included. A smaller, faster version is available as Horizon-Labs/multilingual-toxicity-small.
- Each label gets an independent probability (sigmoid), as in Detoxify. A common choice is to flag a comment when
toxicity >= 0.5, but tune the threshold for your platform. Scores for non-English text tend to be lower than for English (see F1 @ 0.5 below), so a lower threshold may suit multilingual content. Check it on your own data. - For harmful requests to an LLM (weapons, self-harm, etc.) rather than rude comments, see our content-safety-guard.
onnx/model_quantized.onnx(int8 embeddings, 641 MB): probabilities differ from fp32 by 0.0008 on average over 448 test texts.
Usage
from transformers import pipeline
clf = pipeline("text-classification", model="Horizon-Labs/multilingual-toxicity-base", top_k=None)
print(clf("Halt die Klappe, du Vollidiot."))
# [[{'label': 'toxicity', 'score': ...}, {'label': 'insult', 'score': ...}, ...]]
transformers.js:
import { pipeline } from "@huggingface/transformers";
const clf = await pipeline("text-classification", "Horizon-Labs/multilingual-toxicity-base", { dtype: "q8" });
console.log(await clf("Eres un inútil, lárgate de aquí.", { top_k: null }));
Evaluation
The benchmarks were used only for evaluation, and training texts that also occur in them were removed.
- Civil Comments test (English, 20,000 random comments): ROC AUC per label, with labels binarised at >= 0.5 as in the Jigsaw competitions. The table gives the mean over the 6 labels that have at least 20 positives (severe_toxicity has none at that threshold), and the toxicity AUC. Models without a label are scored on the labels they have.
- TextDetox (textdetox/multilingual_toxicity_dataset): binary toxic/non-toxic in 15 languages, 1,000 balanced texts each. The score is ROC AUC of each model's toxicity probability, plus F1 at 0.5.
| model | licence | Civil Comments, mean AUC (6 labels) | Civil Comments, toxicity AUC | TextDetox (15 languages), AUC | TextDetox, F1 @ 0.5 | labels |
|---|---|---|---|---|---|---|
| this model (308M) | Apache-2.0 | 0.988 | 0.971 | 0.854 | 0.433 | 7 Detoxify labels, multilingual |
| Horizon-Labs/multilingual-toxicity-small (141M) | Apache-2.0 | 0.988 | 0.972 | 0.835 | 0.547 | |
| unitary/unbiased-toxic-roberta (125M) | Apache-2.0 | 0.985 | 0.969 | 0.588 | 0.086 | 7 Detoxify labels (+ identity labels), English |
| unitary/toxic-bert (110M) | Apache-2.0 | 0.946 | 0.926 | 0.623 | 0.142 | 6 labels, English |
| s-nlp/roberta_toxicity_classifier (125M) | OpenRAIL++ | 0.975 | 0.975 | 0.642 | 0.092 | binary, English |
| martin-ha/toxic-comment-model (67M) | none given | 0.943 | 0.943 | 0.572 | 0.160 | binary, English |
| unitary/multilingual-toxic-xlm-roberta (278M) | Apache-2.0 | 0.961 | 0.961 | 0.678 | 0.352 | 1 label, multilingual |
| citizenlab/distilbert-base-multilingual-cased-toxicity (135M) | none given | 0.843 | 0.843 | 0.698 | 0.277 | binary, multilingual |
| textdetox/xlmr-large-toxicity-classifier (560M) | OpenRAIL++ | 0.869 | 0.869 | 0.922 | 0.878 | binary, multilingual; trained on TextDetox data |
- The table shows the released checkpoint. Means over two training seeds: small Civil Comments mean AUC .988 / toxicity .972 / TextDetox .838; base .988 / .972 / .855.
- Models ahead of this one: Civil Comments mean AUC: none (single-label models are scored on their one label, so compare the toxicity column too); Civil Comments toxicity AUC: roberta_toxicity_classifier; TextDetox AUC: xlmr-large-toxicity-classifier. The English models were trained on Jigsaw/Civil Comments data, the same source as this test set. xlmr-large-toxicity-classifier was trained on the TextDetox data.
Per label (Civil Comments, AUC):
| label | this model | unbiased-toxic-roberta |
|---|---|---|
| toxicity | 0.971 | 0.969 |
| obscene | 0.993 | 0.991 |
| threat | 0.993 | 0.985 |
| insult | 0.980 | 0.980 |
| identity_attack | 0.990 | 0.987 |
| sexual_explicit | 0.998 | 0.996 |
Per TextDetox language (AUC):
| language | this model (AUC) | multilingual-toxic-xlm-roberta | xlmr-large-toxicity (in-domain) |
|---|---|---|---|
| Amharic | 0.716 | 0.548 | 0.982 |
| Arabic | 0.787 | 0.413 | 0.989 |
| German | 0.859 | 0.662 | 0.997 |
| English | 0.993 | 0.992 | 1.000 |
| Spanish | 0.907 | 0.819 | 0.998 |
| French | 0.987 | 0.981 | 0.971 |
| Hebrew | 0.788 | 0.501 | 0.701 |
| Hindi | 0.910 | 0.644 | 0.995 |
| Hinglish (romanised Hindi) | 0.765 | 0.631 | 0.756 |
| Italian | 0.875 | 0.818 | 0.805 |
| Japanese | 0.846 | 0.506 | 0.852 |
| Russian | 0.973 | 0.807 | 1.000 |
| Tatar | 0.818 | 0.626 | 0.806 |
| Ukrainian | 0.808 | 0.680 | 0.996 |
| Chinese | 0.774 | 0.547 | 0.979 |
Training
- Data: 518,854 English comments from Civil Comments (CC0; enriched for toxic ones), with their fractional annotator labels. Qwen3.8-27B (Apache-2.0) translated 390,862 of them into 33 languages, about 12,000 per language and half of them toxic. It was told to keep insults, profanity and threats intact, and the labels are copied to each translation.
- Model: mmBERT-base with 7 sigmoid outputs, binary cross-entropy on the fractional labels (as in Detoxify), max length 256 tokens. The checkpoint was chosen by mean AUC on held-out comments and their translations.
- Code:
code/in this repository.
Limitations
- Toxicity labels are subjective and culture-dependent. Translated comments keep English annotators' judgements, and slurs or insults that exist only in other languages are under-represented.
- Like other toxicity models, it can over-flag mentions of identity groups and reclaimed or quoted language, and miss implicit or sarcastic abuse. Do not use it as the only basis for decisions about people.
- Accuracy is lower for low-resource languages, where both translation and the base model are weaker (see the per-language table).
- Downloads last month
- -
Model tree for Horizon-Labs/multilingual-toxicity-base
Base model
jhu-clsp/mmBERT-base