Toxic Talking Detector
A fine-tuned DistilBERT model for multi-label toxicity classification, exported to ONNX (INT8 quantized) for lightweight deployment.
Model Description
This model is fine-tuned from distilbert-base-uncased on the Jigsaw Unintended Bias in Toxicity Classification dataset. It predicts continuous toxicity scores (0-1) across 7 categories for a given text input.
Labels:
toxicitysevere_toxicityobscenethreatinsultidentity_attacksexual_explicit
Intended Use
Designed for conversation moderation systems — score individual messages or aggregate scores across a multi-turn conversation to detect escalating toxicity.
Files
| File | Description |
|---|---|
model_int8.onnx |
Quantized ONNX model (~67MB) |
vocab.txt |
Tokenizer vocabulary |
tokenizer.json |
Fast tokenizer config |
tokenizer_config.json |
Tokenizer settings |
special_tokens_map.json |
Special token mapping |
label_order.json |
Output label order (maps model output indices to category names) |
Training Details
- Base model:
distilbert-base-uncased - Dataset: Jigsaw Civil Comments (~319K balanced subsample)
- Task: Multi-label classification with soft/continuous targets (BCEWithLogitsLoss)
- Epochs: 2
- Eval F1 (micro): 0.840
- Eval F1 (macro): 0.616
How to Use
With ONNX Runtime (lightweight, no PyTorch needed)
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
import onnxruntime as ort
import numpy as np
import json
REPO_ID = "bsgcasa/toxic-talking-detector"
model_path = hf_hub_download(repo_id=REPO_ID, filename="model_int8.onnx")
label_order_path = hf_hub_download(repo_id=REPO_ID, filename="label_order.json")
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)
with open(label_order_path) as f:
LABELS = json.load(f)
session = ort.InferenceSession(model_path)
def predict(text: str) -> dict:
enc = tokenizer(text, return_tensors="np", padding="max_length", truncation=True, max_length=128)
ort_inputs = {
"input_ids": enc["input_ids"].astype(np.int64),
"attention_mask": enc["attention_mask"].astype(np.int64),
}
logits = session.run(["logits"], ort_inputs)[0]
probs = 1 / (1 + np.exp(-logits))[0]
return {label: float(p) for label, p in zip(LABELS, probs)}
print(predict("You're completely useless, shut up."))
Limitations
- Trained on English text only.
- Scores individual messages; conversation-level aggregation must be handled by the calling application.
- Inherits any biases present in the Civil Comments dataset (e.g. topic/identity term correlations).
License
Apache 2.0