Toxic Talking Detector

A fine-tuned DistilBERT model for multi-label toxicity classification, exported to ONNX (INT8 quantized) for lightweight deployment.

Model Description

This model is fine-tuned from distilbert-base-uncased on the Jigsaw Unintended Bias in Toxicity Classification dataset. It predicts continuous toxicity scores (0-1) across 7 categories for a given text input.

Labels:

  • toxicity
  • severe_toxicity
  • obscene
  • threat
  • insult
  • identity_attack
  • sexual_explicit

Intended Use

Designed for conversation moderation systems — score individual messages or aggregate scores across a multi-turn conversation to detect escalating toxicity.

Files

File Description
model_int8.onnx Quantized ONNX model (~67MB)
vocab.txt Tokenizer vocabulary
tokenizer.json Fast tokenizer config
tokenizer_config.json Tokenizer settings
special_tokens_map.json Special token mapping
label_order.json Output label order (maps model output indices to category names)

Training Details

  • Base model: distilbert-base-uncased
  • Dataset: Jigsaw Civil Comments (~319K balanced subsample)
  • Task: Multi-label classification with soft/continuous targets (BCEWithLogitsLoss)
  • Epochs: 2
  • Eval F1 (micro): 0.840
  • Eval F1 (macro): 0.616

How to Use

With ONNX Runtime (lightweight, no PyTorch needed)

from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
import onnxruntime as ort
import numpy as np
import json

REPO_ID = "bsgcasa/toxic-talking-detector"

model_path = hf_hub_download(repo_id=REPO_ID, filename="model_int8.onnx")
label_order_path = hf_hub_download(repo_id=REPO_ID, filename="label_order.json")
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)

with open(label_order_path) as f:
    LABELS = json.load(f)

session = ort.InferenceSession(model_path)

def predict(text: str) -> dict:
    enc = tokenizer(text, return_tensors="np", padding="max_length", truncation=True, max_length=128)
    ort_inputs = {
        "input_ids": enc["input_ids"].astype(np.int64),
        "attention_mask": enc["attention_mask"].astype(np.int64),
    }
    logits = session.run(["logits"], ort_inputs)[0]
    probs = 1 / (1 + np.exp(-logits))[0]
    return {label: float(p) for label, p in zip(LABELS, probs)}

print(predict("You're completely useless, shut up."))

Limitations

  • Trained on English text only.
  • Scores individual messages; conversation-level aggregation must be handled by the calling application.
  • Inherits any biases present in the Civil Comments dataset (e.g. topic/identity term correlations).

License

Apache 2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support