hatespeech / README.md
Isa0's picture
Update README and sanitize test cases
2e3e2b8
|
Raw History Blame Contribute Delete
5.88 kB
---
license: cc0-1.0
language:
- en
tags:
- text-classification
- onnx
- onnxruntime
- hate-speech-detection
- offensive-language
- int8
- quantization
pipeline_tag: text-classification
widget:
- text: "I really love this community, everyone is so supportive and kind!"
example_title: "Neutral"
- text: "Shut up, you are being so damn annoying and stupid."
example_title: "Offensive Language"
- text: "Those people are subhuman and don't belong in our country, kick them all out."
example_title: "Hate Speech"
---
# Hate Speech & Offensive Language Classifier (ONNX & INT8)
A lightweight, high-performance text classification model fine-tuned to distinguish between **targeted hate speech**, **offensive language (profanity)**, and **neutral content**.
Both the **Full Precision ONNX (FP32)** and an ultra-compact **Dynamic INT8 Quantized (8q)** model are included for production-ready, ultra-low latency inference on CPUs and edge devices.
---
## Key Features
* **Fine-Grained Distinction**: Accurately differentiates between general offensive language/profanity and genuinely dangerous hate speech.
* **Dual ONNX Models**:
* `model/hatespeech.onnx` (FP32, ~255 MB)
* `model/hatespeech_int8.onnx` (INT8 Quantized, **~64 MB**, **~75% size reduction**)
* **Ultra-Low Latency**: ~**17 ms** per sample on standard CPU with ONNX Runtime.
* **Balanced Class Weighting**: Trained with normalized inverse-frequency class weights to combat severe class imbalance (Hate speech is only ~5.8% of the training dataset).
* **Anti-Overfitting Protection**: Stratified split, weight decay ($0.01$), dropout ($0.2$), and early stopping monitoring **Validation Macro F1**.
---
## Classes
| Class ID | Label | Description |
| :---: | :--- | :--- |
| `0` | **Hate Speech** | Targeted hostility, incitement of violence, or dehumanization against protected groups. |
| `1` | **Offensive Language** | Swear words, slang, insults, and vulgarity without targeted hatred. |
| `2` | **Neither** | Neutral, positive, benign, or conversational language. |
---
## Benchmark & Model Specifications
| Property | Raw ONNX | INT8 Quantized (Recommended) |
| :--- | :--- | :--- |
| **File** | `model/hatespeech.onnx` | `model/hatespeech_int8.onnx` |
| **Precision** | Float32 | Quantized Int8 (Weights) |
| **File Size** | ~255 MB | **64.27 MB** |
| **Inference Engine** | ONNX Runtime | ONNX Runtime |
| **Average Latency (CPU)** | ~28 ms | **~17 ms** |
| **Dynamic Inputs** | Dynamic Batch & Sequence Length | Dynamic Batch & Sequence Length |
---
## Quickstart
### 1. Installation
Install dependencies using `pip` or `uv`:
```bash
pip install onnxruntime transformers numpy
# or using uv:
uv add onnxruntime transformers numpy
```
### 2. Standalone Inference with ONNX Runtime
You can run predictions with just **ONNX Runtime** and Hugging Face's `AutoTokenizer`:
```python
import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
LABEL_NAMES = {0: "Hate Speech", 1: "Offensive Language", 2: "Neither"}
# 1. Load ONNX model and tokenizer
model_path = "./model/hatespeech_int8.onnx"
session = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])
tokenizer = AutoTokenizer.from_pretrained("./model")
# 2. Tokenize input text
text = "I really love this community, everyone is so supportive and kind!"
inputs = tokenizer(text, padding=True, truncation=True, max_length=128, return_tensors="np")
# 3. Run inference
ort_inputs = {
"input_ids": inputs["input_ids"].astype(np.int64),
"attention_mask": inputs["attention_mask"].astype(np.int64),
}
logits = session.run(None, ort_inputs)[0]
# 4. Softmax probabilities
exp_logits = np.exp(logits - np.max(logits, axis=-1, keepdims=True))
probs = exp_logits / np.sum(exp_logits, axis=-1, keepdims=True)
pred_id = int(np.argmax(probs, axis=-1)[0])
print(f"Prediction: {LABEL_NAMES[pred_id]} ({probs[0][pred_id] * 100:.2f}%)")
```
---
## CLI & Interactive Usage
This repository includes [`hatespeech.py`](./hatespeech.py) for easy command-line testing:
```bash
# Run benchmark examples
python hatespeech.py
# Predict a custom sentence
python hatespeech.py --text "Stop being so annoying!"
# Start live interactive prompt
python hatespeech.py --interactive
# Use raw FP32 model instead of INT8
python hatespeech.py --raw
```
---
## Training & Architecture
* **Base Model**: `distilbert-base-uncased`
* **Dataset**: Davidson et al. (2017) *Automated Hate Speech and Offensive Language Detection* (~24,783 annotated samples).
* **Data Split**: Stratified 80% Train, 10% Validation, 10% Holdout Test.
* **Loss Function**: `nn.CrossEntropyLoss` with balanced class weights:
$$w_c = \frac{N_{\text{total}}}{N_{\text{classes}} \times N_c}$$
* **Optimizer**: `AdamW` (learning rate: $2 \times 10^{-5}$, weight decay: $0.01$).
* **LR Scheduler**: Linear warmup ($10\%$ of steps) followed by linear decay.
* **Early Stopping**: Monitored on validation Macro F1 score with patience of 2 epochs.
* **Quantization**: Dynamic INT8 quantization executed using `onnxruntime.quantization.quantize_dynamic`.
To reproduce training:
```bash
python train.py --epochs 3 --batch_size 32
```
---
## Citation
If you use this model or dataset in your research, please cite the underlying dataset by Davidson et al.:
```bibtex
@inproceedings{davidson2017automated,
title={Automated Hate Speech Detection and the Problem of Offensive Language},
author={Davidson, Thomas and Warmsley, Dana and Macy, Michael and Weber, Ingmar},
booktitle={Proceedings of the 11th International AAAI Conference on Web and Social Media},
series={ICWSM '17},
pages={512--515},
year={2017}
}
```
---
## License
This repository and model card are released under the [Creative Commons Zero v1.0 Universal (CC0-1.0)](https://creativecommons.org/publicdomain/zero/1.0/) license.