πŸ›‘οΈ Hinglish Toxicity Classifier - mmBERT (darelphilip/hinglisToxicity_darel_mmberta)

A fine-tuned sequence classification model designed for precise multi-label toxicity detection in Hinglish (Hindi-English code-switched) text. It classifies text across 7 distinct toxicity and identity-based discrimination vectors common to South Asian digital spaces.

πŸš€ Live Interactive Demo: Test this model's batch-processing capabilities in real-time on Hugging Face Spaces.


πŸ“– Model Details

πŸ“ Model Description

This model leverages the jhu-clsp/mmBERT-base foundation (ModernBERT architecture) utilizing native Scaled Dot-Product Attention (SDPA) and Gemma 2 tokenization. It was fine-tuned on over 245,000 code-switched online comments to identify specific vectors of abuse, profanity, and harassment. The training process utilized a weighted Binary Cross-Entropy loss function (pos_weight) to penalize false negatives on severe but underrepresented hate speech categories.

  • Developed by: Darel Philip (darelphilip)
  • Contact / Author Email: enigmaticdarel@gmail.com
  • Model Type: Multi-label sequence classification (ModernBERT encoder)
  • Language(s) (NLP): Hinglish (hi-en), Hindi (hi), English (en)
  • License: Apache-2.0
  • Finetuned from model: jhu-clsp/mmBERT-base

🎯 Target Classification Labels

The model outputs independent probabilities for 7 classes:

  1. profanity_vulgarity
  2. targeted_abuse_harassment
  3. discriminatory_hate_speech
  4. caste
  5. communal_religious
  6. regional_xenophobic
  7. misogyny_gender

πŸ’» Uses

βœ… Direct Use

  • Automated Community Moderation: Risk filtering and toxicity scoring in Hinglish forums, discussion boards, and comment streams.
  • Toxicity Auditing: Batch-processing historical data to identify trends in regional abuse or identity-based harassment.

❌ Out-of-Scope Use

  • Not designed for text generation, translation, or open-ended dialogue tasks.
  • Performance may degrade on deeply obscure regional dialects, code-switching with non-Hindi languages, or formal Hindi written entirely in Devanagari script without transliteration.

⚠️ Bias, Risks, and Limitations

  • Contextual Slang: The model is sourced from regional community discussions; highly colloquial slang or reclaimed terminology may occasionally trigger false positives in the profanity_vulgarity category.
  • Uncertainty Zones: Predictions scoring between 0.35 and 0.65 represent statistical uncertainty and are ideal candidates for human-in-the-loop review.

πŸ› οΈ How to Get Started with the Model

Use the code below for multi-label inference:

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

MODEL_ID = "darelphilip/hinglisToxicity_darel_mmberta"
LABEL_COLS = [
    'profanity_vulgarity', 'targeted_abuse_harassment', 'discriminatory_hate_speech',
    'caste', 'communal_religious', 'regional_xenophobic', 'misogyny_gender'
]

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_ID)

text = "kya bakwaas chal raha hai yahan"

inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=160)
with torch.no_grad():
    logits = model(**inputs).logits
    probs = torch.sigmoid(logits).squeeze().tolist()

scores = dict(zip(LABEL_COLS, [round(p, 4) for p in probs]))
print(f"Comment: {text}")
for label, score in scores.items():
    print(f"  {label:<30}: {score:.4f} {'🚨' if score > 0.5 else ''}")
Downloads last month
38
Safetensors
Model size
0.3B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for darelphilip/hinglisToxicity_darel_mmberta

Finetuned
(175)
this model

Spaces using darelphilip/hinglisToxicity_darel_mmberta 2