Image-Only Vision Model (ResNet18) -- Diagnostic/Grad-CAM Use Only

This model does NOT produce the deployed hate-speech classification. It is hosted here solely so the live API can compute a Grad-CAM overlay (transparency/diagnostic visualization) for uploaded images. The actual /predict/image prediction comes from OCR-extracted text run through Prat-04/hate-speech-distilbert.

Why this model isn't used for the actual prediction

Trained and evaluated on the MultiOFF dataset (Suryawanshi et al., 2020) as part of a required image-only/text-only/fusion comparison:

  • Image-only test macro-F1: 0.5117 (barely above majority-class baseline)
  • Text-only test macro-F1: 0.6573
  • Grad-CAM analysis (see source repository) found this model's attention on text-heavy memes mostly reflects generic text-density texture patterns, not genuine visual understanding of hateful content -- it cannot "read" meme text the way OCR + a text model can.

Full comparison, methodology, and the decision rationale are documented in the source repository under "Multimodal Results."

Architecture

ResNet18 (ImageNet-pretrained), early layers frozen, only the last residual block (layer4) fine-tuned. Binary classifier head (Non-offensive/Offensive).

Intended use

Visualizing model attention (Grad-CAM) for transparency in a live demo only. Not intended for any actual classification use -- use the DistilBERT model above for that.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support