leukas/hippo_superset
Viewer β’ Updated β’ 738k β’ 5
This repository contains our 4B model for HIPPO: Harmful Input, Positive and Productive Output. This is part of our work, From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation, where we show that instruction-tuning an LLM for hate speech can dramatically improve its performance on hate speech-related tasks. We've collected 36 different hate speech datasets and converted them into a conversational template. We used it to train our HIPPO 4B and 32B models, which outperform several task-specific models.
Base model
Qwen/Qwen3-4B-Instruct-2507