leukas/hippo_superset
Viewer • Updated • 738k • 9
This repository contains our 32B model for HIPPO: Harmful Input, Positive and Productive Output. This is part of our work, From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation, where we show that instruction-tuning an LLM for hate speech can dramatically improve its performance on hate speech-related tasks. We've collected 36 different hate speech datasets and converted them into a conversational template. We used it to train our HIPPO 4B and 32B models, which outperform several task-specific models.
Base model
Qwen/Qwen3-32B