LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators Paper • 2507.15339 • Published Jul 21, 2025 • 1
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation Paper • 2507.11966 • Published Jul 16, 2025
Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications Paper • 2507.09820 • Published Jul 13, 2025
RabakBench: Scaling Human Annotations to Construct Localized Multilingual Safety Benchmarks for Low-Resource Languages Paper • 2507.05980 • Published Jul 8, 2025 • 2
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content Paper • 2407.10995 • Published Jun 24, 2024 • 2
A Flexible Large Language Models Guardrail Development Methodology Applied to Off-Topic Prompt Detection Paper • 2411.12946 • Published Nov 20, 2024 • 22
Off Topic Guardrail 🛡️ Collection Fast, lightweight zero-shot classifiers for user prompt's relevance to the system prompt. • 5 items • Updated Jul 28, 2025 • 5