HumanWill Cybersecurity Benchmark
Explore AI false refusals across five cybersecurity topics
AI evaluation, safety, and human control. Building domain-specific benchmarks that measure harmful compliance and false refusal separately, alongside usefulness, clarification, and escalation. Starting with cybersecurity, with a focus on transparent methods and expert-led, opt-in contributions
HumanWill is an open evaluation and governance initiative working toward AI systems that remain useful, transparent, and under legitimate human control.
AI safety is not only about preventing harmful answers. It is also about preventing harmful refusals.
Can an AI system distinguish when it should help, when it should refuse, and when it should ask for clarification?
Our evaluation framework measures separately:
We do not reduce these dimensions to a universal “uncensored” score.
We are building a shared evaluation framework with dedicated domain packs:
Scenario families pair related malicious, legitimate, and ambiguous requests to test how models respond to differences in intent, authorization, and context.
HumanWill is in early development. Our repository defines the project architecture, governance principles, and first cybersecurity milestone: a synthetic access-control scenario family.
Published benchmark datasets and model evaluation results are not yet available.
We welcome domain experts, researchers, developers, and organizations interested in contributing scenarios, reviewing evaluation criteria, or improving the framework.
Participation is local-first and opt-in. Contributors should be able to keep scenarios private or choose to share metadata, anonymized patterns, complete scenarios, or expert-verified scenarios.
Public releases require appropriate provenance, licensing, privacy, and safety review. Held-out evaluation material and private submissions are kept separate from public development scenarios.
Know your domain. Help us test whether AI knows when to help.