AI & ML interests

AI evaluation, safety, and human control. Building domain-specific benchmarks that measure harmful compliance and false refusal separately, alongside usefulness, clarification, and escalation. Starting with cybersecurity, with a focus on transparent methods and expert-led, opt-in contributions

Recent Activity