AI & ML interests

Mechanistic interpretability, LLM security, indirect prompt injection detection

No public activity