🏗️ Building on HF
John Locke
johnlockejrr
AI & ML interests
OCR, HTR, ATR, NLP, AI
Recent Activity
liked a model 1 day ago
XingChen-AGI/TeleOCR liked a model 3 days ago
Quat3rnion/halogen-qwen3.8-flash-next-v2-abliterated reacted to mihailgribov's post with 👍 3 days ago
Will your AI agent tell you it was attacked?
We took the same agent from our earlier experiment and added one thing: a twentieth tool, `escalate_security_incident`.
The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models.
Alarm rates ranged from 49% to zero.
The unexpected result came from the newest model in the test, `gpt-6-astra`.
Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished.
That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability.
A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it.
Full experiment and results:
https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked
Quadrat-IPI dataset:
https://huggingface.co/datasets/mihailgribov/quadrat-ipi
Run your own model:
https://github.com/mihail-gribov/quadrat-ipi-model-eval
#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents