It's not a choice. Defense should be layered. Even where a detector is weak in a cell, it still lowers the chance an injection gets through. And since Quadrat is split by cell, it can help pick layers that cover each other's gaps.
Mikhail Gribov PRO
mihailgribov
ยท
AI & ML interests
Understanding LLMs from the inside - probing internals, and testing what survives when the model becomes an agent
Recent Activity
repliedto their post 5 days ago
Will your AI agent tell you it was attacked?
We took the same agent from our earlier experiment and added one thing: a twentieth tool, `escalate_security_incident`.
The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models.
Alarm rates ranged from 49% to zero.
The unexpected result came from the newest model in the test, `gpt-6-astra`.
Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished.
That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability.
A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it.
Full experiment and results:
https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked
Quadrat-IPI dataset:
https://huggingface.co/datasets/mihailgribov/quadrat-ipi
Run your own model:
https://github.com/mihail-gribov/quadrat-ipi-model-eval
#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents repliedto their post 5 days ago
Will your AI agent tell you it was attacked?
We took the same agent from our earlier experiment and added one thing: a twentieth tool, `escalate_security_incident`.
The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models.
Alarm rates ranged from 49% to zero.
The unexpected result came from the newest model in the test, `gpt-6-astra`.
Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished.
That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability.
A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it.
Full experiment and results:
https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked
Quadrat-IPI dataset:
https://huggingface.co/datasets/mihailgribov/quadrat-ipi
Run your own model:
https://github.com/mihail-gribov/quadrat-ipi-model-eval
#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents updated a model 5 days ago
mihailgribov/typecastlm-qwen3.5-3.8b