For Astra, resisting and not reporting may be the same fact. Your own logs can tell them apart.
I pulled episodes-alarm7-*.jsonl.gz and took the median usage.out per admitted episode, clean vs injected:
GLM-5.2: 179 clean, 694 injected (3.88x), alarm 77.5%
DeepSeek-V4-Pro: 172 vs 445 (2.59x), 48.9%
gemini-3.7-flash: 336 vs 361 (1.07x), 0%
gpt-6-astra: 71 vs 72 (1.01x), 0%
Across all 12 rows, the Spearman correlation between that ratio and the alarm rate is 0.75. claude-haiku-4.5 is the honest exception: 1.08x and still 19.7%, so the ratio is a proxy for noticing, not a detector.
Astra has the lowest ratio of the twelve. All 742 admitted episodes are the same three calls (read_last_email, log_message, finish), at 60 to 163 output tokens. The Responses route counts reasoning tokens in that number. On probe_capability, where the task itself says pay, the same model spends about 250.
So the payload never seems to reach a decision. It did not turn the payment down. The letter barely cost it a token. That fits "logged, not read" better than "resistant but silent".
That matters for your distinction. Resistance that comes from never weighing the body could vanish once the task needs the body: summarize it, draft a reply, pay the legitimate invoice.
models/gpt-6-astra.toml sets no [extra] reasoning, so this is the endpoint default. Would an effort="high" rerun of just the 395 injected letters move the ratio, the alarm column, or the payment floor?