Abstract
Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as agents move from deciding whether to act to executing single actions and dependent workflows. Across ten model-harness configurations, strong static action assessment can coexist with much weaker interactive execution. Failures often begin before execution: agents stop with incomplete investigation or act before required evidence is established. Once required evidence is obtained, single-action execution is usually reliable, while multi-action workflows additionally expose unresolved prerequisites and incomplete execution. For this analysis, we introduce SafeActBench, comprising 656 cases across six operational domains and five protocols that progress from static action judgment and investigated non-action to single- and multi-action workflows. A provenance-bound Evidence Ledger and deterministic trajectory evaluator track what information was established, when actions occurred, and whether downstream dependencies were satisfied. These results show that failures arise not only from missing information, but also from how agents use established evidence when deciding and executing actions.
Community
Tool-using agents can reach the right end state without having established the evidence that justified their actions. We introduce SafeActBench (656 cases, 6 operational domains, 5 protocols from static action judgment to dependency-constrained multi-action workflows) and evaluate 10 model–harness configurations to find where the evidence-to-action chain breaks.
Key findings:
- Static judgment ≠ execution: for three configurations re-evaluated on the same V1 cases, static accuracy is ≥95% while interactive success is ≤52%.
- Failures start before acting: agents stop before completing the required investigation in 21.7–62.9% of V0 episodes, and 37.0–66.9% of V1 action attempts occur before the required evidence is established.
- Missing evidence ≠ stopping: in controlled interventions on 43 V1 cases (three configurations), agents still act in 46.5–53.5% of completed episodes after one decisive record is withheld.
A provenance-bound Evidence Ledger and a deterministic trajectory evaluator (no LLM judge) check whether each consequential action was supported by evidence established beforehand.
interesting work!
Get this paper in your agent:
hf papers read 2610.07753 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper