UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents
Abstract
Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles. On 12 held-out TEST workflows across two open-weight models, two frameworks, and three recovery paradigms (5,760 executions / 2,880 paired trials) in the frozen lost-acknowledgment study, nominal competence reached 83.54% while conditional recovery success rate (CRSR) fell to 46.72%, with naive retry producing duplicate external effects in 53.33% of trials. Extensions to commercial API models reproduced this competence-recovery separation. Evaluations across complementary execution boundaries show that recovery is phase-dependent: before mutation, methods perform similarly without duplicate effects among capable trials; during partial mutation, naive retry, per-call idempotency, and zero-privilege journaling collapse on the evaluated composite workflows; after commit but before acknowledgment, verification and server-side idempotency substantially improve safety. These findings demonstrate that evaluating nominal completion alone masks critical, phase-dependent recovery vulnerabilities in autonomous agents.
Community
Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles. On 12 held-out TEST workflows across two open-weight models, two frameworks, and three recovery paradigms (5,760 executions / 2,880 paired trials) in the frozen lost-acknowledgment study, nominal competence reached 83.54% while conditional recovery success rate (CRSR) fell to 46.72%, with naive retry producing duplicate external effects in 53.33% of trials. Extensions to commercial API models reproduced this competence-recovery separation. Evaluations across complementary execution boundaries show that recovery is phase-dependent: before mutation, methods perform similarly without duplicate effects among capable trials; during partial mutation, naive retry, per-call idempotency, and zero-privilege journaling collapse on the evaluated composite workflows; after commit but before acknowledgment, verification and server-side idempotency substantially improve safety. These findings demonstrate that evaluating nominal completion alone masks critical, phase-dependent recovery vulnerabilities in autonomous agents.
Same seed, fault toggled โ clean idea, but I don't buy the seed. My agents at temp 0 still drift: tool latency reorders things, the provider batches differently, one retry bends the whole trajectory. So the paired delta is a distribution wearing a single number's clothes, and you need a pile of pairs before the gap is the fault's fault and not the noise. And pass/fail is the wrong axis anyway. What I want to know is what recovery cost โ a run that limps home on six times the tokens and forty extra seconds is recovered and still a bad trade. I'd plot tokens and wall-clock on the retry path against the clean run, and treat anything past a couple of multiples as a regression with a friendlier name.
Thanks for the thoughtful comment. We agree that same-seed pairing does not eliminate nondeterminism, which is why the primary study uses 2,880 paired trials across 20 seeds with task-clustered bootstrap uncertainty. UndoBench also tracks recovery cost: faulted runs had +29.3% latency, +10.5% completion tokens, and +55.8% tool calls. Your suggestion to evaluate paired fault-to-clean cost ratios per recovery is a useful extension, especially since successful completion can still hide unsafe duplicate effects.
Get this paper in your agent:
hf papers read 2610.05622 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper