Papers
arxiv:2609.32754

Adaptive Consistency Graph for Long-Horizon Agents

Published on Sep 26
· Submitted by
jc
on Sep 29
Authors:
,
,

Abstract

Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long sequences of dependent actions and tool calls. During execution, task requirements, historical evidence, and the current execution state may gradually become disconnected, so later decisions can drift from the original objective. We study this problem by introducing the Adaptive Consistency Graph (ACG) for long-horizon execution. ACG incrementally organizes execution evidence and its provenance in a persistent graph, then constructs a temporary requirement-centered view for each decision under a bounded context budget. Rather than replacing the base agent's planner or tool executor, ACG provides a structured and traceable context view for each decision. In the matched evaluation, ACG improves GPT-5.6-luna's average success from 44.5\% with ReAct to 50.2\%, with the largest gain on BrowseComp-Plus (73.5\% versus 62.4\%). We further analyze trajectory structure and inference cost to characterize this improvement.

Community

How can AI agents go further on long tasks—and stay connected to their original goals?
As tasks grow longer, agents need to connect the original requirements, accumulated evidence, and current execution state. Even a locally reasonable decision can gradually lead execution away from the goal.
We introduce ACG (Adaptive Consistency Graph): a framework that organizes execution evidence and its sources into a persistent graph, then dynamically selects task-relevant information for each decision within a limited context budget. Past evidence remains traceable, helping ground the next action.
📊 In the paper’s matched evaluation, ACG improves GPT-5.6-luna’s equally weighted average success rate across three benchmarks from 44.5% to 50.2%. On BrowseComp-Plus, success rises from 62.4% to 73.5%—a gain of 11.1 percentage points.
Our research explores a central question: How can agents maintain the connection between goals, evidence, and actions throughout long-horizon execution?

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.32754
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.32754 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.32754 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.32754 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.