Abstract
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.
Community
We release VeriHarness, an agentic verification framework that leverages disagreement resolution, consensus challenging, and evidence-backed revision to reliably improve long-horizon LLM agent performance without reference answers or grading rubrics.
The counterintuitive part is treating disagreement between rollouts as a reliability signal — that only holds if the rollouts are actually independent, and with one base model and one prompt they mostly aren't. Sample the same weights at temperature 0.8 five times and you're measuring decoding noise, not uncertainty; the variance collapses the moment you drop the temperature, which is a decent hint it was never epistemic. Real independence means different models, different prompts, different tool orderings — and then you have to say what the disagreement is even about. The other thing I'd want measured: does the verifier share the generator's blind spots? Same family, same training distribution, and it'll confidently pass exactly the failures the generator produces most fluently. I'd trust a false-negative rate on adversarial cases far more than an agreement rate on the easy ones.
Get this paper in your agent:
hf papers read 2610.00972 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper