Papers
arxiv:2605.00914

The Cost of Consensus: Isolated Self-Correction Prevails Over Unguided Homogeneous Multi-Agent Debate

Published on Apr 29
Authors:
,

Abstract

Multi-agent debate, where teams of LLMs iteratively exchange rationales and vote on answers, is widely deployed under the assumption that peer review filters hallucinations. Yet the failure dynamics of homogeneous debate remain poorly understood, therefore we report findings from a controlled empirical study of teams of N{=}10 homogeneous agents (Qwen2.5-7B, Llama-3.1-8B, Ministral-3-8B) across R{=}3 debate rounds on two high-difficulty benchmarks (GSM-Hard and MMLU-Hard). We compare peer debate against isolated self-correction and a stochastic noise control that injects rationales from unrelated problems. We decompose debate failure into three model-dependent pathways: sycophantic conformity, where agents uncritically adopt majority answers (modal adoption up to 85.5%); contextual fragility, where peer rationales destabilize previously correct reasoning (vulnerability rate up to 70.0%); and consensus collapse, where plurality voting discards correct answers already present in the generation pool (oracle gap up to 32.3 percentage points). Ablations over communication density (K in {2,4,9}) and sampling temperature (T in {0.4, 0.7}) show that conformity reaches high levels at minimal peer exposure (K{=}2) and intensifies with greater initial diversity. Across all configurations, debate consumes 2.1-3.4times more tokens (up to 28,631 tokens per problem) than self-correction for equal or lower accuracy. Our results indicate that, within the 7-8B parameter class, homogeneous teams without structured roles do not benefit from unguided peer exchange, and that isolated self-correction consistently offers a more favorable cost-accuracy tradeoff.

Community

This paper received the industry spotlight at ACM CAIS and was presented at AI Engineer World's Fair in San Francisco that attracts more than 6000 developers!

We built controlled experiments around teams of 10 LLM agents (Qwen2.5-7B, Llama-3.1-8B, Ministral-3-8B) debating hard math and STEM problems. The headline finding: debate burned 2.1–3.4× more tokens compared to self-correction and never outperformed it. Three compounding failures (sycophancy, contextual fragility, consensus collapse) flip multi-agent debate from "wisdom of the crowd" into expensive groupthink.

📅 CAIS 2026 · May 26–29 · DoubleTree by Hilton, San Jose
🌐 https://www.caisconf.org
📄 Paper: https://lnkd.in/dMenz6xe
💻 Code: https://lnkd.in/dq8zCkkK

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2605.00914
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2605.00914 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2605.00914 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2605.00914 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.