Papers
arxiv:2609.36104

An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures

Published on Sep 28
Authors:
,

Abstract

Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task-averaged number conceals. Proposer-Critic captures the arithmetic gains, scaling steepest and, in aggregate, surpassing every other architecture at the largest budget (item-clustered intervals exclude zero), though it ranks among the weakest elsewhere, and no architecture wins across tasks. We explain these trajectories with an exact generate-transform decomposition. Partitioning any workflow into proposal coverage and a downstream transform, any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. The decomposition diagnoses each task: arithmetic offers coverage headroom that a critic-guided transform converts, whereas the multiple-choice benchmarks either saturate in coverage or fail to convert it, and on open-ended code generative recovery nearly vanishes so accuracy tracks coverage. At equal call budgets token cost still varies 2.1x. Extra calls therefore create candidate opportunity that only some architectures, on some tasks, convert. Team scaling is a task- and architecture-specific bet, not a uniform lever.

Community

👋 Author here. TL;DR: scaling an agent team from 3 to 30 LLM calls helps a lot on some tasks and barely at all on others, and no orchestration architecture wins everywhere. We explain why with an exact generate–transform decomposition.

Setup: 8 orchestration architectures × five instruction-tuned 7–9B models × five short-answer benchmarks + an executable-code benchmark, up to 30 calls.

Findings

  • From 3 → 30 calls, accuracy rises by up to 17 points on GSM8K / GSM-Hard but at most 4 on ARC, GPQA and MMLU, for every architecture.
  • Proposer–Critic scales steepest on arithmetic and, in aggregate, surpasses every other architecture at the largest budget (item-clustered intervals exclude zero), but ranks among the weakest elsewhere.
  • Decomposition: partition any workflow into proposal coverage and a downstream transform; any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. Arithmetic offers coverage headroom a critic-guided transform converts; MCQ either saturates coverage or fails to convert it; on code, generative recovery nearly vanishes and accuracy tracks coverage.
  • At equal call budgets, token cost still varies 2.1×.

Happy to discuss the decomposition or how to apply it to your own workflow!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.36104
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.36104 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.36104 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.36104 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.