Abstract
Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library. Downstream agents adopt agent- and human-written libraries alike but underuse them, reimplementing capabilities the library already provides. Our failure analysis finds that downstream agents write extra code mainly because agent-written libraries are rigid or hard to use, not because capabilities are missing. We also experiment with giving designers more prescriptive, agent-first guidance and having them test their library with subagents; this improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.
Community
We are fast approaching a point where the primary users of software libraries are agents, not human engineers. If agents are the users, then a library should be judged by how little code agents need to write correct programs with it, not by how it reads to a human. That is why I am excited to release LibraryDesignBench, a two-phase benchmark that scores an agent-written library solely by how much it helps the future agents that use it.
Asking agents to write libraries for other agents
LibraryDesignBench asks agents to write effective libraries, and we intentionally give them design flexibility. We never prescribe signatures their library must expose. Nor do we ever guide them to specific abstractions or patterns they should use. We intentionally give them underspecified, vague, and ambiguous library specifications because frontier agents need to envision how future agents will use their library, what they will need, and what design will work best. LibraryDesignBench gives agents this vast freedom because the evaluation needs to serve as a test bed for understanding what patterns agents actually prefer, not what we think they will prefer.
Evaluating A Library Based on How Agents Use It
A library that implements something correctly does not immediately provide any value – it needs to be correct and make future code simpler. Thus, the only way to measure a library's quality is to observe how much future agents benefit from its design decisions. LibraryDesignBench does this in two phases:
- Design Phase: The agent implements a full library from an intentionally non-prescriptive specification.
- Evaluation Phase: We evaluate the library by observing multiple different implementer agents attempt to solve problems using it.
The only aspect we check in the design phase is if the library is installable in their internet-restricted environment, as we pre-install all libraries for evaluation since trying to install a library is not part of the signal we care about. We then task three agents (5.6 Luna on Codex, DeepSeek v4.1 Flash on mini-SWE-agent, and GLM 5.3 Flash on mini-SWE-agent) with solving programming problems with the library in as little code as possible. Each problem $\times$ agent $\times$ library is run in its own environment without internet access, with high reasoning, and with the library pre-installed. We use a highly prescriptive prompt to ensure agents attempt to fully exploit the library.
We now take these solutions and score each with:
Pass rate is the percentage of tests passed, and we square it to penalize incorrectness more harshly than simplicity. Simplicity is defined as:
$\mathcal{M}$ is the set of static metrics we use to compare how far off the solution written with the library is from the golden reference, $y^*$, written with the real production library. Each ratio is capped at 1, so a solution that beats the reference gets no extra credit. The four metrics are:
- Source Lines of Code: how much code the agent wrote.
- Cyclomatic Complexity: how many branches and loops the agent needed.
- Cognitive Complexity: how hard the code is to follow, with extra weight on nesting.
- Halstead Volume: size in operators and operands, which dense one-liners can't game.
Agents Copy Human Designs But Worse
We evaluate 11 designer setups, covering 9 frontier models with some run in more than one harness, all at high reasoning. Each setup attempts each of the 15 library design tasks three times, which produces 45 libraries. The three implementer agents then use each library to solve every problem in its task. Across all 242 problems, that comes to 242 problems × 3 libraries × 3 implementers = 2,178 trials per setup.
We compare against two baselines that use the same 2,178 trials. In the first, implementers have no library and use a separate prompt. In the second, they have the human-written production library pre-installed. Neither baseline has three designed libraries to vary, so we instead run each problem-implementer pair three times.
Opus 5.5 designs libraries that help downstream agents more than the human-written production libraries do. With them, implementers pass just as many tests while writing simpler code. Fable 5.1 roughly matches production. Correctness barely separates designers, since every setup passes about the same share of tests, so the ranking comes down to how much code implementers still have to write. At the other end, DeepSeek V4 Pro's library actively hurts: implementers do worse with it than with no library at all. The harness also matters. Fable's libraries are much more useful when designed in mini-SWE-agent than in Claude Code, and Astra's are slightly more useful in mini-SWE-agent than in Codex.
Learn More
The full leaderboards and tasks are at ldbench.com. You can view the full repo here and the task repo here. You can read the full technical report here. If you have a task you would like to see, open an issue on the repo or get involved in the Discord.
This would not be possible without my amazing collaborators: Alex L. Zhang, Avi Trost, Vincent Sunn Chen, Frederic Sala, Aws Albarghouthi, and Ludwig Schmidt. I also thank John Yang, Parth Asawa, Xavier Garcia, Ryan Carelli, Arun Kumar, Floriad Brand, and Nick Roberts for their helpful feedback and discussions. LDB is supported by DARPA, NSF, Prime Intellect, and Snorkel AI through the Open Benchmarks Grant.
Get this paper in your agent:
hf papers read 2609.36730 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper