Papers
arxiv:2609.30541

AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework

Published on Sep 24
· Submitted by
Aparajith
on Oct 8
Authors:
,
,
,

Abstract

Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and retains modifications that improve a held-out scalar metric -- to automate this exploration. We report on twelve weeks of running this paradigm at production scale, where iterations consume hours of multi-GPU compute, evaluation involves competing criteria, and campaigns span weeks across many training jobs. Across two independently developed representation-learning systems for a book recommendation pipeline, we ran 220+ experiments and observed five recurring failure modes absent from the original setting: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation. We contribute a three-principle scaffolding design -- prevent, persist, redirect -- that maps each failure mode to a structural remedy and whose instantiation scales with iteration cost. The framework produced a 1.82x Recall@6 lift and a 2.1x coherence lift over hand-tuned baselines, and the agent autonomously designed a text-only fallback that expanded catalog coverage by 5.8x. The two systems span nearly three orders of magnitude in per-iteration cost yet exhibit the same failure modes, suggesting these are structural properties of production-scale autonomous research rather than artifacts of either application.

Community

Paper author Paper submitter
•
edited about 21 hours ago

Excited to share AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework, accepted at IEEE ICDM 2026!

What breaks when an LLM agent runs ML research at production cost? We ran Karpathy's AutoResearch loop for 12 weeks, 220+ experiments, on two very different systems at Amazon Books. One learns retrieval
embeddings (9–19 hours per iteration on multi-GPU). The other learns hierarchical Semantic IDs with RQ-VAE (6–52 minutes on one GPU). Their per-iteration cost differs by ~760×, yet they hit the same
failure modes.

✨ Key highlights:
*A taxonomy of 5 failure modes: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation
*A prevent / persist / redirect scaffold, built for System A as a three-agent extension (Researcher, Code Fixer, Criticizer)
*1.82× Recall@6 over the hand-tuned baseline, and 2.1× weighted coherence on Semantic IDs
*The agent designed its own text-only fallback and grew catalog coverage 5.8×, with no human direction
*The Code Fixer fixed 14 latent bugs in its first pass, including a multi-GPU bug that kept an 8-GPU job on one device for 10+ hours
*Cost decides the design: when iterations are expensive, prevent. When they are cheap, revert.

📄 Paper: https://arxiv.org/abs/2609.30541

Would love to hear other's thoughts!

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.30541
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.30541 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.30541 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.30541 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.