Production-scale AutoResearch exposes five failure modes and a three-principle fix

Twelve weeks of running Karpathy's iterative LLM search across two Amazon book recommendation systems yielded 1.82x and 2.1x metric lifts, while revealing structural failure modes.

Big Tech

Amazon

Research Digest··3 min read
The authors applied AutoResearch, an LLM that iteratively edits training scripts and keeps changes that improve a held-out metric, to two production embedding systems at Amazon.

The authors ran Andrej Karpathy's AutoResearch paradigm in two independently developed representation-learning systems for Amazon's book recommendation pipeline.

Why this paper

From Amazon

In one line

AutoResearch at production scale reveals five failure modes fixed by a prevent-persist-redirect multi-agent framework.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.