AI coding agents still miss production inference failures

SWE-Serve tests repository-scale inference engineering and finds that end-to-end serving checks reject many patches that pass narrower tests.

Big Tech
Jennifer Williams · Dave Farris · Jeff Farris · Jiantao Jiao

NVIDIA · University of California, Berkeley

Research Digest··2 min read
The authors built 53 tasks from recent production changes to the SGLang inference-serving system and evaluated 11 models across 31 model-effort configurations.

The authors derived 53 repository-grounded tasks from recent SGLang changes, covering six families of inference engineering work across model support, runtime execution, and public APIs.

Why this paper

From NVIDIA and University of California, Berkeley

In one line

SWE-Serve reveals that AI agents often pass local tests but fail production end-to-end tests for inference serving tasks.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.

AI coding agents still miss production inference failures | Zotpaper