The authors derived 53 repository-grounded tasks from recent SGLang changes, covering six families of inference engineering work across model support, runtime execution, and public APIs.
AI coding agents still miss production inference failures
SWE-Serve tests repository-scale inference engineering and finds that end-to-end serving checks reject many patches that pass narrower tests.
Big Tech
Jennifer Williams · Dave Farris · Jeff Farris · Jiantao Jiao
NVIDIA · University of California, Berkeley
Research Digest··2 min read
The authors built 53 tasks from recent production changes to the SGLang inference-serving system and evaluated 11 models across 31 model-effort configurations.
Why this paper
From NVIDIA and University of California, Berkeley
In one line
SWE-Serve reveals that AI agents often pass local tests but fail production end-to-end tests for inference serving tasks.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§