The authors built FDE-Bench from hand-authored deployment scenarios covering Docker images, multi-service Compose stacks, and Kubernetes.
LLM agents deploy repairs better than systems built from scratch
FDE-Bench tests whether agents can produce replayable Docker, Compose, and Kubernetes configurations that build, become ready, behave correctly, and meet specifications.
Top University
Weihang Ding · Junfei Zhan · Yueting Li · Qirong Guo
University of California, Berkeley · Imperial College London · The Hong Kong University of Science and Technology (Guangzhou)
Research Digest··2 min read
Ding et al.
Why this paper
From University of California, Berkeley and 2 others
In one line
FDE-Bench reveals that readiness checks are the primary failure point for LLM agents in deployment tasks.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§