Agent workloads expose blind spots in standard LLM inference benchmarks

AgentPerfBench replays multi-turn workload distributions and profiles GPU kernels to reveal bottlenecks hidden by chatbot-style tests.

Top University
Cheuk Hang Lau · Zeyu Cao · Kevin Wong Cheuk Yin · Yao Lai · Haoran Wu · Nicholas D. Lane · +3 more

Imperial College London · University of Cambridge · University of Oxford

Research Digest··2 min read
Lau and colleagues built an inference benchmark from traces of agentic tasks, including SWE-Bench and TerminalBench, where requests span multiple turns and accumulate context.

The authors constructed workload profiles from empirical distributions of input length, output length and turn count in agent traces.

Why this paper

From Imperial College London and 2 others

In one line

Existing chat benchmarks miss multi-turn context growth and hardware saturation, so they underestimate performance requirements for agentic LLM inference.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (hardware H100)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.