Coding-agent benchmarks should mirror the task flows of their target users

A study of 4,782 real IDE sessions shows that users mix and switch task types, and the authors propose SWE-TaskFlow to calibrate benchmarks to measured interaction distributions.

Industry
Igor Slinko · Yaroslav Golubev · Sergey Titov

JetBrains Research

Research Digest··3 min read
The authors collected 4,782 agent sessions of software engineers in JetBrains IDEs and analyzed the long sessions with at least three user messages, which account for 33% of the sample.

The study starts from Production Sessions: 4,782 logged agent sessions of software engineers working in JetBrains IDEs.

Why this paper

From JetBrains Research

In one line

Coding-agent benchmarks should match the task flows of their target users.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe