The study starts from Production Sessions: 4,782 logged agent sessions of software engineers working in JetBrains IDEs.
Coding-agent benchmarks should mirror the task flows of their target users
A study of 4,782 real IDE sessions shows that users mix and switch task types, and the authors propose SWE-TaskFlow to calibrate benchmarks to measured interaction distributions.
Industry
Igor Slinko · Yaroslav Golubev · Sergey Titov
JetBrains Research
Research Digest··3 min read
The authors collected 4,782 agent sessions of software engineers in JetBrains IDEs and analyzed the long sessions with at least three user messages, which account for 33% of the sample.
Why this paper
From JetBrains Research
In one line
Coding-agent benchmarks should match the task flows of their target users.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§