Process-based evaluation reveals where computer-use agents go wrong

OSWorld-Pro tracks sequential subgoals rather than only final outcomes, exposing input errors, irrelevant actions, and incomplete progress.

Big Tech
Zhilin Wang · Shaokun Zhang · Yifan Zhang · Hao Zhang · Jin Xu · Binfeng Xu · +6 more

NVIDIA

Research Digest··2 min read
The authors introduce OSWorld-Pro, a benchmark that decomposes more than 300 computer-use tasks into over 2,800 sequentially dependent subgoals.

The authors extended the OSWorld computer-use benchmark with more than 300 tasks divided into over 2,800 subgoals.

Why this paper

From NVIDIA

In one line

OSWorld-Pro evaluates computer-use agents on intermediate subgoals to reveal specific failure modes.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.