The authors extended the OSWorld computer-use benchmark with more than 300 tasks divided into over 2,800 subgoals.
Process-based evaluation reveals where computer-use agents go wrong
OSWorld-Pro tracks sequential subgoals rather than only final outcomes, exposing input errors, irrelevant actions, and incomplete progress.
Big Tech
Zhilin Wang · Shaokun Zhang · Yifan Zhang · Hao Zhang · Jin Xu · Binfeng Xu · +6 more
NVIDIA
Research Digest··2 min read
Thread:Process Agent Benchmarks
The authors introduce OSWorld-Pro, a benchmark that decomposes more than 300 computer-use tasks into over 2,800 sequentially dependent subgoals.
Why this paper
From NVIDIA
In one line
OSWorld-Pro evaluates computer-use agents on intermediate subgoals to reveal specific failure modes.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§