AgentHorizon is constructed from human-recorded trajectories for pairs of similar instructions.
New benchmark tests AI judges on long computer-use tasks by swapping instructions
AgentHorizon pairs human-recorded trajectories with similar but incompatible instructions, revealing that even the best judge reaches only 80.9% balanced accuracy.
Top University
Xing Han Lù · Dheeraj Vattikonda · Sina Hajimiri · Fatemeh Pesaran Zadeh · Parishad BehnamGhader · Ghazwa Darwiche · +4 more
ServiceNow Research · McGill University · Mila – Quebec AI Institute · ÉTS Montréal · Seoul National University
Research Digest··3 min read
The authors introduce AgentHorizon, a benchmark of 1,373 instruction-trajectory pairs built from 166 hours of human-recorded computer-use sessions across three operating systems.
Why this paper
From McGill University and 7 others
In one line
A benchmark of 1,373 paired computer-use tasks shows current judges, best GPT-5.5 at 80.9% balanced accuracy, often miss subtle instruction violations.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§