GUI agents sacrifice future options in long visual puzzles

LongPuzzleBench shows that agents often choose moves for immediate visible progress rather than preserving later solvability.

Chinese Tech
Bingo Zhang · Haochuan Lu · Zongjie Li · Genjian Li · Ari Yu Zhang · Chaozheng Wang

Vera Praxis · Tencent · The Hong Kong University of Science and Technology · Independent Researcher · The Chinese University of Hong Kong

Research Digest··2 min read
Zhang et al.

The authors created LongPuzzleBench, comprising 114 levels in six puzzle games, grouped into 16 objectives.

Why this paper

From Tencent and 4 others

In one line

GUI agents judge each move by visible progress, not future options, causing failure on long-horizon visual puzzles.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.