A single visual agent spans mobile, desktop, web, and tool use

MintAct combines interface grounding and multi-step action across several digital environments while matching similarly sized domain-specific agents.

Big Tech
Mingfei Gao · Rui Tian · Haiming Gang · Bohan Zhai · Le Zhang · Yuanzheng Gong · +8 more

Apple

Research Digest··2 min read
Gao and colleagues trained MintAct vision-language agents at 2B, 4B, and 8B parameter scales using infrastructure that runs hundreds of heterogeneous digital environments concurrently.

The authors developed one model family for UI grounding—locating interface elements from visual input—multi-step navigation on mobile, desktop, and web platforms, and visual tool use.

Why this paper

From Apple · Part of Cross-Device Agent Benchmarks, now 3 papers

In one line

MintAct is a family of vision-language models that unifies UI grounding, multi-step navigation, and visual tool use across mobile, desktop, and web at 2B-8B scales.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.