The authors developed one model family for UI grounding—locating interface elements from visual input—multi-step navigation on mobile, desktop, and web platforms, and visual tool use.
A single visual agent spans mobile, desktop, web, and tool use
MintAct combines interface grounding and multi-step action across several digital environments while matching similarly sized domain-specific agents.
Big Tech
Mingfei Gao · Rui Tian · Haiming Gang · Bohan Zhai · Le Zhang · Yuanzheng Gong · +8 more
Apple
Research Digest··2 min read
Gao and colleagues trained MintAct vision-language agents at 2B, 4B, and 8B parameter scales using infrastructure that runs hundreds of heterogeneous digital environments concurrently.
Why this paper
From Apple · Part of Cross-Device Agent Benchmarks, now 3 papers
In one line
MintAct is a family of vision-language models that unifies UI grounding, multi-step navigation, and visual tool use across mobile, desktop, and web at 2B-8B scales.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§