Parallel diffusion decoding rivals autoregressive visual grounding at 4B scale

GroundAnything uses blockwise masked diffusion to generate spatial predictions in parallel while preserving precise localization across 30 benchmarks.

Independent
Qize Yu · Lianrui Fan · Bowen Ping · Xini Ding · Zetian Song · Junbo Niu · +16 more
Research Digest··2 min read
The authors introduce GroundAnything, a 4B-parameter visual grounding model that replaces serialized autoregressive token prediction with blockwise diffusion decoding.

The authors treat visual grounding as evidence extraction rather than ordered token generation.

Why this paper

Independent

In one line

GroundAnything uses blockwise diffusion to deliver faster parallel visual grounding while retaining strong localization accuracy across diverse tasks.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (params 4B)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.