The authors treat visual grounding as evidence extraction rather than ordered token generation.
Parallel diffusion decoding rivals autoregressive visual grounding at 4B scale
GroundAnything uses blockwise masked diffusion to generate spatial predictions in parallel while preserving precise localization across 30 benchmarks.
Independent
Qize Yu · Lianrui Fan · Bowen Ping · Xini Ding · Zetian Song · Junbo Niu · +16 more
Research Digest··2 min read
The authors introduce GroundAnything, a 4B-parameter visual grounding model that replaces serialized autoregressive token prediction with blockwise diffusion decoding.
Why this paper
Independent
In one line
GroundAnything uses blockwise diffusion to deliver faster parallel visual grounding while retaining strong localization accuracy across diverse tasks.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 4B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§