Starting from a 2B-parameter VLA policy (OpenVLA with a flow-matching action expert), the authors added a lightweight affordance head that predicts pixel-wise interaction regions from intermediate backbone features.
Wiring details determine whether affordance heads help or hurt VLA policies
A stop-gradient and an actively initialized residual bridge lift LIBERO success from 93.1% to 96.2%, matching far more complex architectures.
Big Tech
Zijian An · Linhan Wang · Jiayan Wang · Shijie Geng · Ran Yang · Yiming Feng · +1 more
Drexel University · Virginia Tech · Amazon Store Foundation AI
Research Digest··3 min read
The authors systematically ablate how an affordance head is injected into a vision-language-action policy.
Why this paper
From Amazon Store Foundation AI and 2 others
In one line
Stop-gradient affordance head with active initialized residual bridge boosts VLA policy success to 96.2% on LIBERO.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§