The authors formulated the AttenFeed module by adding a non-linear activation function to standard attention, building on an interpretation of attention as an FFN without activation.
Relaxing the Attention-FFN split in Vision Transformers can boost small model performance
The authors show that replacing alternating attention and feed-forward layers with a unified AttenFeed module improves accuracy at smaller scales, with the gap narrowing as models grow.
Academic
Junhyeok Kim · Jinyeong Kim · Jae Wan Park · Seong Jae Hwang
Yonsei University
Research Digest··2 min read
The authors introduce AttenFeed, a module designed to subsume the functions of both Attention and Feed-Forward Network (FFN) layers, and use it to build uViT, a Vision Transformer with no strict Attention-FFN separation.
Why this paper
From Yonsei University
In one line
In vision transformers, the rigid alternating Attention-FFN structure is not necessary and can impair performance at smaller model scales.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§