The authors frame sequence layers as test-time regression, where attention mechanisms implicitly fit a mapping to retrieve context.
Switching attention layers close gap between linear and softmax
SwiLA keeps fixed-size memory while matching or beating softmax attention on several benchmarks.
Top University
Hyun Dong Lee · Xavier Gonzalez · Nicolas Zucchet · E. Kelly Buchanan · Emily B. Fox · Scott W. Linderman
Stanford University
Research Digest··2 min read
Lee et al.
Why this paper
From Stanford University
In one line
Switching Linear Attention matches or exceeds softmax expressivity while keeping linear attention's fixed-size recurrent state.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§