The authors propose a diagnostic, effective depth (D_eff), which converts the layer-wise residual stream into a similarity autocorrelation profile and aggregates it into one number.
Effective depth shows transformer layers are redundantly correlated
A new scalar diagnostic aggregates layer-wise similarity and finds 15 of 16 language models below a structural reference, indicating correlated updates rather than unused depth.
Big Tech
Barak Gahtan · Ido Galil · Alex M. Bronstein
Technion Israel Institute of Technology · Nvidia · ISTA Institute of Science and Technology Austria
Research Digest··3 min read
Gahtan et al.
Why this paper
From Nvidia and 2 others
In one line
Transformer residual streams have an effective depth below 2 layers, regardless of their nominal layer count.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§