Wang et al.
Chunk-level sparse autoencoders discover more reliable semantic features than token-level counterparts
By encoding mean-pooled activations over chunks and using reconstruction and neighbor prediction, the method improves high-level feature discovery, reasoning detection, and steering.
Chinese Tech
Xu Wang · Yifan Yang · TingHao YU · Difan Zou
The University of Hong Kong · Tencent
Research Digest··2 min read
The authors introduce chunk-level sparse autoencoders (SAEs) that encode mean-pooled activations over contiguous token chunks instead of individual tokens.
Why this paper
From Tencent and The University of Hong Kong
In one line
Chunk-level sparse autoencoders trained on mean-pooled activations learn more reliable semantic features than token-level SAEs.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§