Chunk-level sparse autoencoders discover more reliable semantic features than token-level counterparts

By encoding mean-pooled activations over chunks and using reconstruction and neighbor prediction, the method improves high-level feature discovery, reasoning detection, and steering.

Chinese Tech
Xu Wang · Yifan Yang · TingHao YU · Difan Zou

The University of Hong Kong · Tencent

Research Digest··2 min read
The authors introduce chunk-level sparse autoencoders (SAEs) that encode mean-pooled activations over contiguous token chunks instead of individual tokens.

Wang et al.

Why this paper

From Tencent and The University of Hong Kong

In one line

Chunk-level sparse autoencoders trained on mean-pooled activations learn more reliable semantic features than token-level SAEs.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.