The authors examined six pretrained decoder-only language models from the LLaMA, Qwen, DeepSeek and Mistral families, spanning 1 billion to 32 billion parameters.
One Hidden Channel Gates Extreme Activations in Large Language Models
Across six decoder-only models, the authors identify a fixed embedding channel that determines whether an early feed-forward network produces unusually large activations.
Top University
Minjia Mao · Shi Chen · Bowen Yin · Xiao Fang
University of Delaware · Peking University
Research Digest··3 min read
Mao and colleagues trace massive activations, exceptionally large values in a few hidden dimensions, to a single input channel of an early feed-forward network.
Why this paper
From Peking University and University of Delaware
In one line
A single fixed channel in an early feed-forward network's input embedding decides whether massive activations appear in large language models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (3 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§