The authors start from the learning-dynamics framework of Ren & Sutherland (2025) but move beyond coarse example-level interaction to a token- and layer-wise view.
Continual learning's core challenges stem from a single evolving update-behavior interaction
The authors derive a token- and layer-wise decomposition that unifies data attribution, forgetting, and plasticity loss as distinct regimes of the same learning dynamics
Big Tech
Yi Ren · Wenlong Deng · Guanzhe Hong · Clare Lyle · Yarin Gal
University of Oxford · University of British Columbia · Google DeepMind
Research Digest··3 min read
Ren et al.
Why this paper
From Google DeepMind and 2 others
In one line
A token-level update-behavior interaction unifies data selection, forgetting mechanisms, and plasticity loss in continually trained language models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§