The authors propose WISE (witness-integrated set effects), which computes set-level causal effects by averaging over component subsets while using witness pinning to hold downstream activations at clean values, preventing self-repair from outside the set.
New method discovers language model circuits while accounting for component interactions
JuntaLearner, based on witness-integrated set effects, achieves higher circuit recognition scores across multiple models and tasks compared to attribution baselines.
Top University
Sankaran Vaidyanathan · Rafal Urbaniak · Emily Bunnapradist · Michelangelo Naim · Daniel Waxman
Basis Research Institute · University of Massachusetts Amherst · MIT
Research Digest··3 min read
The authors introduce witness-integrated set effects (WISE), a family of causal estimands that average effects over component sets while blocking self-repair, and JuntaLearner, a gradient-based circuit discovery method that learns to rank components by their causal impact across varying set sizes.
Why this paper
From MIT and 2 others
In one line
WISE and JuntaLearner discover language model circuits that account for interactions between components.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§