The authors treat an intervention as a signed dose, λ, along an axis representing the strength of a counterfactual contrast.
Apparent model self-repair follows pre-existing counterweight responses
Across several language models and tasks, downstream components responded to graded interventions according to a simple affine relationship.
Industry
Areeb Ahmad · Pratinav Seth · Vinay Kumar Sankarapu
Lexsi Labs
Research Digest··3 min read
Ahmad, Seth and Sankarapu reinterpret apparent self-repair after model ablations as the ordinary action of counterweights already embedded in a network.
Why this paper
From Lexsi Labs
In one line
Self-repair in language models is a fixed counterweight that responds linearly to the strength of the removed signal.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§