Activation steering strengthens token-level supervision for self-distilled reasoning

A frozen teacher guided by activation differences between successful and unsuccessful trajectories outperformed reference-conditioned self-distillation across the evaluated reasoning tasks.

Big Tech
Zhexi Lu · Subhajit Chaudhury · Tejaswini Pedapati · Keerthiram Murugesan · Lei Yu

Rensselaer Polytechnic Institute · IBM Research

Research Digest··2 min read
Lu et al.

The authors sampled multiple reasoning trajectories from each base model and verified their final answers.

Why this paper

From IBM Research and Rensselaer Polytechnic Institute

In one line

Activation-Conditioned Self-Distillation steers a frozen self-teacher using contrasts from verified correct trajectories, improving math and code accuracy beyond reference-conditioned self-distillation.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.