The authors introduce test-time self-distillation, which uses two fixed counterfactual prompts rather than expert demonstrations.
Counterfactual prompts let beam search self-distill language models at inference
The method contrasts answer probabilities under excellent and poor reasoning prompts, then uses that signal to steer decoding without training.
Chinese Tech
Su Ee Tan · Xiaotong Ji · Rasul Tutunov · Haitham Bou-Ammar · Matthieu Zimmer
Huawei Noah’s Ark Lab · UCL Centre for AI
Research Digest··3 min read
Tan and colleagues translate self-distillation from a training procedure into an inference-time decoding method.
Why this paper
From Huawei Noah’s Ark Lab and UCL Centre for AI
In one line
Test-time self-distillation using counterfactual contexts and beam search improves language model generation without retraining.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§