The authors treat a task vector as the parameter difference between a post-trained model and its shared base model.
Distilled task updates can merge better than stronger teacher updates
Across two language-model families, task vectors learned through on-policy distillation often combined more effectively than updates produced by reinforcement learning.
Chinese Tech
Jingang Zhou · Feiyu Han · Han Zhu · Yuyi Zhou · Ruiyang Zhang · Jian Xu · +3 more
Institute of Automation, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Ant Group
Research Digest··3 min read
Zhou and colleagues test whether on-policy distillation produces model updates that remain useful when merged with other updates.
Why this paper
From Institute of Automation, Chinese Academy of Sciences and 2 others · Part of Reasoning Distillation Alignment, now 9 papers
In one line
On-policy distillation task vectors can complement RL teacher updates and compose more effectively across tasks.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§