The authors developed Counterfactual Constraint-Conditioned On-Policy Distillation, or CC-OPD.
Ablating individual constraints improves distillation for complex instruction following
CC-OPD derives token-level training rewards by measuring how a frozen teacher's predictions change when each instruction constraint is removed.
Chinese Tech
Yanzhao Zheng · Yuanqiang Yu · Tianze Xu · Chao Ma · Zhentao Zhang · Jihuai Zhu · +3 more
Alibaba Group
Research Digest··2 min read
The authors introduce a form of on-policy distillation designed for prompts containing many simultaneous constraints.
Why this paper
From Alibaba Group
In one line
CC-OPD improves multi-constraint instruction following by using per-constraint probability differentials as token-level rewards.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§