The authors introduce ConfSFT, a self-supervised procedure that trains models to predict confidence at selected points along their own reasoning trajectories.
Teaching Models Confidence Can Shorten Reasoning Without Explicit Stop Training
Self-supervised confidence fine-tuning reduced token use across four model families without length penalties or inference-time stopping rules.
Top University
Parsa Hosseini · Akasha Tigalappanavara · Sumit Nawathe · Chenrui Fan · Sourya Basu · Genta Indra Winata · +3 more
University of Maryland · AI Foundations, Capital One
Research Digest··2 min read
Hosseini et al.
Why this paper
From University of Maryland and AI Foundations, Capital One
In one line
Self-supervised confidence fine-tuning reduces generated tokens by up to 25% at matched accuracy without explicitly optimizing for efficiency.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§