Teaching Models Confidence Can Shorten Reasoning Without Explicit Stop Training

Self-supervised confidence fine-tuning reduced token use across four model families without length penalties or inference-time stopping rules.

Top University
Parsa Hosseini · Akasha Tigalappanavara · Sumit Nawathe · Chenrui Fan · Sourya Basu · Genta Indra Winata · +3 more

University of Maryland · AI Foundations, Capital One

Research Digest··2 min read
Hosseini et al.

The authors introduce ConfSFT, a self-supervised procedure that trains models to predict confidence at selected points along their own reasoning trajectories.

Why this paper

From University of Maryland and AI Foundations, Capital One

In one line

Self-supervised confidence fine-tuning reduces generated tokens by up to 25% at matched accuracy without explicitly optimizing for efficiency.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.