Companion model sharpens LLM confidence without harming task performance

The authors introduce CoCal, which trains a lightweight companion from rollout hidden states and verifier-derived correctness supervision while leaving task optimization unchanged.

Research Lab
Shiyu Ni · Keping Bi · Jiafeng Guo · Yilong Xu · Jingtong Wu · Zengxin Han · +1 more

State Key Laboratory of AI Safety · Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences

Research Digest··3 min read
The authors propose concurrent confidence calibration, where confidence is learned alongside capability improvement during reinforcement learning from verifiable rewards (RLVR).

The authors trained Qwen3-8B and Qwen3-14B models using RLVR on mathematical reasoning tasks.

Why this paper

From Institute of Computing Technology, Chinese Academy of Sciences and 2 others

In one line

CoCal trains a lightweight companion on rollout hidden states to calibrate LLM confidence without sacrificing task performance.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.