The authors trained Qwen3-8B and Qwen3-14B models using RLVR on mathematical reasoning tasks.
Companion model sharpens LLM confidence without harming task performance
The authors introduce CoCal, which trains a lightweight companion from rollout hidden states and verifier-derived correctness supervision while leaving task optimization unchanged.
Research Lab
Shiyu Ni · Keping Bi · Jiafeng Guo · Yilong Xu · Jingtong Wu · Zengxin Han · +1 more
State Key Laboratory of AI Safety · Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences
Research Digest··3 min read
The authors propose concurrent confidence calibration, where confidence is learned alongside capability improvement during reinforcement learning from verifiable rewards (RLVR).
Why this paper
From Institute of Computing Technology, Chinese Academy of Sciences and 2 others
In one line
CoCal trains a lightweight companion on rollout hidden states to calibrate LLM confidence without sacrificing task performance.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§