The author toggled Qwen3 between thinking and non-thinking modes while keeping its weights fixed, then compared answer diversity and the probability that two wrong samples produced the same answer.
Reasoning makes language models repeat the same wrong answers
With model weights held fixed, enabling reasoning increased agreement among incorrect samples and weakened the premise behind self-consistency voting.
Academic
Asaad Althoubi
Oklahoma State University
Research Digest··2 min read
Althoubi compared reasoning and non-reasoning modes across five benchmarks using 74,944 samples, isolating the effect of reasoning while holding model weights fixed.
Why this paper
From Oklahoma State University
In one line
Reasoning makes independently sampled wrong answers converge, undermining self-consistency, while confidence-weighted voting fails to improve on plain majority voting.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§