The authors analyze chain-of-thought monitoring as a function of task difficulty and model size.
Complex hidden reasoning leaves traces, though monitors may miss them
The authors establish when covert computation must reveal information in chain-of-thought traces, then show that this information can remain computationally unreadable.
Independent
Mohammadali Mohammadkhani · Madhava Krishna · Yash Sarrof · Michael Hahn
Research Digest··2 min read
Mohammadkhani and colleagues study whether reasoning models can solve covert tasks while concealing the relevant computation from chain-of-thought monitors.
Why this paper
Independent
In one line
Simple computation can be covert, but beyond a model-size-dependent threshold complex hidden computation necessarily leaks near-linear information into the chain of thought, though not necessarily readably.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§