The authors investigated reinforcement learning with verifiable rewards (RLVR), where models are trained to produce outputs that can be automatically checked for correctness.
RL reasoning gains live in a low-dimensional, low-variance activation subspace
By steering activations with trainable vectors, the authors map where RLVR improvements reside and use that geometry to stabilize training for 2,000 steps.
Chinese Tech
Yuchen Cai · Ding Cao · Qixiang Yin · Xin Xu · Kai Yang · Siye Wu · +7 more
USTC · Tencent Hunyuan · BUPT
Research Digest··3 min read
Cai et al.
Why this paper
From Tencent Hunyuan and 2 others · Released code
In one line
RLVR-induced gains in LLMs are explained by a low-dimensional activation manifold where effective control directions lie in the low-variance complement of the principal subspace.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§