The authors analyze how training-inference mismatch and PPO clipping prevent a common offset in advantage estimates from cancelling, causing update drift.
Twin critics calibrate token-level advantages from a single rollout, boosting RL training efficiency
The proposed T5 method uses a conditional-moment saddle-point objective to combine two advantage estimates, achieving a 7.8% performance gain and up to 63.4% faster training steps compared to a critic-free baseline.
Tsinghua University · Tencent · Shanghai Jiao Tong University · Central South University · The Hong Kong University of Science and Technology
Why this paper
From Institute of Automation, Chinese Academy of Sciences and 7 others
In one line
Twin critics calibrate token-level advantages in reinforcement mid-training, improving mean benchmark performance by 7.8% and cutting training-step time by up to 63.4% versus critic-free methods.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.