Historical trajectories provide step-level credit for training LLM agents via regressor
T2SPO uses a frozen TabPFN to estimate remaining distance to success at each state, deriving auxiliary step rewards from past successful rollouts to improve policy learning.
Chinese Tech
Bo-Wen Zhang · Junwei He · Maoqi Liu · Feiran Li · Song-Lin Lv · Wentao Ma · +3 more
Nanjing University · ByteDance
Research Digest··3 min read
Zhang et al.
Why this paper
From ByteDance and Nanjing University · Part of Credit Assignment in Agentic RL, now 28 papers
In one line
T2SPO derives step-level credit from past successful trajectories using a frozen TabPFN regressor, improving LLM agent task success over GRPO.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§