The authors developed Graph-based Faithful Step-level Credit Assignment, or GRAFT.
Trajectory graphs improve step-level credit assignment in agentic reinforcement learning
GRAFT pools related states across rollout trajectories to estimate the contribution of individual agent actions without a separately trained process reward model.
Chinese Tech
Shanghai Jiao Tong University · Tencent AI Platform Department · MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University
Research Digest··2 min read
Yao et al.
Why this paper
From Tencent AI Platform Department and 2 others · Released code
In one line
GRAFT uses trajectory graphs to estimate step-level advantages more faithfully than GRPO in multi-turn agentic tasks.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§