The authors first conducted a pilot study on rollouts from a supervised fine-tuned model, showing that existing credit-assignment methods (like DAPO and GiGPO) often yield zero group-relative advantages or cannot distinguish command contributions.
Dependency-aware credit assignment improves reinforcement learning for terminal agents
DepGPO uses execution dependency graphs to redistribute advantages to commands that actually affect task verification.
Chinese Tech
Yu Li · Guangfeng Cai · Long-Fei Li · Shuo Han · Shengtian Yang · Han Luo · +2 more
Southeast University · Huawei Noah’s Ark Lab
Research Digest··3 min read
The authors propose Dependency-Aware Group Policy Optimization (DepGPO), which constructs command dependency graphs from execution traces and traces backward from the resources inspected by the task verifier.
Why this paper
From Huawei Noah’s Ark Lab and Southeast University
In one line
DepGPO assigns credit based on command read-write dependencies, improving RL training for terminal agents.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§