Dependency-aware credit assignment improves reinforcement learning for terminal agents

DepGPO uses execution dependency graphs to redistribute advantages to commands that actually affect task verification.

Chinese Tech
Yu Li · Guangfeng Cai · Long-Fei Li · Shuo Han · Shengtian Yang · Han Luo · +2 more

Southeast University · Huawei Noah’s Ark Lab

Research Digest··3 min read
The authors propose Dependency-Aware Group Policy Optimization (DepGPO), which constructs command dependency graphs from execution traces and traces backward from the resources inspected by the task verifier.

The authors first conducted a pilot study on rollouts from a supervised fine-tuned model, showing that existing credit-assignment methods (like DAPO and GiGPO) often yield zero group-relative advantages or cannot distinguish command contributions.

Why this paper

From Huawei Noah’s Ark Lab and Southeast University

In one line

DepGPO assigns credit based on command read-write dependencies, improving RL training for terminal agents.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.