Shared rollout states improve credit assignment for long-horizon agents

CRBC recursively combines evidence from intersecting rollouts while weighting alternative actions and transitions by their observed frequencies.

Top University
Yangyang Ren · Haodong Zhu · Linlin Yang · Sheng Xu · Peichao Lai · Baochang Zhang

Beihang University · Zhongguancun Academy · Communication University of China · Peking University · Hangzhou Innovation Institute of Beihang University

Research Digest··2 min read
Ren et al.

The authors introduce Cross-Rollout Bellman Closure, or CRBC, for assigning credit to individual actions in long-horizon language-model agents.

Why this paper

From Beihang University and 4 others · Part of Credit Assignment in Agentic RL, now 30 papers

In one line

Cross-Rollout Bellman Closure improves long-horizon agentic RL by merging rollouts and computing a Bellman fixed point for step-level credit.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.