The authors augmented Independent Proximal Policy Optimization (IPPO) with two world models trained offline.
Counterfactual world models improve fully decentralized multi-agent learning
CASTLE uses frozen models of local dynamics and action-specific social consequences to guide agents without communication or a centralized critic.
Big Tech
Fernando Martinez · Tao Li · Yingdong Lu · Juntao Chen
Fordham University · City University of Hong Kong · IBM Research
Research Digest··2 min read
Martinez et al.
Why this paper
From IBM Research and 2 others
In one line
Agents learn more effectively by prospectively comparing consequences of candidate actions using counterfactual world models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§