Counterfactual world models improve fully decentralized multi-agent learning

CASTLE uses frozen models of local dynamics and action-specific social consequences to guide agents without communication or a centralized critic.

Big Tech
Fernando Martinez · Tao Li · Yingdong Lu · Juntao Chen

Fordham University · City University of Hong Kong · IBM Research

Research Digest··2 min read
Martinez et al.

The authors augmented Independent Proximal Policy Optimization (IPPO) with two world models trained offline.

Why this paper

From IBM Research and 2 others

In one line

Agents learn more effectively by prospectively comparing consequences of candidate actions using counterfactual world models.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (3 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.