The authors built 100 human-annotated collaborative tasks and 100 augmented variants in an MMORPG sandbox.
Current LLM agents struggle to sustain long-horizon team collaboration
AgentWorld tests teams of 3 to 20 agents over more than 50 interaction rounds, with the best evaluated model completing only 52.0 percent of tasks.
Top University
OpenAgents · Columbia University · University of Pennsylvania · Seoul National University · Penn State University
Research Digest··2 min read
The authors introduce AgentWorld, an MMORPG-based benchmark designed to isolate collaboration among independently acting LLM agents with asymmetric roles and abilities.
Why this paper
From Columbia University and 4 others
In one line
AgentWorld shows that multi-agent LLMs achieve at most 52% success on long-horizon collaboration tasks with 25-55 rounds.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§