The authors paired open-weight and commercial language models with open-source, mixed or closed agent frameworks.
Cheap open survey agents evade any single pollution check
Across nine configurations, locally run open agents completed surveys competitively with commercial systems while producing distinct and inconsistently detected traces.
Research Lab
Center for Humans and Machines · Max Planck Institute for Human Development · International Max Planck Research School on Learning, Institutions, and Future Evolution (LIFE) · University of Basel · Adaptive Rationality Center
Research Digest··2 min read
Rilla and colleagues tested nine autonomous agent configurations on an online survey, running each configuration 40 times and evaluating procedural, behavioral and response-based detection checks.
Why this paper
From Max Planck Institute for Human Development and 5 others
In one line
Cheap open-source LLM agents evade detection differently from commercial ones, requiring multilayered detection.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§