ExplorationBench comprises two sandboxes: AlienCode, with 31 discovery targets and 70 tasks, and AlienLogic, with 24 discovery targets and 70 tasks.
New benchmark measures whether AI systems can explore, not just recall, unfamiliar rules
ExplorationBench presents two executable sandboxes whose rules contradict everyday knowledge, so success requires active discovery rather than memorized answers.
Chinese Tech
Fudan University · Hunyuan Team · Tencent · Tsinghua University
Research Digest··3 min read
The authors introduce ExplorationBench, a benchmark that evaluates AI systems' ability to explore unknown environments by forming hypotheses, designing experiments, and iterating on results.
Why this paper
From Hunyuan Team and 3 others
In one line
ExplorationBench shows that leading AI systems can learn unfamiliar executable rules through exploration, but their gains vary across runs and can reverse with further exploration.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§