New benchmark measures whether AI systems can explore, not just recall, unfamiliar rules

ExplorationBench presents two executable sandboxes whose rules contradict everyday knowledge, so success requires active discovery rather than memorized answers.

Chinese Tech

Fudan University · Hunyuan Team · Tencent · Tsinghua University

Research Digest··3 min read
The authors introduce ExplorationBench, a benchmark that evaluates AI systems' ability to explore unknown environments by forming hypotheses, designing experiments, and iterating on results.

ExplorationBench comprises two sandboxes: AlienCode, with 31 discovery targets and 70 tasks, and AlienLogic, with 24 discovery targets and 70 tasks.

Why this paper

From Hunyuan Team and 3 others

In one line

ExplorationBench shows that leading AI systems can learn unfamiliar executable rules through exploration, but their gains vary across runs and can reverse with further exploration.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.