← All threads

Research Ideation Evaluation

Methods and benchmarks for evaluating AI systems' ability to generate novel research ideas and recover scientific contributions.

4 papers

Where this stands

The written synthesis of this thread is for subscribers. Subscribe.

How this thread developed

  1. Oct 2026 · Massachusetts Institute of Technology, Hong Kong Polytechnic University

    Small test-time-trained guidance model beats adapting large LLM generators for discovery

    Proposes test-time RL to train a small guidance model, improving performance on discovery tasks relevant to research ideation evaluation.

  2. 1 further paper

    Oct 2026 · Carnegie Mellon University, Stanford University

    AI agents can predict accurately but rarely produce scientific insights, benchmark shows.

    introduces EurekaBench, a benchmark for evaluating AI agents' ability to generate scientific insights and satisfy constraints.

  3. Oct 2026 · Delft University of Technology

    Test-set tuning modestly inflates scores but rarely changes model rankings

    Provides empirical evidence on the inflation of benchmark scores from test-set tuning, informing evaluation methodology for research ideas.

3 of 4 papers shown