Research Ideation Evaluation
Methods and benchmarks for evaluating AI systems' ability to generate novel research ideas and recover scientific contributions.
4 papers
Where this stands
The written synthesis of this thread is for subscribers. Subscribe.
How this thread developed
Oct 2026 · Massachusetts Institute of Technology, Hong Kong Polytechnic University
Small test-time-trained guidance model beats adapting large LLM generators for discovery
Proposes test-time RL to train a small guidance model, improving performance on discovery tasks relevant to research ideation evaluation.
1 further paper
Oct 2026 · Carnegie Mellon University, Stanford University
AI agents can predict accurately but rarely produce scientific insights, benchmark shows.
introduces EurekaBench, a benchmark for evaluating AI agents' ability to generate scientific insights and satisfy constraints.
Oct 2026 · Delft University of Technology
Test-set tuning modestly inflates scores but rarely changes model rankings
Provides empirical evidence on the inflation of benchmark scores from test-set tuning, informing evaluation methodology for research ideas.
3 of 4 papers shown