The authors built an environment from 2,113 training and 265 held-out Polymarket questions.
Training Forecasters to Search Improves Calibration and Reduces Tool Use
Reinforcement learning over historical prediction questions taught a language model to gather evidence more selectively and match larger frontier models at much lower inference cost.
Industry
Yusuf Afifi · Artur Kiulian · Anton Polishko · Mykola Khandoga · Hamudi Naanaa · Alina Krasnobrizha
Future Principle
Research Digest··3 min read
Afifi et al.
Why this paper
From Future Principle · Released code
In one line
Training language models to gather their own evidence during reinforcement learning improves forecasting and beats frontier models at lower cost.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§