Training Forecasters to Search Improves Calibration and Reduces Tool Use

Reinforcement learning over historical prediction questions taught a language model to gather evidence more selectively and match larger frontier models at much lower inference cost.

Industry
Yusuf Afifi · Artur Kiulian · Anton Polishko · Mykola Khandoga · Hamudi Naanaa · Alina Krasnobrizha

Future Principle

Research Digest··3 min read
Afifi et al.

The authors built an environment from 2,113 training and 265 held-out Polymarket questions.

Why this paper

From Future Principle · Released code

In one line

Training language models to gather their own evidence during reinforcement learning improves forecasting and beats frontier models at lower cost.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.