The authors first strengthen an existing concept-guided sampling pipeline by generating all concepts in a single autoregressive trajectory (rather than iteratively), increasing diversity.
Training a concept generator improves LLM reasoning over repeated sampling.
A small model trained with reinforcement learning learns to propose diverse problem-solving concepts, boosting a frozen large model's accuracy on hard math problems.
Big Tech
Ismail Labiad · Matthieu Kowalski · Marc Schoenauer · Rémi Munos · Julia Kempe
Meta FAIR · Université Paris-Saclay · Inria · NYU Courant Institute
Research Digest··2 min read
, hints, strategies) that maximize the downstream success of a frozen, larger answer generator.
Why this paper
From Meta FAIR and 3 others
In one line
Training a small concept generator with reinforcement learning nearly doubles a larger frozen model's reasoning pass@k on hard math problems.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§