The study evaluates model pairs from four families across three agentic benchmarks, producing 42 base versus post-training comparisons.
RL post-training improves agent consistency but narrows solution coverage
Across agentic benchmarks, base models explored more successful strategies under repeated sampling, while post-trained models performed better on individual attempts.
Top University
Changdae Oh · Qi Zeng · Qi Qi · Andrey Zhmoginov · Deren Lei · Yun He · +4 more
Meta Superintelligence Labs · University of Wisconsin–Madison · NYU · Stanford University
Research Digest··3 min read
The authors compare 14 base and post-trained model pairs on three interactive agent benchmarks, using a lightweight harness to let base language models call tools and act across multiple turns.
Why this paper
From NYU and 3 others · Released code
In one line
Post-training RL sharpens LLMs, improving pass@1 but reducing pass@K coverage on agentic tasks.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§