RL post-training improves agent consistency but narrows solution coverage

Across agentic benchmarks, base models explored more successful strategies under repeated sampling, while post-trained models performed better on individual attempts.

Top University
Changdae Oh · Qi Zeng · Qi Qi · Andrey Zhmoginov · Deren Lei · Yun He · +4 more

Meta Superintelligence Labs · University of Wisconsin–Madison · NYU · Stanford University

Research Digest··3 min read
The authors compare 14 base and post-trained model pairs on three interactive agent benchmarks, using a lightweight harness to let base language models call tools and act across multiple turns.

The study evaluates model pairs from four families across three agentic benchmarks, producing 42 base versus post-training comparisons.

Why this paper

From NYU and 3 others · Released code

In one line

Post-training RL sharpens LLMs, improving pass@1 but reducing pass@K coverage on agentic tasks.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.