Fisher-Rao prompt curricula make GRPO training more rollout-efficient

ARCUS selects prompts in a statistically natural pass-rate coordinate, improving mathematical reasoning accuracy while using substantially fewer model rollouts.

Top University
Mei Okonkwo · Pixel Nomand · Julian Berg · Elena Voss · Lena Park · Marcus Hale · +2 more

University of Wisconsin–Madison · University of Washington

Research Digest··2 min read
Okonkwo et al.

The authors represent a prompt's pass rate, p, using the arc-length coordinate ψ = arcsin√p on the Fisher-Rao geometry of Bernoulli outcomes.

Why this paper

From University of Washington and University of Wisconsin–Madison

In one line

ARCUS reduces GRPO rollouts by 48-57% and improves accuracy by 2.8-2.9 points using Fisher-Rao arc length for prompt selection.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.