Adaptive prefill-decode execution and elastic GPU reuse boost agentic RL throughput

PEARL achieves 2.17-2.79x throughput gains by dynamically choosing between colocated and disaggregated inference while reusing idle training GPUs for rollout.

Chinese Tech
Jiaan Zhu · Wei Gao · Youhui Bai · Zewen Jin · Ju Huang · Siran Yang · +3 more

University of Science and Technology of China · Hong Kong University of Science and Technology · Alibaba Group

Research Digest··2 min read
The authors introduce PEARL, a system that coordinates external elastic GPU resources, temporary reuse of idle training GPUs, and adaptive prefill-decode (PD) configuration for asynchronous agentic reinforcement learning.

The authors designed PEARL to address throughput imbalances between the rollout and training stages of agentic RL.

Why this paper

From Alibaba Group and 2 others

In one line

PEARL achieves 2.17 to 2.79 times the throughput of fixed-resource ROLL across LLMs for agentic RL.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.