The authors designed PEARL to address throughput imbalances between the rollout and training stages of agentic RL.
Adaptive prefill-decode execution and elastic GPU reuse boost agentic RL throughput
PEARL achieves 2.17-2.79x throughput gains by dynamically choosing between colocated and disaggregated inference while reusing idle training GPUs for rollout.
Chinese Tech
Jiaan Zhu · Wei Gao · Youhui Bai · Zewen Jin · Ju Huang · Siran Yang · +3 more
University of Science and Technology of China · Hong Kong University of Science and Technology · Alibaba Group
Research Digest··2 min read
The authors introduce PEARL, a system that coordinates external elastic GPU resources, temporary reuse of idle training GPUs, and adaptive prefill-decode (PD) configuration for asynchronous agentic reinforcement learning.
Why this paper
From Alibaba Group and 2 others
In one line
PEARL achieves 2.17 to 2.79 times the throughput of fixed-resource ROLL across LLMs for agentic RL.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§