The authors first had humans identify bottlenecks in long-horizon manipulation tasks.
Targeted subtask reinforcement learning improves long-horizon robot manipulation
PARTS trains residual corrections only at recurring bottlenecks, substantially increasing full-task success with limited robot practice and human resetting.
Top University
Sichang Su · Benjamin Yang · Zhiyun Deng · Boyuan Liang · Yip Fun Yeung · Zelin Wang · +1 more
UT Austin · Autel US · UC Berkeley
Research Digest··2 min read
Su et al.
Why this paper
From UC Berkeley and 2 others · Part of Credit Assignment in Agentic RL, now 17 papers
In one line
PARTS fine-tunes a frozen pretrained policy with residual RL on targeted bottleneck subtasks, using learned selectors and success verifiers to improve long-horizon manipulation with minimal human intervention.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§