Bidirectional start-goal curriculum accelerates sparse-reward reinforcement learning without demonstrations.

Expanding from both the goal and the initial state simultaneously, the method consistently outperforms single-sided curricula on navigation, locomotion, and manipulation tasks.

Big Tech
Juri Pfammatter · Kaixian Qu · Clemens Schwarke · Victor Klemm · Marco Hutter

ETH Zürich · NVIDIA

Research Digest··3 min read
The authors introduce BVER, a curriculum learning algorithm that grows intermediate start states from the goal and intermediate goals from the initial state distribution, biasing both toward unexplored regions.

BVER transfers the bidirectional expansion of RRT-Connect to curriculum generation for goal-conditioned RL.

Why this paper

From NVIDIA and ETH Zürich

In one line

Bidirectional Voronoi-biased exploration curricula, growing starts from goals and goals from starts, accelerate sparse-reward goal-conditioned reinforcement learning.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.