Guided by failures, choosing initial states improves robot learning per supervised episode

Mulligan retries failed and untried object placements each round and, combined with value-based action selection, lifts real-task success by 10 to 34 percentage points over HG-DAgger.

Top University
Lars Ankile · Perry Dong · Rohan Bhowmik · Aneesh Muppidi · David D. Yuan · Shuran Song · +1 more

Stanford University

Research Digest··3 min read
The authors introduce Mulligan, a data-collection strategy for batch-online robot learning that picks each round's initial states from observed failures and untried regions.

Supervised deployment keeps a human in the loop: an operator places objects, takes control when the policy goes off track, and the resulting corrections are added to the training set.

Why this paper

From Stanford University

In one line

Mulligan selects initial states from failures to improve robot learning faster than uniform collection.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.