The authors formulate functional web self-refinement as a reinforcement learning problem where the generator produces a page, the critic evaluates it without execution outcomes, and the refiner revises it.
Learning the refinement loop improves web generation beyond final pages
WebLoop jointly optimizes generation, critique, and refinement via execution-grounded reinforcement learning, achieving substantial gains on two benchmarks.
Chinese Tech
Yuxin Meng · Ruixu Zhang · Junjie Wang · Yuhan Suo · Yuhan Sun · Ruining Hu · +7 more
Tsinghua University · Huawei Noah's Ark Lab · East China Normal University · Tongji University · Beihang University
Research Digest··3 min read
The authors introduce WebLoop, a framework that trains a generator, critic, and refiner as roles of a shared policy for functional web generation.
Why this paper
From Huawei Noah's Ark Lab and 6 others
In one line
Learning generation, critique, and refinement jointly in a loop improves functional web generation more than optimizing each role in isolation.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§