Flexible action chunks boost world model planning for distant goals

FlexiWorld uses variable-length action chunks and mixed-span goal supervision to achieve 89.29% mean success across four benchmarks, outperforming fixed-chunk baselines.

Big Tech
Shidu Ren · Qilin Gu · Zhenghao Ni · Junhan Sun · Jiaqi Wang · Damien Scieur · +1 more

University of Toronto · Zhejiang University · Tencent Jarvis Lab · Mila & Universite de Montreal · Samsung SAIL

Research Digest··3 min read
The authors introduce FlexiWorld, a JEPA-based world model that jointly learns latent prediction and goal-conditioned action generation from action chunks of varying length and goal spans.

FlexiWorld builds on Joint Embedding Predictive Architectures (JEPAs) that predict future latent states without pixel reconstruction.

Why this paper

From Samsung SAIL and 5 others

In one line

FlexiWorld achieves 89.29% mean success via variable-length action chunks and mixed-span goal supervision.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.