One-step action generation and asynchronous pipeline achieve real-time world action models

RealtimeWAM uses teacher-anchored consistency distillation and cross-expert pipelining to accelerate Mixture-of-Transformers based world action models by up to 25x on H100 GPUs with less than 1% accuracy loss

Chinese Tech
Chengtao Lv · Jinyang Du · Shuyi Feng · Yang Yong · Shiqiao Gu · Shunzi Yang · +4 more

Nanyang Technological University · Beihang University · Sensetime · Continental Automotive Singapore

Research Digest··2 min read
The authors present RealtimeWAM, a post-training framework that accelerates World Action Models (WAMs) by enabling one-step action generation and overlapping execution of video and action experts.

The authors identify two bottlenecks in Mixture-of-Transformers (MoT) based WAMs: multi-step action denoising (intra-expert iteration) and sequential execution of video and action experts (inter-expert waiting).

Why this paper

From Sensetime and 3 others

In one line

RealtimeWAM enables near-lossless real-time action generation with 25x speedup via one-step denoising and asynchronous pipelining.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.