The authors identify two bottlenecks in Mixture-of-Transformers (MoT) based WAMs: multi-step action denoising (intra-expert iteration) and sequential execution of video and action experts (inter-expert waiting).
One-step action generation and asynchronous pipeline achieve real-time world action models
RealtimeWAM uses teacher-anchored consistency distillation and cross-expert pipelining to accelerate Mixture-of-Transformers based world action models by up to 25x on H100 GPUs with less than 1% accuracy loss
Chinese Tech
Chengtao Lv · Jinyang Du · Shuyi Feng · Yang Yong · Shiqiao Gu · Shunzi Yang · +4 more
Nanyang Technological University · Beihang University · Sensetime · Continental Automotive Singapore
Research Digest··2 min read
The authors present RealtimeWAM, a post-training framework that accelerates World Action Models (WAMs) by enabling one-step action generation and overlapping execution of video and action experts.
Why this paper
From Sensetime and 3 others
In one line
RealtimeWAM enables near-lossless real-time action generation with 25x speedup via one-step denoising and asynchronous pipelining.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§