The authors introduce Flow-in-Stage Transformer, or FiST, which repeatedly applies a shared Transformer to move a noisy latent representation toward a generated image latent.
Three-stage latent generator reaches 1.11 FID on ImageNet
FiST combines teacher distillation with end-to-end adversarial refinement to generate class-conditional 256-pixel images in two or three stages.
Chinese Tech
Xiaoran Xu · Yujing Wang
ByteDance
Research Digest··3 min read
Xu and Wang train a few-step image generator by first imitating a pretrained teacher, then optimizing the entire sequence of latent transitions against real ImageNet data.
Why this paper
From ByteDance
In one line
A distill-then-refine latent rollout approach with a staged Transformer achieves FID 1.11 on ImageNet 256x256 without Fréchet losses.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§