Three-stage latent generator reaches 1.11 FID on ImageNet

FiST combines teacher distillation with end-to-end adversarial refinement to generate class-conditional 256-pixel images in two or three stages.

Chinese Tech
Xiaoran Xu · Yujing Wang

ByteDance

Research Digest··3 min read
Xu and Wang train a few-step image generator by first imitating a pretrained teacher, then optimizing the entire sequence of latent transitions against real ImageNet data.

The authors introduce Flow-in-Stage Transformer, or FiST, which repeatedly applies a shared Transformer to move a noisy latent representation toward a generated image latent.

Why this paper

From ByteDance

In one line

A distill-then-refine latent rollout approach with a staged Transformer achieves FID 1.11 on ImageNet 256x256 without Fréchet losses.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.