The authors designed BudgetPix to retrofit pretrained pixel-space diffusion transformers with adaptive tokenization.
Adaptive tokenization cuts compute for pixel-space diffusion models by up to 75% with no fidelity loss.
Kara et al. propose BudgetPix, a framework that dynamically allocates tokens based on visual complexity, matching full-budget baselines for text-to-image generation at 25% of the original compute.
Big Tech
Ozgur Kara · Yujia Chen · Daniel Watson · David Forsyth · James Matthew Rehg · Wen-Sheng Chu · +1 more
University of Illinois Urbana-Champaign · Google
Research Digest··3 min read
Kara et al.
Why this paper
From Google and University of Illinois Urbana-Champaign
In one line
BudgetPix matches full-compute image fidelity using 25% of the compute by adaptively tokenizing based on visual complexity.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§