Text queries retrieve SAE features for steering diffusion transformers

D-Scope matches target descriptions to visual centroids of activating patches, retrieving decoder directions that steer generation under spatially masked interventions.

Top University
Xinyue Xu · Jiahao Zhang · Lijie Hu · Peter Hase · Hao Wang

Pivotal Research · Mohamed bin Zayed University of Artificial Intelligence · Schmidt Sciences · Stanford University · University of Illinois at Urbana-Champaign

Research Digest··3 min read
The authors introduce D-Scope, a framework that retrieves sparse autoencoder (SAE) decoder directions from text queries by aggregating SigLIP 2 embeddings of highly activating image patches into visual centroids.

The authors propose D-Scope (Diffusion Scope) to connect SAE feature interpretation with generation control in diffusion transformers (DiTs).

Why this paper

From Stanford University and 4 others

In one line

D-Scope retrieves sparse autoencoder decoder directions for diffusion transformers by matching text queries to visual centroids of activating patches, enabling interpretable generation steering.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.