DASA uses a frozen reference model to obtain activation-gradient feedback, signals indicating how changes to internal activations could reduce training loss.
Unreadable synthetic embeddings can fine-tune language models as effectively
DASA optimized continuous inputs for model-useful updates, matching or exceeding natural-language training data across varied tasks.
Research Lab
Jinhao Zhang · Zeyu Liu · Zicheng Yan · Yunquan Zhang · Daning Cheng · Song Tang
Beijing University of Posts and Telecommunications · Institute of Computing Technology, Chinese Academy of Sciences · University of Science and Technology of China · University of Shanghai for Science and Technology
Research Digest··2 min read
Zhang et al.
Why this paper
From Institute of Computing Technology, Chinese Academy of Sciences and 3 others
In one line
Human-readable text is not necessary for effective LLM fine-tuning; synthetic embeddings work as well or better.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§