The authors propose DirectSpeech2LLM, which connects a pretrained speech foundation model (SFM) directly to a frozen LLM without any projection module or adapter.
Simple training method stops speech-LLMs from ignoring new instructions
DirectSpeech2LLM preserves LLM generalization to unseen tasks after training only on ASR data.
Big Tech
Hemant Yadav · Sunayana Sitaram · Roger Zimmermann · Rajiv Ratn Shah
IIIT Delhi · Microsoft Research · National University of Singapore
Research Digest··3 min read
The authors present DirectSpeech2LLM, an end-to-end framework that uses a distance-based CTC loss to align speech embeddings with the LLM's input space along geometric and temporal dimensions.
Why this paper
From Microsoft Research and 2 others
In one line
DirectSpeech2LLM preserves an LLM's instruction-following ability on unseen speech tasks without task-specific training.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§