← All threads

Speculative Decoding Efficiency

Methods for improving the speed and memory efficiency of speculative decoding in language models, especially at long sequence lengths.

1 papers

Where this stands

The written synthesis of this thread is for subscribers. Subscribe.