Speculative Decoding: How to Get Free Tokens
中文摘要
探讨推测解码(Speculative Decoding)如何通过解决内存瓶颈来加速LLM生成,提升推理速度。
English Summary
This article explains how Speculative Decoding overcomes memory bottlenecks to accelerate LLM token generation.
Original Excerpt
LLM decoding is slow. Like embarrassingly slow. And if you’ve read my blog before this you know that the reason is memory. Generating one… Continue reading on Medium »