Shipping LLMs (Part 3/6): Speculative Decoding vs Quantization
中文摘要
量化解决内存带宽,投机解码优化自回归。按此顺序叠加两者,可将大模型推理成本降低3至4倍。
English Summary
Quantization optimizes memory bandwidth and speculative decoding accelerates autoregression; stack them in order to reduce LLM inference costs by 3–4x.
原文节选
Quantization fixes memory bandwidth. Speculative decoding fixes autoregression. Stack them for 3–4x cheaper LLM inference, in this order. Continue reading on Medium »