Your LLM Isn’t Slow Because of the Model. It’s Slow Because of Physics.
中文摘要
LLM推理慢并非模型问题,而是受限于内存带宽等物理层面的瓶颈。
English Summary
LLM inference is slow due to memory bandwidth and physical constraints, not the model itself.
原文节选
I kept hearing the same phrase over and over again: “LLM inference is memory-bound.” Continue reading on Medium »