Google 论文:小量化 LLM 冷启动瓶颈在模型加载
中文摘要
Google研究表明,小量化LLM冷启动延迟主要源于模型加载,增加内存可显著提升性能。
English Summary
Google finds model loading causes most cold start latency for small quantized LLMs; increasing memory can significantly boost performance.
原文节选
Google 新论文指出,对于部署在无服务器 CPU 上的小型量化 LLM,55-70% 的冷启动延迟仅仅来自模型加载。也就是说,瓶颈往往在于搬运模型权重,而非生成 token。 推理本身的问题,还不如把模型加载进内存来得大。 给同一个模型分配 8 GB 的 Cloud Run 内存而非 4 GB,可解锁约 2 倍的 CPU,使热推理时间几乎减半。