返回首页
量子位··国内权威源

把记忆交给CPU,大模型会变快

中文摘要

将KV缓存移至CPU并让GPU专攻Token生成,可大幅提升大模型运行效率。

English Summary

Offloading memory to the CPU allows the GPU to focus on token generation, significantly accelerating large model performance.

原文节选

让GPU去忙生成Token