把记忆交给CPU,大模型会变快
中文摘要
将KV缓存移至CPU并让GPU专攻Token生成,可大幅提升大模型运行效率。
English Summary
Offloading memory to the CPU allows the GPU to focus on token generation, significantly accelerating large model performance.
原文节选
让GPU去忙生成Token
中文摘要
将KV缓存移至CPU并让GPU专攻Token生成,可大幅提升大模型运行效率。
English Summary
Offloading memory to the CPU allows the GPU to focus on token generation, significantly accelerating large model performance.
让GPU去忙生成Token