How a 24GB GPU Can Run a 100B LLM
中文摘要
MoE模型通过在内存中保留特定专家并从SSD流式传输权重,使24GB显存能运行千亿级模型。
English Summary
MoE models allow 24GB GPUs to run 100B LLMs by keeping active experts in RAM and streaming weights from SSD memory.
原文节选
MoE models allow us to run only certain experts in ram and stream the relevant weights from SSD memory. Continue reading on The AI Brief »