返回首页
AI on Medium··行业媒体

How a 24GB GPU Can Run a 100B LLM

中文摘要

MoE模型通过在内存中保留特定专家并从SSD流式传输权重,使24GB显存能运行千亿级模型。

English Summary

MoE models allow 24GB GPUs to run 100B LLMs by keeping active experts in RAM and streaming weights from SSD memory.

原文节选

MoE models allow us to run only certain experts in ram and stream the relevant weights from SSD memory. Continue reading on The AI Brief »