How Do LLMs Run on Limited GPU Memory? The Magic of Quantization
中文摘要
本文探讨量化技术如何让拥有数千亿参数的大型语言模型在仅有4GB或8GB显存的设备上运行。
English Summary
This article explains how quantization enables massive LLMs with billions of parameters to run on devices with limited GPU memory, such as 4GB or 8GB.
原文节选
Seeing folks running billion / trillion parameters model on few gigs of RAM say, 4GB, 8GB etc. Continue reading on Medium »