I Fine-Tuned an LLM on a 4GB Laptop GPU. Here’s the Roofline Math Big Labs Don’t Publish
中文摘要
作者分享了在4GB显存笔记本上微调大模型的经验,并通过屋顶模型数学分析揭示了容易被忽视的推理瓶颈。
English Summary
The author explains fine-tuning LLMs on a 4GB laptop GPU, using roofline math to reveal memory bandwidth bottlenecks often overlooked by large labs.
原文节选
Why the real inference bottleneck was never the GPU you couldn’t afford — it was the one you didn’t understand. Continue reading on Medium »