I Ran a 35B AI Model on My 6GB Laptop GPU. Then I Found Something Nobody Warned Me About.
中文摘要
作者在6GB显存笔记本上运行35B模型时遭遇意外崩溃,揭示了常规性能测试中被忽略的技术隐患。
English Summary
Running a 35B model on a 6GB GPU caused an unexpected crash, uncovering technical risks that most local LLM performance guides overlook.
Original Excerpt
Most “run a big LLM locally” posts stop at “it worked, here’s the tokens/sec.” This one doesn’t. I hit a real crash, root-caused it, and… Continue reading on Medium »