I Ran Every Local LLM on a 24GB GPU So You Don’t Have To
中文摘要
作者在 24GB GPU 上测试了多种本地大语言模型,旨在评估性能并为读者节省测试精力。
English Summary
The author tested various local LLMs on a 24GB GPU to evaluate performance, saving others the effort of testing them personally.
原文节选
It was 2 AM on a Tuesday and I was watching a 7B parameter model generate a three-line email for 90 seconds. The fan on my RTX 3090 was… Continue reading on Medium »