Back to Home
AI on Medium··Industry Media

My “Faster” Local LLM Streams More Data Per Token, Not Less. I Was Wrong About Why.

中文摘要

修剪20%的专家使LLM提速40-70%。提速并非源于减少PCIe数据传输,而是因为每个token传输了更多数据。

English Summary

Pruning 20% of experts boosted LLM speed by 40-70%. The speedup comes from streaming more data per token, not from reducing PCIe data transfer.

Original Excerpt

I pruned 20% of my model’s experts and it ran 40–70% faster. I told everyone it was because a smaller model moves less data over the PCIe… Continue reading on Medium »