返回首页
AI on Medium··行业媒体

Why Your AI Is Slow: Understanding the Inference KV Cache

中文摘要

AI语言模型利用推理KV缓存避免重复读取对话,提高效率。然而,即使有此优化,长上下文仍会增加基础设施成本。

English Summary

Language models use an inference KV cache to avoid rereading conversations, speeding up AI. However, longer contexts still raise infrastructure costs despite this optimization.

原文节选

How language models avoid rereading the entire conversation — and why longer contexts still increase infrastructure costs. Continue reading on Medium »