Why Your AI Is Slow: Understanding the Inference KV Cache
中文摘要
AI语言模型利用推理KV缓存避免重复读取对话,提高效率。然而,即使有此优化,长上下文仍会增加基础设施成本。
English Summary
Language models use an inference KV cache to avoid rereading conversations, speeding up AI. However, longer contexts still raise infrastructure costs despite this optimization.
Original Excerpt
How language models avoid rereading the entire conversation — and why longer contexts still increase infrastructure costs. Continue reading on Medium »