返回首页
Towards Data Science··行业媒体

Can an LLM Forget the Right Things?

中文摘要

一款专为机器人设计的LLM推理运行时,支持33ms实时响应,按语义进行KV缓存驱逐,并采用纯手写CUDA编写。

English Summary

A specialized LLM runtime for robotics ensures 33ms real-time deadlines, evicts KV cache by meaning, and is written entirely in hand-written CUDA.

原文节选

Most LLM inference runtimes have no idea a physical deadline exists. This one refuses admission rather than miss a 33ms robot control cycle, evicts KV cache by meaning instead of age, and is written entirely in hand-written CUDA — no cuBLAS, no libtorch. The post Can an LLM Forget the Right Things? appeared first on Towards Data Science.