返回首页
RadarAI··论文与技术

103μs 降到 18μs 背后,Cursor 为何剑指英伟达,重写 GPU?

中文摘要

Cursor 通过 Pull 调度、Megakernel 及 Ring Buffer,将 MoE 延迟从 103μs 降至 18μs,展现了应用层优化 GPU 内核的路径。

English Summary

Cursor used Pull scheduling, Megakernels, and Ring Buffers to reduce MoE latency from 103μs to 18μs, demonstrating application-layer GPU kernel optimization.

原文节选

📌 一句话摘要 Cursor 通过 Pull 调度、Megakernel 与 Ring Buffer 等创新,将 MoE 通信延迟从 103μs 降至 18μs,展示了应用层深度优化 GPU 内核的路径。 📝 详细摘要 文章详细拆解了 Cursor 开源的 Mixture-of-Kittens (MoK) 如何在 Blackwell 架构的 GB300 NVL72 集群中优化 MoE 的 token 调度与跨 GPU ...