MoE 模型推理的计算与数据搬运:从 Prefill 到 Decode 的四阶段解析
中文摘要
SemiAnalysis将MoE推理分为Prefill、Midfill、Decode attention及Decode experts四阶段,并指出各阶段计算、内存与网络需求差异显著。
English Summary
SemiAnalysis breaks MoE inference into four stages—Prefill, Midfill, Decode attention, and Decode experts—highlighting their distinct computational, memory, and networking requirements.
原文节选
SemiAnalysis 长文拆解 MoE 前沿模型的推理流程,将其分为 Prefill、Midfill、Decode attention、Decode experts 四个阶段,指出四者对计算、内存与网络的需求差异显著。