返回首页
Towards Data Science··行业媒体

Speculative Decoding on CPUs: Nearly 4x Faster Token Generation with DFlash

中文摘要

DFlash在CPU实现推测解码,令牌生成速度提升近4倍,不改模型输出。利用闲置算力,吞吐量达3.92倍。

English Summary

DFlash enables speculative decoding on CPUs, accelerating token generation by nearly 4x without altering model output. Tests show 3.92x throughput with Qwen3.5-9B, utilizing underused CPU compute efficiently.

原文节选

Speculative decoding can turn underused CPU compute into faster token generation, without changing the model's output. In our vLLM tests, DFlash delivered 3.92x the autoregressive throughput with Qwen3.5-9B on Intel Xeon 6 at concurrency 1. We break down where the speedup comes from, explain the acceptance metrics, and show what determines whether speculation pays off. The post Speculative Decoding on CPUs: Nearly 4x Faster Token Generation with DFlash appeared first on Towards Data Science.