Chasing the Batch-1 Latency Floor: Ling-3.0-flash at 0.78 ms TPOT on 4× B200
中文摘要
Ling-3.0-flash在4块B200上实现0.78毫秒TPOT,专注于用户真实体验的单令牌延迟,而非批量基准测试。
English Summary
Ling-3.0-flash achieved an ultra-low 0.78 ms TPOT on 4x B200 GPUs, prioritizing real-user single-token latency over batching benchmarks.
Original Excerpt
Most inference benchmarks reward batching. Real users do not experience a batch. They experience the next token. Continue reading on Medium »