返回首页
Machine Learning Mastery··论文与技术

Measuring Performance of Transformer Inference

中文摘要

中文:本章介绍衡量Transformer模型推理性能的八个部分,包括延迟、内存使用及多GPU/多机性能。

English Summary

English: This chapter outlines eight aspects of measuring Transformer inference performance, covering latency, memory usage, and multi-GPU/multi-machine setups.

原文节选

This chapter is divided into eight parts; they are: • Metrics for LLM Inference • Measuring a Single Request • Warmup and Synchronization • Measuring GPU Work with CUDA Events • Measuring Memory Usage • Measuring Concurrent Requests • Multiple GPUs and Multiple Machines • Cost per Token The most common inference metrics are: • Latency: How long a request takes from start to finish.