返回首页
AI on Medium··行业媒体

LLM Inference Cost on H100 in 2026: What It Actually Costs to Serve a Model at Scale

中文摘要

H100服务Llama 4 70B模型,单流95 token/s,批量可达380。文章预测2026年推理成本,提及每小时2.85美元。

English Summary

H100 inference for Llama 4 70B yields 95 tokens/s (single) or 380/s (batched). A Medium article details 2026 costs, noting a $2.85 hourly rate.

原文节选

An H100 running Llama 4 70B single-stream generates roughly 95 tokens a second. Batched to 8 concurrent requests, 380. At a $2.85 hourly… Continue reading on Medium »