LLM Inference Cost on H100 in 2026: What It Actually Costs to Serve a Model at Scale
中文摘要
H100服务Llama 4 70B模型,单流95 token/s,批量可达380。文章预测2026年推理成本,提及每小时2.85美元。
English Summary
H100 inference for Llama 4 70B yields 95 tokens/s (single) or 380/s (batched). A Medium article details 2026 costs, noting a $2.85 hourly rate.
原文节选
An H100 running Llama 4 70B single-stream generates roughly 95 tokens a second. Batched to 8 concurrent requests, 380. At a $2.85 hourly… Continue reading on Medium »