Inference Cost Overview
Every active model across MI355X, B200, B300, GB200 and GB300 at a glance.
GPU rental cost / 1M output tokens · 8K→1K · @50 tok/s/user · ↓ lower is better
Database snapshot through Jul 30
| Model | B200 · Reference | MI355X | B300 | GB200 | GB300 | Details |
|---|---|---|---|---|---|---|
DeepSeek V4 Pro 1.6T | Approximately $4.55. Estimated from validated benchmark runs.Jul 12 vLLM · FP4 | MoRI SGLang · FP4 | Dynamo SGLang · FP4 | Approximately $4.55. Estimated from validated 8P+8D and 8P+32D runs.About the same cost as B200May 19 Dynamo vLLM · FP4 | Dynamo SGLang · FP4 | View details |
Kimi K3 2.8T | No 8K/1K results | View details | ||||
Kimi K2.5/2.6/2.7-Code 1T | Approximately $2.31. Estimated from validated benchmark runs.Jun 7 vLLM · FP4 · Standard decode | vLLM · FP4 · Standard decode | vLLM · FP4 · Standard decode | Dynamo vLLM · FP4 · Standard decode | Dynamo vLLM · FP4 · Standard decode | View details |
MiniMax M3 428B | Approximately $1.01. Estimated from validated benchmark runs.Jul 29 vLLM · FP8 | vLLM · FP4 | vLLM · FP4 | Approximately $1.21. Estimated from validated 8P+16D and 4P+16D runs.19% more expensive than B200Jul 29 Dynamo vLLM · FP8 · Standard decode | Approximately $1.23. Estimated from validated 4P+8D and 6P+16D runs.21% more expensive than B200Jul 29 Dynamo vLLM · FP8 · Standard decode | View details |
GLM5.2 | No 8K/1K results | View details | ||||
Qwen3.5 397B | Approximately $0.62. Estimated from validated benchmark runs.Jul 5 SGLang · FP4 | SGLang · FP4 | SGLang · FP4 | Dynamo SGLang · FP8 | Dynamo SGLang · FP4 · Standard decode | View details |
- View detailsB200Approximately $4.55. Estimated from validated benchmark runs.Jul 12vLLM · FP4MI355XMoRI SGLang · FP4B300Dynamo SGLang · FP4GB200Approximately $4.55. Estimated from validated 8P+8D and 8P+32D runs.About the same cost as B200May 19Dynamo vLLM · FP4GB300Dynamo SGLang · FP4
Kimi K3 2.8T
No 8K/1K results
View detailsKimi K2.5/2.6/2.7-Code 1T
View detailsB200Approximately $2.31. Estimated from validated benchmark runs.Jun 7vLLM · FP4 · Standard decodeMI355XvLLM · FP4 · Standard decodeB300vLLM · FP4 · Standard decodeGB200Dynamo vLLM · FP4 · Standard decodeGB300Dynamo vLLM · FP4 · Standard decodeMiniMax M3 428B
View detailsB200Approximately $1.01. Estimated from validated benchmark runs.Jul 29vLLM · FP8MI355XvLLM · FP4B300vLLM · FP4GB200Approximately $1.21. Estimated from validated 8P+16D and 4P+16D runs.19% more expensive than B200Jul 29Dynamo vLLM · FP8 · Standard decodeGB300Approximately $1.23. Estimated from validated 4P+8D and 6P+16D runs.21% more expensive than B200Jul 29Dynamo vLLM · FP8 · Standard decodeGLM5.2
No 8K/1K results
View detailsQwen3.5 397B
View detailsB200Approximately $0.62. Estimated from validated benchmark runs.Jul 5SGLang · FP4MI355XSGLang · FP4B300SGLang · FP4GB200Dynamo SGLang · FP8GB300Dynamo SGLang · FP4 · Standard decode
Priority: speculative FP4 → speculative FP8 → standard FP4 → standard FP8.
Cost = 3-yr rental $/GPU/hr ÷ output tok/s per deployed GPU. All percentages compare against B200.
Directional platform comparison: cells pick each platform’s best observed envelope, so dates, engines, precisions and decode methods may differ.
Disaggregated results include both prefill and decode GPUs in the denominator.
∞ = no comparable result
Tier values use the best observed platform serving envelope; ≈ marks estimates between validated runs. No extrapolation.