Inference Cost Overview
Every active model across MI355X, B200, B300, GB200 and GB300 at a glance.
GPU rental cost / 1M output tokens · 8K→1K · @30 tok/s/user · ↓ lower is better
Database snapshot through Jul 30
| Model | B200 · Reference | MI355X | B300 | GB200 | GB300 | Details |
|---|---|---|---|---|---|---|
DeepSeek V4 Pro 1.6T | Approximately $3.34. Estimated from validated benchmark runs.Jul 12 vLLM · FP4 | SGLang · FP4 | vLLM · FP4 | Dynamo vLLM · FP4 | Dynamo SGLang · FP4 | View details |
Kimi K3 2.8T | No 8K/1K results | View details | ||||
Kimi K2.5/2.6/2.7-Code 1T | no exact @30 result | Approximately $2.04. Estimated from validated benchmark runs.Jul 5 vLLM · FP4 · Standard decode | Approximately $1.96. Estimated from validated benchmark runs.Jun 7 vLLM · FP4 · Standard decode | no exact @30 result | Approximately $0.62. Estimated from validated 16P+8D and 32P+24D runs.Jun 20 Dynamo vLLM · FP4 · Standard decode | View details |
MiniMax M3 428B | Approximately $0.87. Estimated from validated benchmark runs.Jul 29 vLLM · FP8 | vLLM · FP4 | vLLM · FP4 | Dynamo vLLM · FP8 · Standard decode | Dynamo vLLM · FP8 · Standard decode | View details |
GLM5.2 | No 8K/1K results | View details | ||||
Qwen3.5 397B | Approximately $0.86. Estimated from validated benchmark runs.Jul 5 SGLang · FP8 | SGLang · FP4 | SGLang · FP8 | no exact @30 result | no exact @30 result | View details |
- View detailsB200Approximately $3.34. Estimated from validated benchmark runs.Jul 12vLLM · FP4MI355XSGLang · FP4B300vLLM · FP4GB200Dynamo vLLM · FP4GB300Dynamo SGLang · FP4
Kimi K3 2.8T
No 8K/1K results
View detailsKimi K2.5/2.6/2.7-Code 1T
View detailsB200no exact @30 resultMI355XApproximately $2.04. Estimated from validated benchmark runs.Jul 5vLLM · FP4 · Standard decodeB300Approximately $1.96. Estimated from validated benchmark runs.Jun 7vLLM · FP4 · Standard decodeGB200no exact @30 resultGB300Approximately $0.62. Estimated from validated 16P+8D and 32P+24D runs.Jun 20Dynamo vLLM · FP4 · Standard decodeMiniMax M3 428B
View detailsB200Approximately $0.87. Estimated from validated benchmark runs.Jul 29vLLM · FP8MI355XvLLM · FP4B300vLLM · FP4GB200Dynamo vLLM · FP8 · Standard decodeGB300Dynamo vLLM · FP8 · Standard decodeGLM5.2
No 8K/1K results
View detailsQwen3.5 397B
View detailsB200Approximately $0.86. Estimated from validated benchmark runs.Jul 5SGLang · FP8MI355XSGLang · FP4B300SGLang · FP8GB200no exact @30 resultGB300no exact @30 result
Priority: speculative FP4 → speculative FP8 → standard FP4 → standard FP8.
Cost = 3-yr rental $/GPU/hr ÷ output tok/s per deployed GPU. All percentages compare against B200.
Directional platform comparison: cells pick each platform’s best observed envelope, so dates, engines, precisions and decode methods may differ.
Disaggregated results include both prefill and decode GPUs in the denominator.
∞ = no comparable result
Tier values use the best observed platform serving envelope; ≈ marks estimates between validated runs. No extrapolation.