Inference Cost Overview
Every active model across MI355X, B200, B300, GB200 and GB300 at a glance.
GPU rental cost / 1M output tokens · 8K→1K · @75 tok/s/user · ↓ lower is better
Database snapshot through Jul 30
| Model | B200 · Reference | MI355X | B300 | GB200 | GB300 | Details |
|---|---|---|---|---|---|---|
DeepSeek V4 Pro 1.6T | Approximately $3.14. Estimated from validated benchmark runs.Jun 12 TRTLLM · FP4 | ATOM¹ · FP4 | Dynamo SGLang · FP4 | Dynamo SGLang · FP4 | Dynamo TRTLLM · FP4 | View details |
Kimi K3 2.8T | No 8K/1K results | View details | ||||
Kimi K2.5/2.6/2.7-Code 1T | Approximately $2.47. Estimated from validated 4P+20D runs.Jul 22 Dynamo TRTLLM · FP4 · Standard decode | ATOM¹ · FP4 · Standard decode | vLLM · FP4 · Standard decode | Dynamo TRTLLM · FP4 · Standard decode | Dynamo TRTLLM · FP4 · Standard decode | View details |
MiniMax M3 428B | Approximately $0.83. Estimated from validated benchmark runs.Jul 29 vLLM · FP4 | ATOM¹ · FP4 | vLLM · FP4 | Dynamo vLLM · FP8 · Standard decode | Approximately $2.90. Estimated from validated 2P+8D and 2P+16D runs.247% more expensive than B200Jul 29 Dynamo vLLM · FP8 · Standard decode | View details |
GLM5.2 | No 8K/1K results | View details | ||||
Qwen3.5 397B | Approximately $0.65. Estimated from validated benchmark runs.Jun 24 TRTLLM · FP4 | SGLang · FP4 | SGLang · FP4 | Approximately $0.76. Estimated from validated 24P+16D and 16P+16D runs.16% more expensive than B200Jul 26 Dynamo SGLang · FP8 | Approximately $0.84. Estimated from validated 28P+16D and 24P+16D runs.28% more expensive than B200Jul 30 Dynamo SGLang · FP8 | View details |
- View detailsB200Approximately $3.14. Estimated from validated benchmark runs.Jun 12TRTLLM · FP4MI355XATOM¹ · FP4B300Dynamo SGLang · FP4GB200Dynamo SGLang · FP4GB300Dynamo TRTLLM · FP4
Kimi K3 2.8T
No 8K/1K results
View detailsKimi K2.5/2.6/2.7-Code 1T
View detailsB200Approximately $2.47. Estimated from validated 4P+20D runs.Jul 22Dynamo TRTLLM · FP4 · Standard decodeMI355XATOM¹ · FP4 · Standard decodeB300vLLM · FP4 · Standard decodeGB200Dynamo TRTLLM · FP4 · Standard decodeGB300Dynamo TRTLLM · FP4 · Standard decodeMiniMax M3 428B
View detailsB200Approximately $0.83. Estimated from validated benchmark runs.Jul 29vLLM · FP4MI355XATOM¹ · FP4B300vLLM · FP4GB200Dynamo vLLM · FP8 · Standard decodeGB300Approximately $2.90. Estimated from validated 2P+8D and 2P+16D runs.247% more expensive than B200Jul 29Dynamo vLLM · FP8 · Standard decodeGLM5.2
No 8K/1K results
View detailsQwen3.5 397B
View detailsB200Approximately $0.65. Estimated from validated benchmark runs.Jun 24TRTLLM · FP4MI355XSGLang · FP4B300SGLang · FP4GB200Approximately $0.76. Estimated from validated 24P+16D and 16P+16D runs.16% more expensive than B200Jul 26Dynamo SGLang · FP8GB300Approximately $0.84. Estimated from validated 28P+16D and 24P+16D runs.28% more expensive than B200Jul 30Dynamo SGLang · FP8
Priority: speculative FP4 → speculative FP8 → standard FP4 → standard FP8.
Cost = 3-yr rental $/GPU/hr ÷ output tok/s per deployed GPU. All percentages compare against B200.
Directional platform comparison: cells pick each platform’s best observed envelope, so dates, engines, precisions and decode methods may differ.
Disaggregated results include both prefill and decode GPUs in the denominator.
∞ = no comparable result
Tier values use the best observed platform serving envelope; ≈ marks estimates between validated runs. No extrapolation.