Inference Cost Overview

Every active model across MI355X, B200, B300, GB200 and GB300 at a glance.

GPU rental cost / 1M output tokens · 8K→1K · @30 tok/s/user · ↓ lower is better

Database snapshot through Jul 30

  • DeepSeek V4 Pro 1.6T

    B200
    Approximately $3.34. Estimated from validated benchmark runs.Jul 12
    vLLM · FP4
    MI355X
    Approximately $1.72. Estimated from validated benchmark runs.48% cheaper than B200Jul 14
    SGLang · FP4
    B300
    Approximately $2.53. Estimated from validated benchmark runs.24% cheaper than B200Jul 12
    vLLM · FP4
    GB200
    Approximately $1.80. Estimated from validated 8P+8D runs.46% cheaper than B200May 19
    Dynamo vLLM · FP4
    GB300
    Approximately $0.81. Estimated from validated 32P+8D and 24P+8D runs.76% cheaper than B200Jun 3
    Dynamo SGLang · FP4
    View details
  • Kimi K2.5/2.6/2.7-Code 1T

    B200
    no exact @30 result
    MI355X
    Approximately $2.04. Estimated from validated benchmark runs.Jul 5
    vLLM · FP4 · Standard decode
    B300
    Approximately $1.96. Estimated from validated benchmark runs.Jun 7
    vLLM · FP4 · Standard decode
    GB200
    no exact @30 result
    GB300
    Approximately $0.62. Estimated from validated 16P+8D and 32P+24D runs.Jun 20
    Dynamo vLLM · FP4 · Standard decode
    View details
  • MiniMax M3 428B

    B200
    Approximately $0.87. Estimated from validated benchmark runs.Jul 29
    vLLM · FP8
    MI355X
    Approximately $1.37. Estimated from validated benchmark runs.57% more expensive than B200Jul 2
    vLLM · FP4
    B300
    Approximately $0.64. Estimated from validated benchmark runs.26% cheaper than B200Jul 29
    vLLM · FP4
    GB200
    Approximately $0.66. Estimated from validated 20P+16D and 12P+16D runs.24% cheaper than B200Jul 29
    Dynamo vLLM · FP8 · Standard decode
    GB300
    Approximately $0.63. Estimated from validated 12P+8D and 6P+8D runs.28% cheaper than B200Jul 29
    Dynamo vLLM · FP8 · Standard decode
    View details
  • Qwen3.5 397B

    B200
    Approximately $0.86. Estimated from validated benchmark runs.Jul 5
    SGLang · FP8
    MI355X
    Approximately $0.76. Estimated from validated benchmark runs.12% cheaper than B200Jul 16
    SGLang · FP4
    B300
    Approximately $0.96. Estimated from validated benchmark runs.13% more expensive than B200May 21
    SGLang · FP8
    GB200
    no exact @30 result
    GB300
    no exact @30 result
    View details

Priority: speculative FP4 → speculative FP8 → standard FP4 → standard FP8.

Cost = 3-yr rental $/GPU/hr ÷ output tok/s per deployed GPU. All percentages compare against B200.

Directional platform comparison: cells pick each platform’s best observed envelope, so dates, engines, precisions and decode methods may differ.

Disaggregated results include both prefill and decode GPUs in the denominator.

∞ = no comparable result

Tier values use the best observed platform serving envelope; ≈ marks estimates between validated runs. No extrapolation.