Inference Cost Overview

Every active model across MI355X, B200, B300, GB200 and GB300 at a glance.

GPU rental cost / 1M output tokens · 8K→1K · @50 tok/s/user · ↓ lower is better

Database snapshot through Jul 30

  • DeepSeek V4 Pro 1.6T

    B200
    Approximately $4.55. Estimated from validated benchmark runs.Jul 12
    vLLM · FP4
    MI355X
    Approximately $2.25. Estimated from validated 8P+8D runs.51% cheaper than B200Jul 27
    MoRI SGLang · FP4
    B300
    Approximately $1.01. Estimated from validated 24P+8D and 16P+8D runs.78% cheaper than B200Jul 30
    Dynamo SGLang · FP4
    GB200
    Approximately $4.55. Estimated from validated 8P+8D and 8P+32D runs.About the same cost as B200May 19
    Dynamo vLLM · FP4
    GB300
    Approximately $0.88. Estimated from validated 24P+8D and 16P+8D runs.81% cheaper than B200Jun 3
    Dynamo SGLang · FP4
    View details
  • Kimi K2.5/2.6/2.7-Code 1T

    B200
    Approximately $2.31. Estimated from validated benchmark runs.Jun 7
    vLLM · FP4 · Standard decode
    MI355X
    Approximately $3.38. Estimated from validated benchmark runs.46% more expensive than B200Jul 5
    vLLM · FP4 · Standard decode
    B300
    Approximately $2.75. Estimated from validated benchmark runs.19% more expensive than B200Jun 7
    vLLM · FP4 · Standard decode
    GB200
    Approximately $1.14. Estimated from validated 12P+16D and 4P+16D runs.51% cheaper than B200Jun 21
    Dynamo vLLM · FP4 · Standard decode
    GB300
    Approximately $1.20. Estimated from validated 12P+16D and 8P+24D runs.48% cheaper than B200Jun 20
    Dynamo vLLM · FP4 · Standard decode
    View details
  • MiniMax M3 428B

    B200
    Approximately $1.01. Estimated from validated benchmark runs.Jul 29
    vLLM · FP8
    MI355X
    Approximately $1.42. Estimated from validated benchmark runs.41% more expensive than B200Jul 2
    vLLM · FP4
    B300
    Approximately $0.78. Estimated from validated benchmark runs.23% cheaper than B200Jul 29
    vLLM · FP4
    GB200
    Approximately $1.21. Estimated from validated 8P+16D and 4P+16D runs.19% more expensive than B200Jul 29
    Dynamo vLLM · FP8 · Standard decode
    GB300
    Approximately $1.23. Estimated from validated 4P+8D and 6P+16D runs.21% more expensive than B200Jul 29
    Dynamo vLLM · FP8 · Standard decode
    View details
  • Qwen3.5 397B

    B200
    Approximately $0.62. Estimated from validated benchmark runs.Jul 5
    SGLang · FP4
    MI355X
    Approximately $0.90. Estimated from validated benchmark runs.46% more expensive than B200Jul 16
    SGLang · FP4
    B300
    Approximately $0.59. Estimated from validated benchmark runs.5% cheaper than B200Jul 4
    SGLang · FP4
    GB200
    Approximately $0.66. Estimated from validated 32P+16D runs.7% more expensive than B200Jul 26
    Dynamo SGLang · FP8
    GB300
    Approximately $0.59. Estimated from validated 24P+16D and 20P+16D runs.5% cheaper than B200Jun 25
    Dynamo SGLang · FP4 · Standard decode
    View details

Priority: speculative FP4 → speculative FP8 → standard FP4 → standard FP8.

Cost = 3-yr rental $/GPU/hr ÷ output tok/s per deployed GPU. All percentages compare against B200.

Directional platform comparison: cells pick each platform’s best observed envelope, so dates, engines, precisions and decode methods may differ.

Disaggregated results include both prefill and decode GPUs in the denominator.

∞ = no comparable result

Tier values use the best observed platform serving envelope; ≈ marks estimates between validated runs. No extrapolation.