Inference Cost Overview

Every active model across MI355X, B200, B300, GB200 and GB300 at a glance.

GPU rental cost / 1M output tokens · 8K→1K · @75 tok/s/user · ↓ lower is better

Database snapshot through Jul 30

  • DeepSeek V4 Pro 1.6T

    B200
    Approximately $7.15. Estimated from validated benchmark runs.Jul 12
    vLLM · FP4
    MI355X
    Approximately $6.18. Estimated from validated 8P+8D runs.14% cheaper than B200Jul 27
    MoRI SGLang · FP4
    B300
    Approximately $1.51. Estimated from validated 8P+8D and 4P+8D runs.79% cheaper than B200Jul 30
    Dynamo SGLang · FP4
    GB200
    Approximately $0.96. Estimated from validated 40P+16D and 32P+16D runs.87% cheaper than B200Jun 24
    Dynamo SGLang · FP4
    GB300
    Approximately $1.11. Estimated from validated 16P+8D and 8P+8D runs.84% cheaper than B200Jun 3
    Dynamo SGLang · FP4
    View details
  • Kimi K2.5/2.6/2.7-Code 1T

    B200
    Approximately $3.53. Estimated from validated benchmark runs.Jun 7
    vLLM · FP4 · Standard decode
    MI355X
    Approximately $5.87. Estimated from validated benchmark runs.66% more expensive than B200Jul 5
    vLLM · FP4 · Standard decode
    B300
    Approximately $2.91. Estimated from validated benchmark runs.17% cheaper than B200Jun 7
    vLLM · FP4 · Standard decode
    GB200
    Approximately $3.62. Estimated from validated 12P+16D and 4P+16D runs.About the same cost as B200Jun 21
    Dynamo vLLM · FP4 · Standard decode
    GB300
    Approximately $3.57. Estimated from validated 8P+24D and 4P+32D runs.About the same cost as B200Jun 20
    Dynamo vLLM · FP4 · Standard decode
    View details
  • MiniMax M3 428B

    B200
    Approximately $0.83. Estimated from validated benchmark runs.Jul 29
    vLLM · FP4
    MI355X
    Approximately $1.55. Estimated from validated benchmark runs.86% more expensive than B200Jul 2
    vLLM · FP4
    B300
    Approximately $1.03. Estimated from validated benchmark runs.23% more expensive than B200Jul 29
    vLLM · FP4
    GB200
    Approximately $2.38. Estimated from validated 4P+16D runs.185% more expensive than B200Jul 29
    Dynamo vLLM · FP8 · Standard decode
    GB300
    Approximately $2.90. Estimated from validated 2P+8D and 2P+16D runs.247% more expensive than B200Jul 29
    Dynamo vLLM · FP8 · Standard decode
    View details
  • Qwen3.5 397B

    B200
    Approximately $0.78. Estimated from validated benchmark runs.Jul 5
    SGLang · FP4
    MI355X
    Approximately $1.13. Estimated from validated benchmark runs.46% more expensive than B200Jul 16
    SGLang · FP4
    B300
    Approximately $0.74. Estimated from validated benchmark runs.5% cheaper than B200Jul 4
    SGLang · FP4
    GB200
    Approximately $0.76. Estimated from validated 24P+16D and 16P+16D runs.About the same cost as B200Jul 26
    Dynamo SGLang · FP8
    GB300
    Approximately $0.84. Estimated from validated 28P+16D and 24P+16D runs.8% more expensive than B200Jul 30
    Dynamo SGLang · FP8
    View details

Priority: speculative FP4 → speculative FP8 → standard FP4 → standard FP8.

Cost = 3-yr rental $/GPU/hr ÷ output tok/s per deployed GPU. All percentages compare against B200.

Directional platform comparison: cells pick each platform’s best observed envelope, so dates, engines, precisions and decode methods may differ.

Disaggregated results include both prefill and decode GPUs in the denominator.

∞ = no comparable result

Tier values use the best observed platform serving envelope; ≈ marks estimates between validated runs. No extrapolation.