Inference Cost Overview

Every active model across MI355X, B200, B300, GB200 and GB300 at a glance.

GPU rental cost / 1M output tokens · 8K→1K · @75 tok/s/user · ↓ lower is better

Database snapshot through Jul 30

  • DeepSeek V4 Pro 1.6T

    B200
    Approximately $3.14. Estimated from validated benchmark runs.Jun 12
    TRTLLM · FP4
    MI355X
    Approximately $4.18. Estimated from validated benchmark runs.33% more expensive than B200Jul 28
    ATOM¹ · FP4
    B300
    Approximately $1.51. Estimated from validated 8P+8D and 4P+8D runs.52% cheaper than B200Jul 30
    Dynamo SGLang · FP4
    GB200
    Approximately $0.96. Estimated from validated 40P+16D and 32P+16D runs.69% cheaper than B200Jun 24
    Dynamo SGLang · FP4
    GB300
    Approximately $1.10. Estimated from validated 40P+16D and 24P+16D runs.65% cheaper than B200Jun 15
    Dynamo TRTLLM · FP4
    View details
  • Kimi K2.5/2.6/2.7-Code 1T

    B200
    Approximately $2.47. Estimated from validated 4P+20D runs.Jul 22
    Dynamo TRTLLM · FP4 · Standard decode
    MI355X
    Approximately $2.35. Estimated from validated benchmark runs.About the same cost as B200Jul 12
    ATOM¹ · FP4 · Standard decode
    B300
    Approximately $2.91. Estimated from validated benchmark runs.18% more expensive than B200Jun 7
    vLLM · FP4 · Standard decode
    GB200
    Approximately $1.26. Estimated from validated 16P+32D and 8P+32D runs.49% cheaper than B200Jun 20
    Dynamo TRTLLM · FP4 · Standard decode
    GB300
    Approximately $1.38. Estimated from validated 12P+32D and 8P+32D runs.44% cheaper than B200Jun 22
    Dynamo TRTLLM · FP4 · Standard decode
    View details
  • MiniMax M3 428B

    B200
    Approximately $0.83. Estimated from validated benchmark runs.Jul 29
    vLLM · FP4
    MI355X
    Approximately $0.80. Estimated from validated benchmark runs.About the same cost as B200Jul 3
    ATOM¹ · FP4
    B300
    Approximately $1.03. Estimated from validated benchmark runs.23% more expensive than B200Jul 29
    vLLM · FP4
    GB200
    Approximately $2.38. Estimated from validated 4P+16D runs.185% more expensive than B200Jul 29
    Dynamo vLLM · FP8 · Standard decode
    GB300
    Approximately $2.90. Estimated from validated 2P+8D and 2P+16D runs.247% more expensive than B200Jul 29
    Dynamo vLLM · FP8 · Standard decode
    View details
  • Qwen3.5 397B

    B200
    Approximately $0.65. Estimated from validated benchmark runs.Jun 24
    TRTLLM · FP4
    MI355X
    Approximately $1.13. Estimated from validated benchmark runs.72% more expensive than B200Jul 16
    SGLang · FP4
    B300
    Approximately $0.74. Estimated from validated benchmark runs.12% more expensive than B200Jul 4
    SGLang · FP4
    GB200
    Approximately $0.76. Estimated from validated 24P+16D and 16P+16D runs.16% more expensive than B200Jul 26
    Dynamo SGLang · FP8
    GB300
    Approximately $0.84. Estimated from validated 28P+16D and 24P+16D runs.28% more expensive than B200Jul 30
    Dynamo SGLang · FP8
    View details

Priority: speculative FP4 → speculative FP8 → standard FP4 → standard FP8.

Cost = 3-yr rental $/GPU/hr ÷ output tok/s per deployed GPU. All percentages compare against B200.

Directional platform comparison: cells pick each platform’s best observed envelope, so dates, engines, precisions and decode methods may differ.

Disaggregated results include both prefill and decode GPUs in the denominator.

∞ = no comparable result

Tier values use the best observed platform serving envelope; ≈ marks estimates between validated runs. No extrapolation.