Inference Cost Overview

Every active model across MI355X, B200, B300, GB200 and GB300 at a glance.

GPU rental cost / 1M output tokens · 8K→1K · @100 tok/s/user · ↓ lower is better

Database snapshot through Jul 30

  • DeepSeek V4 Pro 1.6T

    B200
    Approximately $7.09. Estimated from validated benchmark runs.Jun 12
    TRTLLM · FP4
    MI355X
    Approximately $6.17. Estimated from validated benchmark runs.13% cheaper than B200Jul 28
    ATOM¹ · FP4
    B300
    Approximately $3.62. Estimated from validated 4P+16D and 4P+24D runs.49% cheaper than B200Jul 30
    Dynamo SGLang · FP4
    GB200
    Approximately $1.19. Estimated from validated 24P+16D and 16P+16D runs.83% cheaper than B200Jun 24
    Dynamo SGLang · FP4
    GB300
    Approximately $1.48. Estimated from validated 24P+16D and 32P+32D runs.79% cheaper than B200Jun 15
    Dynamo TRTLLM · FP4
    View details
  • Kimi K2.5/2.6/2.7-Code 1T

    B200
    Approximately $3.98. Estimated from validated 4P+20D runs.Jul 22
    Dynamo TRTLLM · FP4 · Standard decode
    MI355X
    Approximately $3.28. Estimated from validated benchmark runs.18% cheaper than B200Jul 12
    ATOM¹ · FP4 · Standard decode
    B300
    Approximately $3.38. Estimated from validated benchmark runs.15% cheaper than B200Jun 7
    vLLM · FP4 · Standard decode
    GB200
    Approximately $4.72. Estimated from validated 4P+20D runs.19% more expensive than B200Jun 20
    Dynamo TRTLLM · FP4 · Standard decode
    GB300
    Approximately $5.63. Estimated from validated 4P+20D runs.42% more expensive than B200Jun 22
    Dynamo TRTLLM · FP4 · Standard decode
    View details
  • MiniMax M3 428B

    B200
    Approximately $1.12. Estimated from validated benchmark runs.Jul 29
    vLLM · FP4
    MI355X
    Approximately $0.94. Estimated from validated benchmark runs.15% cheaper than B200Jul 3
    ATOM¹ · FP4
    B300
    Approximately $1.00. Estimated from validated 4P+4D and 2P+4D runs.10% cheaper than B200Jul 26
    Dynamo vLLM · FP4
    GB200
    Approximately $3.79. Estimated from validated 4P+16D runs.239% more expensive than B200Jul 29
    Dynamo vLLM · FP8 · Standard decode
    GB300
    Approximately $4.26. Estimated from validated 2P+16D runs.281% more expensive than B200Jul 29
    Dynamo vLLM · FP8 · Standard decode
    View details
  • Qwen3.5 397B

    B200
    Approximately $0.83. Estimated from validated benchmark runs.Jun 24
    TRTLLM · FP4
    MI355X
    Approximately $1.34. Estimated from validated benchmark runs.62% more expensive than B200Jul 16
    SGLang · FP4
    B300
    Approximately $0.92. Estimated from validated benchmark runs.11% more expensive than B200Jul 4
    SGLang · FP4
    GB200
    Approximately $0.98. Estimated from validated 16P+16D and 12P+16D runs.18% more expensive than B200Jul 26
    Dynamo SGLang · FP8
    GB300
    Approximately $1.06. Estimated from validated 16P+16D and 12P+16D runs.28% more expensive than B200Jul 30
    Dynamo SGLang · FP8
    View details

Priority: speculative FP4 → speculative FP8 → standard FP4 → standard FP8.

Cost = 3-yr rental $/GPU/hr ÷ output tok/s per deployed GPU. All percentages compare against B200.

Directional platform comparison: cells pick each platform’s best observed envelope, so dates, engines, precisions and decode methods may differ.

Disaggregated results include both prefill and decode GPUs in the denominator.

∞ = no comparable result

Tier values use the best observed platform serving envelope; ≈ marks estimates between validated runs. No extrapolation.