DeepSeek V4 Pro 1.6T · Chip comparison

DeepSeek V4 Pro 1.6T — GB200 NVL72 vs GB300 NVL72

Head-to-head AI inference benchmark comparison of GB200 NVL72 (NVIDIA Blackwell) and GB300 NVL72 (NVIDIA Blackwell) on DeepSeek V4 Pro 1.6T. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.

At 64 tok/s/user interactivity on DeepSeek V4 Pro 1.6T, GB200 NVL72 delivers 9108 tok/s/chip at $0.06 per million tokens; GB300 NVL72 delivers 9898 tok/s/chip at $0.06. GB200 NVL72 is 14% cheaper per token; GB300 NVL72 delivers 9% more tok/s/chip at this point.

GB200 NVL72 posts 6055 tok/s/chip for $0.08 per million tokens at 111 tok/s/user on DeepSeek V4 Pro 1.6T; GB300 NVL72 posts 6023 tok/s/chip for $0.11. GB200 NVL72 is 26% cheaper per token; throughput per chip is essentially tied.

Throughput at 157 tok/s/user on DeepSeek V4 Pro 1.6T: GB200 NVL72 hits 523 tok/s/chip, GB300 NVL72 hits 965. Per-million costs land at $0.99 and $0.67 respectively. GB300 NVL72 is 49% cheaper per token; GB300 NVL72 delivers 85% more tok/s/chip. (Numbers reflect the default 8k/1k · fp4 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)

View performance-per-dollar view →

Interpolated from real benchmark data. Edit target interactivity values below to compare at different operating points.
Metric
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Throughput (tok/s/chip)
GB200 NVL72:9108.2GB300 NVL72:9897.8
GB200 NVL72:6055.1GB300 NVL72:6023.1
GB200 NVL72:522.7GB300 NVL72:964.6
Cost ($/M tok)
GB200 NVL72:$0.057GB300 NVL72:$0.065
GB200 NVL72:$0.085GB300 NVL72:$0.107
GB200 NVL72:$0.990GB300 NVL72:$0.665
tok/s/MW
GB200 NVL72:4870672GB300 NVL72:4668766
GB200 NVL72:3238016GB300 NVL72:2841097
GB200 NVL72:279500GB300 NVL72:455004
Concurrency
GB200 NVL72:~9943GB300 NVL72:~1033
GB200 NVL72:~1815GB300 NVL72:~615
GB200 NVL72:~41GB300 NVL72:~31

Inference Performance

Inference performance metrics across different models, hardware configurations, and serving parameters.

Vendor:
Deployment:
Spec Decoding: