DeepSeek V4 Pro 1.6T — B300 vs GB300 NVL72
Head-to-head AI inference benchmark comparison of B300 (NVIDIA Blackwell) and GB300 NVL72 (NVIDIA Blackwell) on DeepSeek V4 Pro 1.6T. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.
B300 / GB300 NVL72 on DeepSeek V4 Pro 1.6T at 72 tok/s/user: 6234 / 9384 tok/s/GPU, $0.10 / $0.08 per million tokens. GB300 NVL72 is 31% cheaper per token; GB300 NVL72 delivers 51% more tok/s/GPU.
Around the middle of the 15–244 tok/s/user interactivity band, at 130 tok/s/user on DeepSeek V4 Pro 1.6T: B300 runs 1365 tok/s/GPU at $0.48/M tokens, GB300 NVL72 runs 3474 at $0.22/M. GB300 NVL72 is 124% cheaper per token; GB300 NVL72 delivers 154% more tok/s/GPU.
Setting 187 tok/s/user as the target on DeepSeek V4 Pro 1.6T, B300 produces 518 tok/s/GPU ($1.24 per million tokens) and GB300 NVL72 produces 502 ($1.44). B300 is 16% cheaper per token; B300 delivers 3% more tok/s/GPU. (Numbers reflect the default 8k/1k · fp4 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)
| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/gpu) | B300:6234.0GB300 NVL72:9384.4 | B300:1365.5GB300 NVL72:3473.9 | B300:518.0GB300 NVL72:501.8 |
| Cost ($/M tok) | B300:$0.103GB300 NVL72:$0.078 | B300:$0.485GB300 NVL72:$0.216 | B300:$1.242GB300 NVL72:$1.436 |
| tok/s/MW | B300:3281057GB300 NVL72:4426587 | B300:718681GB300 NVL72:1638613 | B300:272625GB300 NVL72:236721 |
| Concurrency | B300:~315GB300 NVL72:~1163 | B300:~36GB300 NVL72:~288 | B300:~1GB300 NVL72:~13 |
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.