DeepSeek V4 Pro 1.6T · Chip comparison

DeepSeek V4 Pro 1.6T — B300 vs GB300 NVL72

Head-to-head AI inference benchmark comparison of B300 (NVIDIA Blackwell) and GB300 NVL72 (NVIDIA Blackwell) on DeepSeek V4 Pro 1.6T. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.

B300 / GB300 NVL72 on DeepSeek V4 Pro 1.6T at 71 tok/s/user: 2452 / 9507 tok/s/chip, $0.26 / $0.07 per million tokens. GB300 NVL72 is 284% cheaper per token; GB300 NVL72 delivers 288% more tok/s/chip.

Around the middle of the 13–244 tok/s/user interactivity band, at 129 tok/s/user on DeepSeek V4 Pro 1.6T: B300 runs 1242 tok/s/chip at $0.50/M tokens, GB300 NVL72 runs 3507 at $0.19/M. GB300 NVL72 is 169% cheaper per token; GB300 NVL72 delivers 182% more tok/s/chip.

Setting 187 tok/s/user as the target on DeepSeek V4 Pro 1.6T, B300 produces 542 tok/s/chip ($1.19 per million tokens) and GB300 NVL72 produces 502 ($1.25). B300 is 5% cheaper per token; B300 delivers 8% more tok/s/chip. (Numbers reflect the default 8k/1k · fp4 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)

View performance-per-dollar view →

Interpolated from real benchmark data. Edit target interactivity values below to compare at different operating points.
Metric
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Throughput (tok/s/chip)
B300:2451.9GB300 NVL72:9506.8
B300:1241.7GB300 NVL72:3506.7
B300:541.8GB300 NVL72:501.8
Cost ($/M tok)
B300:$0.259GB300 NVL72:$0.067
B300:$0.502GB300 NVL72:$0.187
B300:$1.186GB300 NVL72:$1.250
tok/s/MW
B300:1290465GB300 NVL72:4484352
B300:653523GB300 NVL72:1654126
B300:285168GB300 NVL72:236721
Concurrency
B300:~23GB300 NVL72:~1113
B300:~5GB300 NVL72:~291
B300:~1GB300 NVL72:~13

Inference Performance

Inference performance metrics across different models, hardware configurations, and serving parameters.

Vendor:
Deployment:
Spec Decoding:
B300 vs GB300 NVL72: DeepSeek V4 Pro Inference Benchmark | InferenceX