MiniMax M3 428B · GPU comparison

MiniMax M3 428B — B300 vs GB200 NVL72

Head-to-head AI inference benchmark comparison of B300 (NVIDIA Blackwell) and GB200 NVL72 (NVIDIA Blackwell) on MiniMax M3 428B. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.

B300 posts 6892 tok/s/GPU for $0.10 per million tokens at 66 tok/s/user on MiniMax M3 428B; GB200 NVL72 posts 4050 tok/s/GPU for $0.15. B300 is 60% cheaper per token; B300 delivers 70% more tok/s/GPU.

Throughput at 106 tok/s/user on MiniMax M3 428B: B300 hits 4403 tok/s/GPU, GB200 NVL72 hits 1912. Per-million costs land at $0.15 and $0.32 respectively. B300 is 120% cheaper per token; B300 delivers 130% more tok/s/GPU.

B300 / GB200 NVL72 on MiniMax M3 428B at 147 tok/s/user: 3218 / 811 tok/s/GPU, $0.20 / $0.78 per million tokens. B300 is 288% cheaper per token; B300 delivers 297% more tok/s/GPU. (Numbers reflect the default 8k/1k · fp8 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)

View performance-per-dollar view →

Interpolated from real benchmark data. Edit target interactivity values below to compare at different operating points.
Metric
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Throughput (tok/s/gpu)
B300:6892.0GB200 NVL72:4050.0
B300:4403.3GB200 NVL72:1912.4
B300:3217.5GB200 NVL72:811.5
Cost ($/M tok)
B300:$0.095GB200 NVL72:$0.152
B300:$0.148GB200 NVL72:$0.325
B300:$0.200GB200 NVL72:$0.777
tok/s/MW
B300:3627393GB200 NVL72:2165792
B300:2317538GB200 NVL72:1022690
B300:1693422GB200 NVL72:433944
Concurrency
B300:~50GB200 NVL72:~222
B300:~21GB200 NVL72:~45
B300:~11GB200 NVL72:~13

Inference Performance

Inference performance metrics across different models, hardware configurations, and serving parameters.

Vendor:
Aggregation:
Spec Decoding: