MiniMax M3 428B — B300 vs GB200 NVL72
Head-to-head AI inference benchmark comparison of B300 (NVIDIA Blackwell) and GB200 NVL72 (NVIDIA Blackwell) on MiniMax M3 428B. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.
B300 posts 6892 tok/s/GPU for $0.10 per million tokens at 66 tok/s/user on MiniMax M3 428B; GB200 NVL72 posts 4050 tok/s/GPU for $0.15. B300 is 60% cheaper per token; B300 delivers 70% more tok/s/GPU.
Throughput at 106 tok/s/user on MiniMax M3 428B: B300 hits 4403 tok/s/GPU, GB200 NVL72 hits 1912. Per-million costs land at $0.15 and $0.32 respectively. B300 is 120% cheaper per token; B300 delivers 130% more tok/s/GPU.
B300 / GB200 NVL72 on MiniMax M3 428B at 147 tok/s/user: 3218 / 811 tok/s/GPU, $0.20 / $0.78 per million tokens. B300 is 288% cheaper per token; B300 delivers 297% more tok/s/GPU. (Numbers reflect the default 8k/1k · fp8 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)
| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/gpu) | B300:6892.0GB200 NVL72:4050.0 | B300:4403.3GB200 NVL72:1912.4 | B300:3217.5GB200 NVL72:811.5 |
| Cost ($/M tok) | B300:$0.095GB200 NVL72:$0.152 | B300:$0.148GB200 NVL72:$0.325 | B300:$0.200GB200 NVL72:$0.777 |
| tok/s/MW | B300:3627393GB200 NVL72:2165792 | B300:2317538GB200 NVL72:1022690 | B300:1693422GB200 NVL72:433944 |
| Concurrency | B300:~50GB200 NVL72:~222 | B300:~21GB200 NVL72:~45 | B300:~11GB200 NVL72:~13 |
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.