MiniMax M3 428B · GPU comparison

MiniMax M3 428B — B300 vs GB300 NVL72

Head-to-head AI inference benchmark comparison of B300 (NVIDIA Blackwell) and GB300 NVL72 (NVIDIA Blackwell) on MiniMax M3 428B. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.

Throughput at 63 tok/s/user on MiniMax M3 428B: B300 hits 7080 tok/s/GPU, GB300 NVL72 hits 3284. Per-million costs land at $0.09 and $0.22 respectively. B300 is 143% cheaper per token; B300 delivers 116% more tok/s/GPU.

B300 / GB300 NVL72 on MiniMax M3 428B at 106 tok/s/user: 4403 / 2046 tok/s/GPU, $0.15 / $0.37 per million tokens. B300 is 147% cheaper per token; B300 delivers 115% more tok/s/GPU.

Toward the upper edge of the 22–190 tok/s/user interactivity band, at 148 tok/s/user on MiniMax M3 428B: B300 runs 3199 tok/s/GPU at $0.20/M tokens, GB300 NVL72 runs 834 at $0.90/M. B300 is 348% cheaper per token; B300 delivers 284% more tok/s/GPU. (Numbers reflect the default 8k/1k · fp8 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)

View performance-per-dollar view →

Interpolated from real benchmark data. Edit target interactivity values below to compare at different operating points.
Metric
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Throughput (tok/s/gpu)
B300:7079.9GB300 NVL72:3283.8
B300:4403.3GB300 NVL72:2045.8
B300:3198.9GB300 NVL72:833.8
Cost ($/M tok)
B300:$0.092GB300 NVL72:$0.224
B300:$0.148GB300 NVL72:$0.365
B300:$0.202GB300 NVL72:$0.902
tok/s/MW
B300:3726248GB300 NVL72:1548952
B300:2317538GB300 NVL72:964981
B300:1683639GB300 NVL72:393316
Concurrency
B300:~54GB300 NVL72:~128
B300:~21GB300 NVL72:~44
B300:~11GB300 NVL72:~14

Inference Performance

Inference performance metrics across different models, hardware configurations, and serving parameters.

Vendor:
Aggregation:
Spec Decoding: