MiniMax M2.5/M2.7 — B200 vs H200
Head-to-head AI inference benchmark comparison of B200 (NVIDIA Blackwell) and H200 (NVIDIA Hopper) on MiniMax M2.5/M2.7. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.
B200 posts 9617 tok/s/chip for $0.05 per million tokens at 43 tok/s/user on MiniMax M2.5/M2.7; H200 posts 2816 tok/s/chip for $0.12. B200 is 140% cheaper per token; B200 delivers 241% more tok/s/chip.
Throughput at 72 tok/s/user on MiniMax M2.5/M2.7: B200 hits 3893 tok/s/chip, H200 hits 1855. Per-million costs land at $0.13 and $0.18 respectively. B200 is 43% cheaper per token; B200 delivers 110% more tok/s/chip.
B200 / H200 on MiniMax M2.5/M2.7 at 102 tok/s/user: 2080 / 1011 tok/s/chip, $0.23 / $0.33 per million tokens. B200 is 42% cheaper per token; B200 delivers 106% more tok/s/chip. (Numbers reflect the default 8k/1k · fp8 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)
| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | B200:9616.9H200:2816.4 | B200:3893.2H200:1855.1 | B200:2080.0H200:1011.3 |
| Cost ($/M tok) | B200:$0.050H200:$0.120 | B200:$0.126H200:$0.181 | B200:$0.232H200:$0.330 |
| tok/s/MW | B200:5623907H200:2055797 | B200:2276751H200:1354090 | B200:1216372H200:738156 |
| Concurrency | B200:~582H200:~30 | B200:~32H200:~12 | B200:~15H200:~5 |
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.