MiniMax M2.5/M2.7 — H100 vs MI300X
Head-to-head AI inference benchmark comparison of H100 (NVIDIA Hopper) and MI300X (AMD CDNA 3) on MiniMax M2.5/M2.7. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.
H100 / MI300X on MiniMax M2.5/M2.7 at 54 tok/s/user: 715 / 966 tok/s/chip, $0.45 / $0.27 per million tokens. MI300X is 65% cheaper per token; MI300X delivers 35% more tok/s/chip.
Around the middle of the 41–95 tok/s/user interactivity band, at 68 tok/s/user on MiniMax M2.5/M2.7: H100 runs 482 tok/s/chip at $0.67/M tokens, MI300X runs 604 at $0.44/M. MI300X is 54% cheaper per token; MI300X delivers 25% more tok/s/chip.
Setting 82 tok/s/user as the target on MiniMax M2.5/M2.7, H100 produces 324 tok/s/chip ($1.00 per million tokens) and MI300X produces 351 ($0.75). MI300X is 34% cheaper per token; MI300X delivers 8% more tok/s/chip. (Numbers reflect the default 1k/1k · fp8 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)
| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | H100:714.6MI300X:965.9 | H100:482.0MI300X:603.6 | H100:323.9MI300X:350.7 |
| Cost ($/M tok) | H100:$0.448MI300X:$0.271 | H100:$0.673MI300X:$0.436 | H100:$1.003MI300X:$0.747 |
| tok/s/MW | H100:521623MI300X:694864 | H100:351807MI300X:434251 | H100:236414MI300X:252309 |
| Concurrency | H100:~53MI300X:~38 | H100:~29MI300X:~18 | H100:~16MI300X:~9 |
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.