MiniMax M2.5/M2.7 · Chip comparison

MiniMax M2.5/M2.7 — H100 vs MI300X

Head-to-head AI inference benchmark comparison of H100 (NVIDIA Hopper) and MI300X (AMD CDNA 3) on MiniMax M2.5/M2.7. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.

H100 / MI300X on MiniMax M2.5/M2.7 at 54 tok/s/user: 715 / 966 tok/s/chip, $0.45 / $0.27 per million tokens. MI300X is 65% cheaper per token; MI300X delivers 35% more tok/s/chip.

Around the middle of the 41–95 tok/s/user interactivity band, at 68 tok/s/user on MiniMax M2.5/M2.7: H100 runs 482 tok/s/chip at $0.67/M tokens, MI300X runs 604 at $0.44/M. MI300X is 54% cheaper per token; MI300X delivers 25% more tok/s/chip.

Setting 82 tok/s/user as the target on MiniMax M2.5/M2.7, H100 produces 324 tok/s/chip ($1.00 per million tokens) and MI300X produces 351 ($0.75). MI300X is 34% cheaper per token; MI300X delivers 8% more tok/s/chip. (Numbers reflect the default 1k/1k · fp8 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)

View performance-per-dollar view →

Interpolated from real benchmark data. Edit target interactivity values below to compare at different operating points.
Metric
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Throughput (tok/s/chip)
H100:714.6MI300X:965.9
H100:482.0MI300X:603.6
H100:323.9MI300X:350.7
Cost ($/M tok)
H100:$0.448MI300X:$0.271
H100:$0.673MI300X:$0.436
H100:$1.003MI300X:$0.747
tok/s/MW
H100:521623MI300X:694864
H100:351807MI300X:434251
H100:236414MI300X:252309
Concurrency
H100:~53MI300X:~38
H100:~29MI300X:~18
H100:~16MI300X:~9

Inference Performance

Inference performance metrics across different models, hardware configurations, and serving parameters.

Vendor:
Deployment:
Spec Decoding: